Artificial Intelligence #kl divergence#attention distillation
StreamKL Delivers up to 43× Speedup in Memory-Efficient Attention Distillation
Researchers propose StreamKL, a fused GPU primitive for Kullback-Leibler divergence in attention distillation. It eliminates quadratic memory materialization, enabling up to 43× and 14× speedups in forward and backward passes, and reduces extra HBM footprint to O(1).
Jun 21, 2026 1 source