Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale
Researchers have introduced a method called Block Diffusion that enables efficient caching for diffusion language models by processing data in segments rather than all at once. By overcoming the technical limitations that previously prevented bidirectional denoisers from using key-value caches, this approach makes large-scale parallel decoding more computationally practical.
Covered by 1 source
- AarXiv CS.AI↗Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko1d ago