Interfaze Ships diffusion-gemma-asr-small, an Open-Source Diffusion ASR Model Transcribing Six Languages via DiffusionGemma’s Parallel Denoising Decoder
Interfaze has released an open-source speech recognition model that uses a diffusion process instead of traditional autoregressive methods to transcribe audio. By utilizing a compact adapter on Google's DiffusionGemma, the system supports six languages while determining processing costs based on the number of denoising steps. This approach offers an alternative architecture for multilingual transcription, decoupling performance from standard sequence generation patterns.
Covered by 1 source
- MMarkTechPost↗Michal SutterJul 3