← Back to Model Beat
Open Source·Jul 3·all news from July 3, 2026

Interfaze Ships diffusion-gemma-asr-small, an Open-Source Diffusion ASR Model Transcribing Six Languages via DiffusionGemma’s Parallel Denoising Decoder

Interfaze has released an open-source speech recognition model that uses a diffusion process instead of traditional autoregressive methods to transcribe audio. By utilizing a compact adapter on Google's DiffusionGemma, the system supports six languages while determining processing costs based on the number of denoising steps. This approach offers an alternative architecture for multilingual transcription, decoupling performance from standard sequence generation patterns.

Covered by 1 source

Related stories

Open SourceRun AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilotJul 7Open SourceFrom Hugging Face to Amazon SageMaker Studio in one clickJul 6 · 2 sourcesOpen SourceLeRobot v0.6.0: Imagine, Evaluate, ImproveJul 7Open SourceSpaceX has an AI device prototype, and it sure sounds phone-ishJul 1 · 5 sources