← Back to Model Beat
Research·Jul 28·all news from July 28, 2026

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

Apple has introduced a new architecture for audio synthesis that enables the generation of expressive speech entirely on device using AFM 3 Core Advanced. By utilizing decoupled temporal depth diffusion transformers, the method reduces memory usage to facilitate real-time performance on local hardware. This advancement allows for high-quality, configurable voice synthesis without the latency or privacy concerns associated with cloud-based processing.

Covered by 2 sources

  • AApple Machine Learning BlogJul 28
  • AarXiv CS.AIDongseong Hwang, Prasanth Yadla, Kaan Elgin, Shifas Padinjaru Veettil, Sivanand Achanta, Dipjyoti Paul, Ramya Rasipuram, Tyler Johnson, Emad Soroush, Chung-Cheng Chiu, Zhifeng ChenJul 28

Related stories

ResearchAdvancing responsible AI across EuropeJul 29 · 24 sourcesResearchAI Investment Boom Faces Reality Check From Markets and RegulatorsJul 29 · 6 sourcesResearchAccelerating scientific discovery with ChatGPT for Academic ResearchersJul 29ResearchMoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action TokenizationJul 28 · 9 sources