Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Apple has introduced a new architecture for audio synthesis that enables the generation of expressive speech entirely on device using AFM 3 Core Advanced. By utilizing decoupled temporal depth diffusion transformers, the method reduces memory usage to facilitate real-time performance on local hardware. This advancement allows for high-quality, configurable voice synthesis without the latency or privacy concerns associated with cloud-based processing.
Covered by 2 sources
- AApple Machine Learning Blog↗2d ago
- AarXiv CS.AI↗Dongseong Hwang, Prasanth Yadla, Kaan Elgin, Shifas Padinjaru Veettil, Sivanand Achanta, Dipjyoti Paul, Ramya Rasipuram, Tyler Johnson, Emad Soroush, Chung-Cheng Chiu, Zhifeng Chen2d ago