Music & Audio3 min reading time
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Apple Research Blog
Read full postApple's Siri Expressive Voices use a memory-efficient audio synthesis model called AFM 3 Core Advanced, running on the Apple Matrix Coprocessor. The model employs a novel detokenizer with decoupled temporal and depth processing, enabling real-time, high-fidelity speech synthesis with low memory usage. This architecture improves voice quality and supports customizable voice features on Apple devices.




