Music & Audio3 min reading time

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

Apple Research Blog
Read full post
Apple's Siri Expressive Voices use a memory-efficient audio synthesis model called AFM 3 Core Advanced, running on the Apple Matrix Coprocessor. The model employs a novel detokenizer with decoupled temporal and depth processing, enabling real-time, high-fidelity speech synthesis with low memory usage. This architecture improves voice quality and supports customizable voice features on Apple devices.

More in Music & Audio

Music & Audio4 min read

Suno trained its v6 AI music models with help from Warner and BMG

Covered by 5 sources
Music & Audio2 min read

Roland is getting into generative AI music with Melody Flip

Covered by 2 sources
Music & Audio5 min read

Google Brings Lyria 3.5 Music Generation to the Gemini App and API

Covered by 2 sources