Artificial Intelligence #text-to-speech#style-captioned
How Do Instructions Shape Speech? New Cross-Attribution Method Reveals Style Control in TTS
A research paper introduces cross-attention attribution for style-captioned text-to-speech, adapting the DAAM framework to speech diffusion models. The method extracts per-token heatmaps across layers and steps, analyzing 3,600 combinations to reveal how caption tokens influence waveforms. Key findings include lower temporal variance for style tokens, correlation with F0 and energy, and peak style conditioning in early ODE steps and deep layers.
Jun 20, 2026 2 sources