Topic
speech processing
ArtBoost: Synthetic Data Augmentation Boosts Acoustic-to-Articulatory Inversion with Limited Real Data
A new data augmentation strategy called ArtBoost leverages large-scale speech-mesh datasets from 3D facial animation to improve acoustic-to-articulatory inversion (AAI) models under limited EMA supervision. The method extracts pseudo articulatory trajectories from facial anchors and pre-trains models before fine-tuning on real data, yielding consistent gains in PCC and RMSE across architectures.
MuVAP: New AI Model Predicts Turn-Taking in Multiparty Conversations Using Audio and Video
Researchers introduce MuVAP, a causal multimodal framework that predicts turn-taking in multiparty conversations using monaural audio and a single camera. The model extends Voice Activity Projection by grounding acoustic predictions in face tracks, and a new 31-hour corpus of unedited conversations supports training.
LM-SPT Uses Semantic Distillation to Improve Speech Tokenization for Language Models
A new speech tokenization method called LM-SPT uses semantic speech-resynthesis distillation to better align discrete speech tokens with language models. The approach outperforms previous semantic-enhanced tokenizers on automatic speech recognition and text-to-speech tasks without sacrificing reconstruction fidelity.
AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction
A team of researchers has developed AP-GRPO, a framework that uses anchor-gated phonetic alignment and policy optimization to reconstruct pathological speech from patients with neurodegenerative and neuromotor disorders. The method preserves reliable audible anchors and aligns recovered content with phonetic cues, improving speech reconstruction across four disease conditions.