Topic
asr
Researchers Analyze Fine-Tuning Strategies for Children's Speech Recognition Across Age, Gender, and Datasets
A new study from researchers including Sapkota and Narayanan provides a comprehensive analysis of fine-tuning strategies for children's automatic speech recognition (ASR), focusing on cross-dataset, age, and gender generalization. Using the TORGO database, they achieved a 4.65% relative improvement in isolated word recognition and 4.63% in sentence recognition for dysarthric speech by employing a Factorized Time Delay Neural Network (F-TDNN) with pitch features.
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
A new code-mixing guided preference-learning framework improves code-switching automatic speech recognition (ASR) by steering synthetic speech generation toward better language-boundary consistency. Fine-tuning Whisper Large with this approach reduced Mixed Error Rate (MER) from 12.1% to 8.9% on the DevMAN set and from 17.8% to 14.2% on the DevSGE set of the SEAME Mandarin-English conversational corpus.