Topic
dysarthric speech
Dysarthric Speech Recognition Improved by 4.65% with F-TDNN Model and Pitch Features
A systematic study by researchers from multiple institutions investigates dysarthric speech recognition using spectral features and acoustic models. The study, published on arXiv, demonstrates that incorporating pitch features and using the Factorized Time Delay Neural Network (F-TDNN) model yields a 4.65% relative improvement in isolated word recognition and a 4.63% relative improvement in sentence recognition for dysarthric speech, compared to previous research.
Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation
A study by Sapkota et al. explores data augmentation techniques for dysarthric automatic speech recognition (ASR) by fine-tuning the end-to-end Wav2Vec2 model. Four methods—Speaking-Rate Modification, Pitch Modification, Formant Modification, and vocal tract Length Perturbation—were tested across severity levels, achieving relative WER reductions of 30.02%, 16.64%, and 15.47% for low, medium, and high severity respectively.