iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition
Home ›› Technology ›› Ai ›› Dysarthric Speech Recognition Improved by 4.65% with F-TDNN Model and Pitch Features

Dysarthric Speech Recognition Improved by 4.65% with F-TDNN Model and Pitch Features

A systematic study by researchers from multiple institutions investigates dysarthric speech recognition using spectral features and acoustic models. The study, published on arXiv, demonstrates that incorporating pitch features and using the Factorized Time Delay Neural Network (F-TDNN) model yields a 4.65% relative improvement in isolated word recognition and a 4.63% relative improvement in sentence recognition for dysarthric speech, compared to previous research.

iG
iGEN Editorial
June 20, 2026
Dysarthric Speech Recognition Improved by 4.65% with F-TDNN Model and Pitch Features

Dysarthric speech, characterized by impaired articulatory precision, poses significant challenges for automatic speech recognition due to pronounced acoustic variability. According to a systematic study published on arXiv, a combination of spectral features and advanced acoustic models can substantially improve recognition performance. The research, conducted by Sapkota, Paban, Kathania, Hemant Kumar, Kurimo, Mikko, Kadiri, Sudarsana Reddy, and Narayanan, Shrikanth, presents a comprehensive investigation of various acoustic features and models tailored to dysarthric speech.

The Challenge of Dysarthric Speech

The primary difficulty in recognizing dysarthric speech arises from impaired articulatory precision, which leads to high acoustic variability. Past research has shown that hybrid Deep Neural Network / Hidden Markov Model (DNN/HMM) sequence discriminative training can improve recognition. Building on this, the current study systematically examined different combinations of acoustic features and acoustic models, focusing on the TORGO database.

Research Methodology

The researchers explored features including spectral characteristics and pitch. They found that incorporating pitch features notably improved recognition performance, especially for sentence recognition tasks. The study implemented methods using the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model. A deliberate selection of the number of overlapping frames between consecutive training example chunks contributed to the improvements.

Key Findings

The experiments demonstrated the potential to enhance F-TDNN performance for dysarthric speech. Compared to previous research, the methods achieved a 4.65% relative improvement in isolated word recognition and a 4.63% relative improvement in sentence recognition. The following table summarizes the improvements:

Task Relative Improvement
Isolated word recognition 4.65%
Sentence recognition 4.63%

These improvements effectively compensate for speech variability, attributed to the deliberate selection of overlapping frames in training.

Implications for Accessibility

The research has implications for healthcare and accessibility technologies. Improved dysarthric speech recognition can enable better voice-controlled systems for individuals with speech impairments, enhancing communication and independence. While this study focuses on the TORGO database, future work could extend to other databases and real-world applications.

Conclusion

This systematic investigation underscores the importance of feature selection and model choice for dysarthric speech recognition. By leveraging pitch features and the F-TDNN model, the research achieves measurable gains that could translate into more robust speech interfaces. Enterprise technology leaders evaluating AI for accessibility may find these developments relevant for inclusive communication tools.


Sources:

Keep Reading

Recommended Stories

AI Is Helping Solve the Genetic Puzzle of Schizophrenia Technology

AI Is Helping Solve the Genetic Puzzle of Schizophrenia

A study published in Nature Genetics used AI-based computational models to analyze data from over 102,000 people, identifying 766 genes associated with schizophrenia, including 641 not found in previous analyses. The research supports the view that schizophrenia arises from a coordinated network of genetic variants, not a single cause.

August 11, 2026
New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models Technology

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models

A new computer vision paper from arXiv investigates how visual tokens are integrated into large language models (LLMs) under two paradigms: in-context prompting and layer-wise injection. The authors find that visual tokens enter the LLM as 'disguised visual context' lacking linguistic structure, then evolve differently depending on the integration architecture. They show that attention allocation alone is insufficient, and performance depends on the quality of visual representations at each layer.

July 8, 2026
Sequential DPO Study Reveals Non-Uniform Forgetting Across Multiple Preference Objectives Technology

Sequential DPO Study Reveals Non-Uniform Forgetting Across Multiple Preference Objectives

A study by Bhandari et al. on sequential Direct Preference Optimization (DPO) finds that later training objectives do not uniformly degrade earlier preferences. Using Llama-3.1-8B-Instruct, the research reveals that forgetting patterns vary from stability to positive transfer depending on objective compatibility and signal strength, offering guidance for multi-objective AI alignment in enterprises.

July 8, 2026
SafeSpec: New Framework Boosts LLM Safety Without Sacrificing Inference Speed Technology

SafeSpec: New Framework Boosts LLM Safety Without Sacrificing Inference Speed

Researchers propose SafeSpec, a safety-aware speculative inference framework that attaches a latent safety head to jointly evaluate semantic validity and safety in a single forward pass. On Qwen3-32B, it reduces attack success rates by 15% while preserving a 2.06x inference speedup on benign workloads, addressing the fundamental incompatibility between existing safety methods and speculative decoding.

June 21, 2026