iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition
Home ›› Technology ›› Ai ›› Llms ›› NVMOS: Novel AI Model Predicts Perceptual Quality of Non-Verbal Vocalizations in Speech

NVMOS: Novel AI Model Predicts Perceptual Quality of Non-Verbal Vocalizations in Speech

A new paper on arXiv introduces NVMOS, the first model purpose-built to assess the perceptual quality of non-verbal vocalizations (NVs) such as laughter, sighs, and coughs in speech. The model was trained on a newly constructed NV-MOS dataset with expert ratings and achieves expert-level agreement with human Mean Opinion Scores. Tests on multimodal LLMs like Gemini showed clear inconsistencies, highlighting the need for specialized NV quality assessment.

iG
iGEN Editorial
June 16, 2026
NVMOS: Novel AI Model Predicts Perceptual Quality of Non-Verbal Vocalizations in Speech

Non-verbal vocalizations (NVs) such as laughter, sighs, and coughs carry critical emotional and intentional cues in speech, yet existing AI systems have struggled to assess their perceptual quality. According to a paper published on arXiv by researchers Mai, Jialong, Jinxin, Xing, Xiaofen, Liu, Wencui, Xu, and Xiangmin, the team has developed NVMOS—to their knowledge the first model that can reliably predict the perceptual quality of NV events in speech.

The Gap in Non-Verbal Vocalization Assessment

Existing speech quality assessment methods typically focus on overall naturalness, the authors explain. Meanwhile, non-verbal text-to-speech (TTS) evaluations mainly examine whether a target NV appears with the correct type and position. The perceptual quality of the NV events themselves has remained largely underexplored.

The NV-MOS Dataset

To fill this gap, the researchers constructed an NV-MOS dataset containing outputs from multiple NV-TTS systems as well as naturally occurring NV samples. Ratings were collected from three acoustic experts on a perceptual quality scale, providing a ground-truth reference for model training and evaluation.

NVMOS Model Architecture and Performance

The proposed NVMOS model incorporates a local NV-event focusing module. Experimental results show that NVMOS reaches expert-level or stronger agreement with human Mean Opinion Scores (MOS). The authors state that it is "to our knowledge the first model that can reliably predict the perceptual quality of NV events in speech."

Multimodal LLM Limitations

The study also evaluated audio-capable multimodal large language models such as Gemini. Clear inconsistencies were found between the scores generated by Gemini and the expert ratings. The authors conclude that general-purpose multimodal models cannot reliably replace human judgments for NV quality assessment.

Method Performance vs Human MOS
Human experts (3 raters) Ground truth reference
NVMOS Expert-level or stronger agreement
Gemini (multimodal LLM) Clear inconsistencies

Implications for Enterprise Voice AI

While the paper does not cite specific commercial applications, the ability to objectively assess NV quality is critical for industries deploying synthetic voices—such as customer service chatbots, virtual assistants, and accessibility tools. The study underscores that off-the-shelf multimodal AI models are insufficient for this specialized task, and dedicated models like NVMOS are necessary to ensure realistic, emotionally appropriate synthetic speech.


Sources:

Keep Reading

Recommended Stories

Neural Audio Codecs' Low Frame Rate Degradation Linked to Training Configuration Technology

Neural Audio Codecs' Low Frame Rate Degradation Linked to Training Configuration

A new study by Gichamba and Busogi investigates the mechanisms behind low frame rate degradation in neural audio codecs. The researchers found that a quality cliff at 6.25 Hz is caused by suboptimal training configuration, not by phonemic collisions or codebook saturation. After correcting the training setup, the codecs perform smoothly down to 3.1 Hz and 1.6 Hz, suggesting that low frame rate efficiency gains are more accessible than previously assumed.

June 17, 2026
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Technology

Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop

Relay, a London startup founded by two former Nothing employees, is building the Relay Q, a portable AI microphone for high-fidelity voice dictation, with software debuting now and hardware due in early 2027. The macOS-first app, powered by Google's Gemini models, adds contextual Skills that automate Slack messages and calendar entries. WIRED's hands-on found the transcription workable but less polished than Google's Pixel 11 Rambler feature, and flagged privacy trade-offs from screen-access permissions.

August 27, 2026
AI could unlock $230 billion annually in upstream oil and gas: McKinsey Technology

AI could unlock $230 billion annually in upstream oil and gas: McKinsey

A McKinsey & Company report estimates artificial intelligence could unlock approximately $230 billion in annual value in global upstream oil and gas at full potential, with $65 billion achievable near-term using current technology. Most of the opportunity is concentrated in a small number of use cases, while oilfield services companies face up to $60 billion of revenue exposure.

August 27, 2026
Nvidia sales soar above $96bn as AI data centre buildout accelerates Technology

Nvidia sales soar above $96bn as AI data centre buildout accelerates

Nvidia reported Q2 revenue of $96bn, more than double year-on-year, with data centre revenue up 117% to $89bn. CEO Jensen Huang said AI has reached its inflection point, while the company guided to $108bn in next-quarter revenue.

August 26, 2026