iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Llms ›› Cough Regression Benchmark Reveals Trade-Offs in Respiratory Acoustic Foundation Models

Cough Regression Benchmark Reveals Trade-Offs in Respiratory Acoustic Foundation Models

A new benchmark from researchers at NC State evaluates five respiratory acoustic foundation models on cough regression tasks—predicting age, BMI, and disease probability from cough audio. The study reveals that smaller MLP heads often outperform linear probes, but full-MLP heads overfit on small clinical data. HeAR and M2D+Resp achieve near-full performance with only 50 samples, while OPERA models require 400. Cross-dataset transfer is asymmetric, with large diverse datasets generalizing better to small clinical populations.

iG
iGEN Editorial
June 16, 2026
Cough Regression Benchmark Reveals Trade-Offs in Respiratory Acoustic Foundation Models

Respiratory acoustic foundation models (FMs) have demonstrated strong performance in cough classification—determining if a cough is indicative of a disease. However, their ability to predict continuous health quantities from cough audio, such as age, BMI, or disease probability, remains largely unexplored. This regression capability has clinical value in settings where physical measurements are unavailable, enabling passive health monitoring. A new preprint by researchers including Sanap, Mayur, Desikan, Prasanna, and Lobaton, Edgar introduces the first multi-model, multi-target cough regression benchmark to evaluate these models.

The Cough Regression Benchmark

The benchmark assesses five foundation models—OPERA-CT, OPERA-CE, OPERA-GT, HeAR, and M2D+Resp—across six targets (age, BMI, disease probability on multiple datasets) under subject-disjoint protocols. The models are tested on three datasets: Coswara, CIDRZ, and CoughVID. Three types of regression heads are compared: linear probing, a small multi-layer perceptron (MLP-small), and a full MLP.

The study, according to the arXiv preprint, finds that MLP-small beats the mean-predictor baseline on all tasks and outperforms linear probing in 23 of 30 model × task cases. However, full MLP overfits on small clinical data but recovers on larger datasets, revealing a dataset-size × head-capacity trade-off.

Model Performance by Target

Model Best Age MAE (Coswara) Best Age MAE (CIDRZ) Key Observation
HeAR 9.12 yr Excluded due to pretraining overlap Leads within-dataset age regression on Coswara
OPERA-GT (favored over OPERA-CT) Margin within seed variance Generative pretraining advantage from breath to cough
M2D+Resp Near-full performance at N=50 Similar Data-efficient on small samples

According to the authors, HeAR leads within-dataset age regression on Coswara with a mean absolute error (MAE) of 9.12 years. However, HeAR's results on CIDRZ are excluded from headline claims due to possible overlap between HeAR's pretraining data and CIDRZ. OPERA-GT is favored over OPERA-CT on age in all three datasets, with the CIDRZ margin within seed variance, extending a generative-pretraining advantage from breath analysis to cough.

Data Efficiency Across Models

Data efficiency varies significantly. HeAR and M2D+Resp reach near-full performance with only N=50 samples, while OPERA models require N=400 samples to achieve comparable results. This makes HeAR and M2D+Resp particularly attractive for deployment in low-data scenarios, such as emerging outbreak monitoring in under-resourced regions.

Asymmetric Cross-Dataset Transfer

Cross-dataset transfer performance is strongly asymmetric. The study reports that large diverse data generalises to small clinical populations (e.g., CoughVID to CIDRZ yields a negative bias of -0.17 years), but transfer in the opposite direction fails (CIDRZ to Coswara leads to a positive bias of +2.43 years, a 26.6% increase). This highlights the importance of using large, diverse training datasets when building regression models for cough audio.

For enterprise technology decision-makers, these findings have practical implications. When deploying respiratory acoustic AI in clinical or remote monitoring systems, choosing the right foundation model and regression head depends on dataset size and target population. HeAR and M2D+Resp offer data efficiency for small-labelled datasets, while OPERA models may benefit from larger datasets. The asymmetric transfer results underscore the need to match training data to the deployment population.


Sources:

Keep Reading

Recommended Stories

UniSinger: First End-to-End Framework Unifies Song Generation and Singing Voice Conversion Technology

UniSinger: First End-to-End Framework Unifies Song Generation and Singing Voice Conversion

Researchers have introduced UniSinger, the first end-to-end framework that unifies song generation and singing voice conversion with accompaniment co-generation. Built on a multimodal diffusion transformer, it enables zero-shot speaker cloning and fine-grained timbre control across tasks. Experiments demonstrate state-of-the-art performance on both tasks, offering new possibilities for intelligent music production.

June 17, 2026
New EEG Benchmark Promises Standardized Evaluation of Foundation Models Technology

New EEG Benchmark Promises Standardized Evaluation of Foundation Models

A new benchmark called EEG-FM-Bench aims to standardize evaluation of electroencephalography foundation models (EEG-FMs). It integrates 14 datasets across 10 paradigms and provides tools for gradient and representation analysis. Early experiments reveal critical insights about multi-task learning, pre-training efficiency, and model scaling.

June 16, 2026
Ensemble Deep Learning Achieves 99.27% Accuracy in Lemon Leaf Disease Detection Technology

Ensemble Deep Learning Achieves 99.27% Accuracy in Lemon Leaf Disease Detection

A study on arXiv presents an ensemble deep learning approach for classifying lemon leaf diseases, achieving 99.27% accuracy. The method combines InceptionV3 and MobileNetV2 with adversarial training and Grad-CAM visualization, using a dataset of 1,354 images across 9 classes.

June 16, 2026
How Multi-Label Classification and Generative AI Scale User Feedback Analysis Technology

How Multi-Label Classification and Generative AI Scale User Feedback Analysis

A research paper on arXiv details how a major software company used supervised machine learning for multi-label topic classification and generative AI for summarization to efficiently process large volumes of user feedback. The study found that sentiment analysis alone does not reliably indicate user satisfaction, emphasizing the need for explicit satisfaction surveys.

June 16, 2026