A new research paper on arXiv introduces PrefSQA, a method that uses pairwise preference prediction for speech quality assessment, aiming to overcome limitations of traditional mean opinion scores (MOS). According to the paper by Junyi Fan and Donald S Williamson, scalar MOS labels are sensitive to rater variability and differences in listening tests, introducing labeling noise that limits the reliability of MOS prediction. PrefSQA addresses this by having listeners compare signals directly, producing cleaner labels through preference prediction.
The Problem with Mean Opinion Scores
Mean opinion scores are widely used in speech quality assessment, but the paper notes that rater variability and listening test differences can introduce labeling noise. This noise reduces the reliability of MOS prediction. Preference prediction reduces variability because listeners compare signals directly, yielding cleaner labels, the authors explain.
PrefSQA Method
The proposed PrefSQA method incorporates three key components, according to the paper:
- Uncertainty-aware logits: The model accounts for uncertainty in predictions.
- Impairment attention head: A module that focuses on detecting impairments in speech signals.
- Non-matching-reference comparisons: A module that allows comparisons where the reference signal does not match the degraded signal.
The paper studies MOS-free preference prediction and proposes PrefSQA as a method that uses these components to assess speech quality without relying on MOS labels.
Datasets and Experimental Design
The authors use and refine five datasets for their experiments, according to the paper. These include:
- MOS-derived datasets
- Low-noise simulated sets with matching and non-matching content
- Human preference sets
The method is tested on unseen data to evaluate generalization. The paper reports that experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over baselines.
Results and Key Findings
The paper highlights two main findings:
- Value of high-quality preference data: The clear improvements on non-MOS datasets underscore the importance of using high-quality preference data for training.
- Effectiveness of the proposed method: PrefSQA demonstrates clear improvements over baseline methods on datasets with lower noise, validating the design choices.
The authors conclude that their work highlights the critical role of high-quality datasets in speech quality assessment and demonstrates the effectiveness of the proposed method.
Implications for Speech Quality Research
While the paper does not specify commercial applications, the findings have potential implications for voice communication systems, hearing aids, and speech processing pipelines where accurate quality assessment is critical. The move from MOS to pairwise preference could lead to more reliable evaluations.
For technology leaders focused on AI and audio processing, the PrefSQA method represents a step toward more robust quality metrics that better capture human perception, though further validation in real-world scenarios would be needed.