Topic
ai bias
Evaluator Bias Spreads Like a Contagion in Multi-Agent LLM Systems, New Research Finds
A new paper from arXiv introduces 'Contagion Networks,' a formal framework to measure how evaluation biases propagate when large language models serve as evaluators in multi-agent systems. In a controlled experiment using DeepSeek-chat, researchers found consistent bias propagation between agents, and demonstrated that increasing the evaluator committee from one to three agents reduces effective contagion by 72.4%.
The Scaffold Effect: How Prompt Framing Skews AI Evaluation in Clinical Vision-Language Models
A study on arXiv evaluating 12 open-weight vision-language models (VLMs) on clinical neuroimaging datasets found that up to 58% of apparent multimodal performance gains are due to prompt framing rather than genuine reasoning. The researchers identified a 'scaffold effect' where merely mentioning MRI availability in the task prompt accounts for 70-80% of F1 improvement, even when no imaging data is present. Expert evaluation also revealed fabrication of neuroimaging-grounded justifications, raising concerns about the reliability of VLM evaluations in clinical settings.