iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Inside the rogue ChatGPT hack of Hugging Face: AI agents operate at superhuman speed but make clumsy mistakes Landstar Expects to Emerge a Winner After Supreme Court’s Montgomery Ruling Widens Broker Liability New Senate bill targets 'chameleon carriers' that reopen to escape penalties Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Inside the rogue ChatGPT hack of Hugging Face: AI agents operate at superhuman speed but make clumsy mistakes Landstar Expects to Emerge a Winner After Supreme Court’s Montgomery Ruling Widens Broker Liability New Senate bill targets 'chameleon carriers' that reopen to escape penalties Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record
Home ›› Technology ›› Ai ›› Llms ›› Evaluator Bias Spreads Like a Contagion in Multi-Agent LLM Systems, New Research Finds

Evaluator Bias Spreads Like a Contagion in Multi-Agent LLM Systems, New Research Finds

A new paper from arXiv introduces 'Contagion Networks,' a formal framework to measure how evaluation biases propagate when large language models serve as evaluators in multi-agent systems. In a controlled experiment using DeepSeek-chat, researchers found consistent bias propagation between agents, and demonstrated that increasing the evaluator committee from one to three agents reduces effective contagion by 72.4%.

iG
iGEN Editorial
June 20, 2026
Evaluator Bias Spreads Like a Contagion in Multi-Agent LLM Systems, New Research Finds

When large language models (LLMs) are used as evaluators in multi-agent systems — for example, in automated quality assurance, supply chain validation, or content moderation — their systematic evaluation biases can spread through the agent network, a new research paper warns. The study, titled "Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems" and authored by Zewen Liu, introduces a formal framework for measuring this phenomenon.

The research, published on arXiv under a Creative Commons license, is directly relevant to enterprise technology leaders deploying LLM-based agents for decision-making tasks where consistency and objectivity are critical.

What Are Contagion Networks?

According to the paper, Contagion Networks provide a mathematical structure to quantify how evaluator biases — systematic preferences or tendencies in scoring or judging — propagate when multiple LLM agents interact. The framework measures this via a Cross-Agent Contagion Matrix, denoted Gamma_N, where N is the number of agents.

The paper defines three propagation regimes determined by the spectral radius rho(Gamma_N): a suppression regime where biases decay, a critical regime, and an amplification regime where biases strengthen.

Experiment and Key Findings

The researchers conducted a controlled 3-agent experiment using DeepSeek-chat with three distinct evaluator bias profiles: structured, balanced, and evidence-based. They measured the Cross-Agent Contagion Matrix Gamma_3 and found that evaluator biases consistently propagated between agents, with contagion coefficients gamma in the range [0.157, 0.352]. This occurred even though all agents used the same underlying model.

Metric Value
Contagion coefficient range (homogeneous model) [0.157, 0.352]
Contagion coefficient from prior cross-model work (MM-EPC) approx 0.85-1.3
Homogeneous vs. cross-model strength 3-5x weaker
Regime for homogeneous agents Suppression
Contagion reduction from committee size k=1 to k=3 72.4%

Notably, the contagion coefficients for homogeneous-model agents were 3-5x weaker than cross-model coefficients reported in prior work (the MM-EPC benchmark, where gamma ≈ 0.85-1.3). This places homogeneous-model systems in the suppression regime, meaning biases are less likely to amplify but still propagate measurably.

Actionable Mitigation: Committee Sizing

The paper offers a concrete, actionable mitigation strategy: increasing the evaluator committee size. Specifically, raising the committee from k=1 to k=3 reduces effective contagion by 72.4%. This finding is critical for enterprise architects designing multi-agent LLM systems, where evaluation tasks — such as code review, contract analysis, or logistics plan scoring — often rely on a single evaluator agent.

Implications for Enterprise Deployments

For CTOs and digital transformation leaders, the research underscores that using multiple, diverse LLM-based evaluators can significantly dampen bias propagation. The open-source Contagion Network experimental framework released by the authors allows organizations to test bias dynamics in their own agent topologies.

While the experiment focused on a single model (DeepSeek-chat), the framework is model-agnostic. Enterprises integrating multiple LLMs — from different vendors or fine-tuned for specific domains — may see stronger contagion effects, as suggested by prior cross-model coefficients.

The study provides a quantitative basis for designing robust multi-agent evaluation systems, a growing area in supply chain automation, trade compliance checks, and other high-stakes enterprise applications where bias can cascade into costly errors.


Sources:

Keep Reading

Recommended Stories

Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability Technology

Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

A research paper proposes a trace-based observability framework for multi-agent LLM systems that diagnoses wasted computation before final evaluation. On 165 GAIA traces, warned failed runs spent 58.1% of tokens after the first warning. A pilot using warnings reduced post-warning token fraction from 0.638 to 0.304, supporting a layered design with cheap online signals and deeper semantic checks.

June 16, 2026
Chinese Open AI Models Rival Silicon Valley, Spark US Policy Backlash Technology

Chinese Open AI Models Rival Silicon Valley, Spark US Policy Backlash

A wave of near-frontier open-source AI models from Chinese labs like Moonshot AI and Alibaba is challenging Silicon Valley's closed-source dominance. The US government has responded with allegations of distillation theft and potential sanctions, while Chinese firms double down on openness to attract global users.

July 22, 2026
China's Moonshot AI claims Kimi K3 can rival OpenAI and Anthropic Technology

China's Moonshot AI claims Kimi K3 can rival OpenAI and Anthropic

Chinese AI startup Moonshot launches Kimi K3, a massive open-source model with 2.8 trillion parameters, claiming it can rival US leaders OpenAI and Anthropic. The model, set for open-source release on July 27, 2026, has topped benchmarks in web interface engineering and triggered sharp stock declines in domestic competitors.

July 17, 2026
Exit-and-Join Dynamics Enable Decentralized Coalition Formation in Multi-Agent Systems Technology

Exit-and-Join Dynamics Enable Decentralized Coalition Formation in Multi-Agent Systems

A new paper by Zhu and Quanyan presents a decentralized dynamical process for coalition formation driven by unilateral exit-and-join decisions. Agents evaluate moves using the Aumann-Dreze value, leading to equilibrium characterizations and stability analysis through Lyapunov representations.

July 8, 2026