Organizations deploying AI agents alongside human workers face a fundamental challenge: more collaborators do not automatically mean better results. In fact, new research shows that without explicit coordination scaffolding, adding humans to AI teams can actually lower performance.
The study, titled Searching for Synergy in Shared Workspace Human-AI Collaboration and posted on arXiv, was conducted by researchers Kotalwar, Nachiket, Das, Rohini, and Rose, Carolyn. It used the Collaborative Gym environment with DiscoveryBench tasks to examine when simulated human-AI teams succeed and when they suffer from process loss.
The Research Setup
The team ran 1,482 sessions of shared-workspace tasks where AI agents and human collaborators had to coordinate responsibilities before submitting a final answer. The environment, Collaborative Gym, is specifically designed for studying human-AI interaction. DiscoveryBench tasks require both automated reasoning and human judgment, reflecting real-world scientific and professional scenarios.
Key Findings on Coordination Overhead
Counterintuitively, adding relevant human collaborators can lower performance when the team lacks structure to coordinate contributions. This phenomenon, termed process loss, turns additional collaborators into coordination overhead. The researchers found that simply increasing the number of participants—even those with relevant expertise—does not guarantee better outcomes. Without clear responsibility signals and routing of expertise, teams underperform relative to smaller or more structured groups.
Scaffolding for Synergy
To address this, the researchers evaluated a scaffolding approach combining shared group memory with simulated human-in-the-loop (HITL) gates. In this setup, selected actions require approval from a designated simulated participant before proceeding. The scaffolding yielded higher mean performance, with the clearest gains observed in three-person teams. The authors attribute this to clearer responsibility signals and stronger routing of expertise to team actions.
| Condition | Mean Performance | Best Team Size | Key Mechanism |
|---|---|---|---|
| No scaffolding | Lower | N/A | Coordinate overhead |
| With shared group memory + HITL gates | Higher | Three-person teams | Clearer responsibility, expertise routing |
The table above summarizes the comparative outcomes reported in the study.
Implications for Enterprise AI Teams
The study's core conclusion carries direct weight for enterprise technology leaders: how human-AI teams coordinate and integrate expertise matters as much as the capability available to them. For supply chain, logistics, and trade technology teams deploying AI for tasks like customs document processing, shipment optimization, or trade finance risk assessment, the findings suggest that organizational design—not just model accuracy—determines ROI. Implementing structured approval workflows and shared memory systems could prevent the coordination overhead observed in unstructured teams. The research provides a empirical foundation for building human-AI teams that truly achieve synergy.