Artificial Intelligence #llms#artificial intelligence
New PhysAssistBench Tests Medical LLMs on Interactive Doctor-Patient-EHR Coordination
Researchers introduce PhysAssistBench, a benchmark for evaluating medical LLMs on interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, it uses a scalable pipeline to create agentic patients. Experiments show leading LLMs remain unreliable, highlighting the need for coordination across knowledge, communication, and systems.
Jun 21, 2026 1 source