Artificial Intelligence #llm safety#red-teaming
New Benchmark Reveals Critical Vulnerabilities in LLM Agents Used for Safety-Critical Systems
A new benchmark called NRT-Bench tests multi-turn red-teaming of LLM agents operating a simulated nuclear power plant. Adaptive attacks cause safety limit breaches in up to 12.1% of sessions, with vulnerabilities nearly disjoint across models.
Jun 20, 2026 1 source