Topic
agent-safety
Artificial Intelligence #ai#agents
LedgerAgent: A New Method for Policy-Adherent Tool-Calling AI Agents in Customer Service
Researchers introduce LedgerAgent, an inference-time method that maintains observed task states in a separate ledger and checks policy constraints before tool calls, improving pass^k metrics across four customer-service domains. The approach addresses common failure modes where agents use stale or incorrect information or violate domain policies.
Jun 20, 2026 1 source
Artificial Intelligence #ai#ai agents
CmdNeedle Reveals Widespread Fragility in AI Agent Command Denylists
A research paper introduces CmdNeedle, an LLM-driven pipeline that systematically detects incompleteness in command denylists used by terminal AI agents. Evaluating 1,709 real-world denylists, the study finds that 69.0–98.6% are fragile, meaning they can be bypassed by alternative commands, undermining security.
Jun 16, 2026 1 source