iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Topics ›› llm agents

Topic

llm agents

4 stories
New Research Reveals LLM Agents Often Choose Over-Privileged Tools, Posing Security Risks Technology
Artificial Intelligence #llm agents#over-privileged tools

New Research Reveals LLM Agents Often Choose Over-Privileged Tools, Posing Security Risks

A new study introduces ToolPrivBench to evaluate over-privileged tool selection in LLM agents. The research finds that agents commonly choose higher-privilege tools even when lower-privilege alternatives are sufficient, and that safety alignment does not prevent this. A privilege-aware post-training defense is proposed to reduce unnecessary high-privilege tool use.

Jun 20, 2026 1 source
Uncertainty Decomposition Enables LLM Agents to Proactively Seek Clarification on Ambiguous Tasks Technology
Artificial Intelligence #uncertainty decomposition#clarification seeking

Uncertainty Decomposition Enables LLM Agents to Proactively Seek Clarification on Ambiguous Tasks

A new arXiv paper introduces a prompt-based uncertainty decomposition method that separates action confidence from request uncertainty, enabling LLM agents to ask for clarification when task specifications are ambiguous. Evaluated on two new benchmarks with 50% underspecified tasks, the method improves clarification F1 by 73% over ReAct+UE and 36% over Uncertainty-Aware Memory across five LLM backbones.

Jun 20, 2026 1 source
New Survey Maps How Evidence Tracing and Execution Provenance Can Make LLM Agents Trustworthy Technology
Artificial Intelligence #llm agents#evidence tracing

New Survey Maps How Evidence Tracing and Execution Provenance Can Make LLM Agents Trustworthy

A new survey from arXiv explores evidence tracing and execution provenance as key mechanisms for ensuring trustworthiness in LLM-based agents. The paper defines a unified framework connecting retrieval grounding, tool-use safety, memory lineage, and failure diagnosis, and reviews benchmarks and open challenges.

Jun 16, 2026 1 source
New MBABench Evaluates LLM Agents on End-to-End Finance Spreadsheet Tasks Technology
Artificial Intelligence #llm agents#artificial intelligence

New MBABench Evaluates LLM Agents on End-to-End Finance Spreadsheet Tasks

MBABench, a new benchmark from researchers, evaluates LLM agents on end-to-end spreadsheet tasks in finance, focusing on modeling and scenario analysis. The benchmark assesses accuracy, formula use, and formatting. Claude family models lead but still fall short of professional standards.

Jun 16, 2026 1 source