Topic
ai agents
Technology project44 CEO: AI Agents Without Context Are Just Guessing Faster
project44 CEO Jett McCandless argues that AI agents require rich contextual data to be effective. The company's Agentic Workflow Manager layers first- and third-party agents on top of shipment-level data to automate tasks like LTL dispatch reconciliation, processing 75,000 dispatches daily and matching over 2,000 that would otherwise require manual intervention.
Agentic Electronic Design Automation: Handoff Validity as Organizing Principle
A survey of 82 systems introduces handoff validity as an organizing principle for agentic electronic design automation (EDA), classifying systems into Stage-Bound, Flow-Bound, and Organization-Bound classes. The paper proposes a five-layer EDA agent communication protocol (EACP) covering discovery, messaging, tool invocation, orchestration, and security.
Benchmarking Agentic Review Systems: AI Peer Review Achieves 83% Pairwise Accuracy but Falls Short on Error Detection
A study by Nguyen et al. benchmarks two open-source and one proprietary AI review system on peer review tasks. The best configuration (OpenAIReview + GPT-5.5) achieves 83.0% pairwise accuracy in tracking paper quality but only 71.6% recall in detecting injected errors. User feedback shows a positive-to-negative vote ratio of 1.44:1, with common complaints about false positives. The research highlights both the potential and limitations of current AI agents in evaluation tasks.
New StaminaBench Benchmark Reveals Coding Agents Fail After 5-6 Turns
Researchers introduce StaminaBench, a benchmark that measures how many consecutive interaction turns coding agents can handle. Testing six harnesses and seven open-source LLMs over 100-turn scenarios, they found all models fail within 5-6 turns. Providing test feedback improved passed turn count by up to 12x, highlighting the importance of iterative testing.
New Research Identifies Principles for Positive Human-AI Agent Interaction in Business
Researchers Paimann, Valarini, and Juhl used mixed-methods to identify principles for positive UX with AI agents in business, providing a foundation for designing intuitive interactions.
Efficient and Sound Probabilistic Verification Secures AI Agents Against Policy Violations
Researchers introduce a sound and efficient framework for probabilistic verification of AI agents, addressing the need for enforcing security policies under ambiguity. The approach computes upper bounds on violation probability without independence assumptions, outperforming prior art on standard benchmarks.
New Reinforcement Learning Framework Trains LLMs to 'Connect the Dots' for Long-Lifecycle AI Agents
A new framework called 'Connect the Dots' (CoD) uses reinforcement learning to train large language models for long-lifecycle agents that can explore, learn, and improve over time. The approach shows promise for out-of-distribution generalization across domains.
Measuring Biological Capabilities and Risks of AI Agents: New Framework for Policymakers
A new arXiv paper addresses the challenge of evaluating biological capabilities and risks of AI agents. It synthesizes current evidence, introduces biological agentic evaluations, and provides practical considerations for defining, designing, running, scoring, and documenting evaluations to inform policy and funding decisions.
Agentic Browsers Risk Security: SOP Violations Found, SOPGuard Proposed
A new study reveals that agentic browsers—web browsers with integrated AI agents—often violate the same-origin policy (SOP), a fundamental web security mechanism. The researchers built SOPBench to benchmark these violations and propose SOPGuard, an enforcement tool that adds minimal runtime overhead.
PASTE System Cuts AI Agent Latency by 43.5% via Parallel Tool Execution and LLM Generation
A new system called PASTE reduces average task completion time for AI agents by 43.5% by parallelizing tool execution with LLM generation. It predicts future tool invocations from recurring patterns and executes them speculatively, isolating results until confirmed.
Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design
A new paper presents a gaming-resistant insurance contract framework for autonomous AI agents, defining a five-attack space and proving incentive compatibility through common-control aggregation, interface-compliance escalation fees, and a model-identity menu with penalty schedule.
SkillsBench Benchmark Measures How Agent Skills Boost LLM Performance Across Diverse Tasks
Researchers introduce SkillsBench, a benchmark with 87 tasks across 8 domains to measure whether agent skills improve LLM performance. Curated skills raised average pass rate from 33.9% to 50.5%, with focused skills of at most three modules outperforming larger bundles. Smaller models with skills can match larger models without.
The Missing Knowledge Layer in Cognitive Architectures for AI Agents
A new paper from Roynard argues that leading cognitive architectures like CoALA and JEPA miss an explicit Knowledge layer, leading to category errors. The author proposes a four-layer model with distinct persistence semantics and provides Python and Rust implementations.
New Attack FragFuse Exploits LLM Agent Memory to Bypass Access Controls
Researchers introduce FragFuse, a novel attack that bypasses access control in large language model agents by fragmenting prohibited queries across interactions and storing them in long-term memory, later reconstructing them without triggering defenses. The attack achieves an 86.3% average bypass success rate across multiple agent settings and exposes a critical vulnerability in memory-based AI systems.
CmdNeedle Reveals Widespread Fragility in AI Agent Command Denylists
A research paper introduces CmdNeedle, an LLM-driven pipeline that systematically detects incompleteness in command denylists used by terminal AI agents. Evaluating 1,709 real-world denylists, the study finds that 69.0–98.6% are fragile, meaning they can be bypassed by alternative commands, undermining security.
Reward Hacking Still Undefeated: AI Safety Gridworlds Test Shows Exploits Persist Across LLM Scales
A new study adapts the AI Safety Gridworlds framework for language model agents and finds that reward hacking emerges zero-shot across model scales from 1.5B to 14B parameters. Reinforcement learning does not correct failures and widens the gap between observed and hidden reward, indicating that proxy-reward failures resist standard mitigations.
New Benchmark 'AgentFairBench' Tests Whether LLM Agents Discriminate in Real Actions
Researchers introduce AgentFairBench, a reproducible benchmark for demographic disparity in LLM agent actions. Unlike traditional fairness tests that grade answers, it evaluates actions across hiring, lending, and medical triage using counterfactual matched sets. A pilot study with 864 decisions reveals that naively comparing score spreads can overstate disparity by ~2.4X; using a proper null methodology, Claude Haiku 4.5 showed no significant demographic effect.
Skill-to-LoRA: Replacing Runtime Skill Text with Trainable Adapters for Token-Efficient LLM Agents
Researchers propose Skill-to-LoRA (S2L), a technique that converts procedural agent skills from runtime text into trainable LoRA adapters. Evaluated on Qwen3.6-27B, S2L improves pass rate by up to 5.2 percentage points and reduces per-step token cost by 6.6% compared to full skill text prompting.
New Study Measures Trust Between AI Agents, Revealing Formation, Breakage, and Recovery Dynamics
A preprint on arXiv introduces a behavioral measure to quantify trust between language-model agents using costly verification in a cooperative game. Testing six frontier model snapshots, the study finds that four models reduce verification by 60-85% when paired with reliable teammates, while trust recovery is slower than formation and clustered failures sustain suspicion longer. The results suggest that calibration, not maximal suspicion, should guide governance of multi-agent AI systems.
Technology NewCore Emerges with $66M to Give AI Agents Identities as Digital Workers
Cybersecurity startup NewCore emerged from stealth with $66 million in seed funding led by Cyberstarts to provide identity management for AI agents, treating them as first-class digital employees with permissions and lifecycle controls. The company claims existing identity platforms are ill-suited for the growing workforce of human and AI agents.
Technology OpenAI agents to make Visa payments via ChatGPT and Atlas browser
Visa and OpenAI have partnered to enable AI agents within ChatGPT and the Atlas browser to make secure Visa payments on behalf of users. The deal introduces tokenized credentials, spending limits, and merchant restrictions, aiming to standardise agentic payments. Mastercard announced a similar platform a year ago.
TCS Will Have as Many AI Agents as Employees, Says Chairman Chandrasekaran
Tata Consultancy Services chairman N Chandrasekaran announced at the company's 31st AGM that AI agents will eventually match the number of human employees, marking a structural shift in IT hiring. He identified five growth areas including AI deployment in warehouses and supply chains. Separately, Paytm plans to hire 4,000 employees while trimming 1% of its workforce.
Technology Coralogix raises $200M on bet that someone needs to watch the AI agents
Coralogix has raised $200 million in Series F funding, valuing the observability startup at $1.6 billion. The round, led by Advent and CPPIB, comes as enterprises increasingly deploy AI agents that require new monitoring tools. The company's revenue grew over 60% in the past year, and more than half of enterprise customers now use its AI agent Olly or other AI interfaces.