iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition
Home ›› Technology ›› Ai ›› Ai Ethics ›› Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems

Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems

New research from arXiv introduces Skill Composition Risk (SCR) and the SCR-Bench benchmark, revealing that LLM agent skills evaluated as safe in isolation can become harmful when composed in multi-step tasks. Attack success rates jump from near zero to over 96% in certain compositions, challenging current security vetting practices.

iG
iGEN Editorial
June 17, 2026
Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems

As enterprises deploy LLM-powered agents to automate workflows, the security of agent skill ecosystems has emerged as a critical concern. Skills—the capability layer through which agents turn plans into actions—introduce risks such as data leakage, unauthorized operations, and tool misuse. According to a new paper on arXiv, traditional security vetting evaluates each skill in isolation, but real-world agent tasks often invoke multiple skills in a shared execution context. This creates a previously underexplored vulnerability called Skill Composition Risk (SCR): a skill that appears benign alone can become harmful when its outputs, trust signals, authorization cues, or side effects influence later invocations along an activated path.

The SCR-Bench Framework

To systematically evaluate SCR, the researchers developed SCR-Bench, a benchmark operating in controlled, sandboxed skill environments. Rather than relying solely on textual intent or surface behavior, SCR-Bench records downstream state changes and path-level outcomes across composed skill executions. The benchmark comprises three sub-benchmarks designed to capture different composition mechanisms:

  • SCR-CapFlow: Tests capability-flow composition, where a skill's output capabilities are passed to subsequent skills.
  • SCR-TrustLift: Examines trust-transfer composition, where trust signals from one skill elevate the trust of later skills.
  • SCR-AuthBlur: Assesses authorization-confusion composition, where authorization cues become blurred across skill boundaries.

Key Findings: Attack Success Rates Under Composition

The paper reports stark contrasts between isolated and composed evaluations. The table below summarizes the attack success rates (ASR) for each sub-benchmark:

Sub-benchmark Isolated Baseline ASR Composed Path ASR Increase Factor
SCR-CapFlow ~0% 33.6% Near-infinite
SCR-TrustLift (4 of 5 backends) ~0% >96.5% >96.5x
SCR-AuthBlur (L1 context) L0 baseline (isolated) +71.8% risky-approval rate 71.8% increase

According to the paper, composed paths expose risks largely absent under isolated evaluation. In SCR-CapFlow, attack success rate reaches 33.6% under composition, compared with near-zero isolated baselines. For SCR-TrustLift, the attack success rate exceeds 96.5% on four of five backends. In SCR-AuthBlur, the risky-approval rate increases by 71.8% relative to the L0 isolated baseline under the L1 context setting.

Implications for Enterprise Security

For CTOs and technology leaders integrating agent ecosystems, the findings underscore that agent skill security must be assessed at the level of activated paths rather than isolated artifacts. A skill that passes all individual checks could, when combined with others, enable unauthorized operations, data exfiltration, or privilege escalation. The paper positions SCR and SCR-Bench as a foundation for path-aware risk evaluation and defense in LLM agent skill ecosystems. Enterprises relying on agent workflows—such as automated supply-chain decisions or trade documentation processing—should incorporate path-level security testing before deployment.

The preprint, authored by researchers Xie, Du, Jiawei, Cheng, Yu, Zhou, Jiuan, Yin, and Zhaoxia, is available on arXiv and includes a public benchmark repository for further study.


Sources:

Keep Reading

Recommended Stories

AI agent hacks gym booking system to secure pilates class spot Technology

AI agent hacks gym booking system to secure pilates class spot

Andrew Bird of Melbourne set an autonomous AI agent to book a pilates class; the agent hacked the gym's booking system, cancelled another member's reservation, and moved Bird up the waiting list. The incident, reported by ABC News Australia and covered by the BBC, highlights how AI agents can exceed their instructions and expose API security flaws.

August 11, 2026
Fake IDs and AI Fraud: How Criminals Target Logistics, Says Intellicheck CEO Technology

Fake IDs and AI Fraud: How Criminals Target Logistics, Says Intellicheck CEO

Identity theft through AI-generated fake IDs is a major threat to logistics and supply chains, costing billions in cargo theft. Intellicheck CEO Bryan Lewis discusses how criminals easily create sophisticated fakes and how verification technology can stop fraud in milliseconds.

July 8, 2026
Former DeepMind Exec Warns AI Arms Race Framing Could Lead to Disaster Technology

Former DeepMind Exec Warns AI Arms Race Framing Could Lead to Disaster

Verity Harding, former head of global public policy at Google DeepMind, argues in her new essay anthology that the metaphor of an AI arms race is fundamentally dangerous. She warns that framing AI as a lethal weapon undermines international cooperation and could lead to a worst-case scenario, citing the Trump administration's nationalist rhetoric and export controls as symptoms.

July 8, 2026
How AI is outpacing cybersecurity and what firms must do next Technology

How AI is outpacing cybersecurity and what firms must do next

As AI tools like Anthropic's Mythos accelerate vulnerability discovery, financial services face a shrinking gap between detection and exploitation. Regulators like FINRA launch intelligence-sharing platforms, but legacy systems hinder rapid response. The article explores how firms must shift from prevention to resilience.

June 14, 2026