iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Maharashtra’s ₹500 crore AI agriculture policy targets data, traceability and farm advisory Commercial LPG prices drop: 19-kg cylinder rate cut by ₹202 in Delhi, ₹209 in Kolkata Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Maharashtra’s ₹500 crore AI agriculture policy targets data, traceability and farm advisory Commercial LPG prices drop: 19-kg cylinder rate cut by ₹202 in Delhi, ₹209 in Kolkata Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue
Home ›› Technology ›› Ai ›› Ai Ethics ›› Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models

Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models

A new method called Safe Trigger leverages the latent safety awareness of Large Reasoning Models to improve safety alignment without external data. Using Supervised Fine-Tuning and Direct Preference Optimization, the approach reduces Attack Success Rate on harmful and jailbreak benchmarks while preserving general performance.

iG
iGEN Editorial
June 16, 2026
Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models

Large Reasoning Models (LRMs) are highly capable at complex tasks, yet remain vulnerable to sophisticated jailbreaks and direct harmful queries. According to a paper on arXiv by Miao, Ke, Li, Jiaxin, Chen, Hongliang, Hu, Yuke, Qin, and Zhan, prior safety alignment methods heavily depend on external manual data annotation. However, the researchers observed that LRMs can inherently identify safety risks when re-presented with original queries alongside their own reasoning trajectories——a capability they term Latent Safety Awareness.

To exploit this, the team proposed a two-stage training approach called Safe Trigger. First, they use Supervised Fine-Tuning (SFT) to explicitly induce safe tags that trigger safety analysis and guidance following the initial reasoning content for unsafe queries. For general queries, standard responses are preserved, ensuring adaptive triggering. Second, they apply Direct Preference Optimization (DPO) to further enhance the correctness and stability of the safety analysis and guidance. Notably, the responses required for both training stages are entirely generated by the models being optimized, eliminating the need for external annotation.

Experimental results demonstrate significant safety enhancement. The Attack Success Rate (ASR) of DeepSeek-R1-Distill-Llama-8B dropped, on average, by 24.65% on harmful benchmarks and by 36.72% on jailbreak benchmarks. The method exerts almost no negative impact on general performance or user experience.

Benchmark Average ASR Reduction
Harmful 24.65%
Jailbreak 36.72%

The paper argues that Safe Trigger method leverages the model's own latent safety awareness, reducing reliance on external data. This approach could be adapted for enterprise AI deployments where safety alignment is critical, such as in supply chain decision-support systems or customer-facing logistics chatbots. The ability to trigger safety analysis without compromising general performance means organizations can deploy LRMs with greater confidence, especially in regulated environments.

Future work may explore extending Safe Trigger to other model families and real-world testing scenarios. The researchers have made their findings available on arXiv for community review.


Sources:

Keep Reading

Recommended Stories

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Technology

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

A new research paper from arXiv shows that reinforcement learning with verifiable rewards (RLVR) can cause large reasoning models to forget foundational capabilities like perception and faithfulness. The authors propose RECAP, a replay strategy with dynamic objective reweighting that preserves general knowledge while maintaining reasoning gains.

June 21, 2026
Yann LeCun's new AI startup AMI Labs raises $1bn to build flexible intelligence beyond LLMs Technology

Yann LeCun's new AI startup AMI Labs raises $1bn to build flexible intelligence beyond LLMs

Yann LeCun, former Meta chief AI scientist, has founded AMI Labs to develop a new AI architecture called JEPA, which aims to overcome the limitations of large language models (LLMs) in understanding the physical world. The startup raised over $1bn in seed funding from Nvidia and Jeff Bezos' private investment fund, marking one of Europe's largest seed rounds.

July 2, 2026
Anthropic Believes Its Own AI Dominance Is the Only Path to Safety Technology

Anthropic Believes Its Own AI Dominance Is the Only Path to Safety

Anthropic, the AI company valued at nearly $1 trillion, holds that advancing AI capabilities and being a market leader are necessary to ensure the technology's safe development. This strategy, described by former employees and analysts, is seen as singular in its conviction.

June 26, 2026
From Construction to Injection: Edit-Based Fingerprints for Large Language Models Technology

From Construction to Injection: Edit-Based Fingerprints for Large Language Models

A new arXiv paper introduces an end-to-end injected fingerprinting framework for large language models (LLMs), addressing the dual challenges of imperceptibility and robustness. The proposed methods—Code-mixing Fingerprints (CF) and Multi-Candidate Editing (MCEdit)—aim to provide reliable ownership verification in black-box deployments without degrading model utility.

June 21, 2026