iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Topics ›› post-training

Topic

post-training

3 stories
AAPA: Adversarially Anchored Preference Alignment Enhances LLM Post-Training Performance Technology
Artificial Intelligence #aapa#adversarial anchoring

AAPA: Adversarially Anchored Preference Alignment Enhances LLM Post-Training Performance

Researchers propose AAPA, a plug-in framework that adds a sentence-level adversarial anchoring signal to existing post-training objectives for large language models. Experiments on instruction-following benchmarks show consistent improvements, with staged AAPA achieving 5.77% gain on Qwen3-0.6B and 3.75% on Qwen3-4B over a strong GRPO baseline.

Jun 20, 2026 1 source
Which Pairs to Compare for LLM Post-Training? Research Reveals Optimal Labeling Strategy Technology
Artificial Intelligence #llm#post-training

Which Pairs to Compare for LLM Post-Training? Research Reveals Optimal Labeling Strategy

A new arXiv paper by researchers Han, Goyal, and Ma addresses the challenge of which comparison pairs to label in preference-based LLM post-training. The study formulates comparison curation as a sampling-design problem and provides theoretical bounds showing how selection affects policy performance. Experiments demonstrate that proposed designs improve sample efficiency over common heuristics.

Jun 20, 2026 1 source
ReGrad: A New AI Paradigm for Continual Learning Without Catastrophic Forgetting Technology
Artificial Intelligence #machine learning#continual learning

ReGrad: A New AI Paradigm for Continual Learning Without Catastrophic Forgetting

A new paper introduces ReGrad (Retrievable Gradients), a paradigm for continual post-training that pre-computes document-specific gradients, stores them in a Gradient Bank, and retrieves query-relevant gradients at inference time for temporary weight adaptation. The method uses bi-level meta-learning to reshape gradients into generalizable signals, outperforming CPT and RAG baselines in experiments.

Jun 16, 2026 1 source