Topic
model training
Artificial Intelligence #artificial intelligence#large reasoning models
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models
A new research paper from arXiv shows that reinforcement learning with verifiable rewards (RLVR) can cause large reasoning models to forget foundational capabilities like perception and faithfulness. The authors propose RECAP, a replay strategy with dynamic objective reweighting that preserves general knowledge while maintaining reasoning gains.
Jun 21, 2026 1 source
Artificial Intelligence #llm#post-training
Which Pairs to Compare for LLM Post-Training? Research Reveals Optimal Labeling Strategy
A new arXiv paper by researchers Han, Goyal, and Ma addresses the challenge of which comparison pairs to label in preference-based LLM post-training. The study formulates comparison curation as a sampling-design problem and provides theoretical bounds showing how selection affects policy performance. Experiments demonstrate that proposed designs improve sample efficiency over common heuristics.
Jun 20, 2026 1 source