Artificial Intelligence #reinforcement learning#theorem proving
Process-Verified Reinforcement Learning for Theorem Proving via Lean: A New Path to AI Reliability
A new arXiv preprint presents process-verified reinforcement learning for theorem proving, using the Lean proof assistant as a symbolic process oracle. By parsing proof attempts into tactic sequences and leveraging Lean's type-theoretic feedback, the method delivers dense, verifier-grounded credit signals. Experiments with STP-Lean and DeepSeek-Prover-V1.5 show tactic-level supervision outperforms outcome-only baselines on MiniF2F and ProofNet benchmarks.
Jul 8, 2026 2 sources