iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Open Science Gains Ground: 10-Year AI Study Shows Sharp Rise in Code and Data Sharing

Open Science Gains Ground: 10-Year AI Study Shows Sharp Rise in Code and Data Sharing

A decade-long analysis of 56,800 AI conference papers shows documentation practices improving dramatically, with code and data sharing nearly sixfold from 11% to 64%. Estimated reproducibility also rose from 28% to 64%, improvements that predated formal reproducibility checklists.

iG
iGEN Editorial
June 16, 2026
Open Science Gains Ground: 10-Year AI Study Shows Sharp Rise in Code and Data Sharing

The reproducibility crisis in artificial intelligence research has prompted major conferences to adopt documentation standards, but a new analysis of 56,800 papers from 2014 to 2024 suggests that the field's improvement in sharing code and data predates and far exceeds the impact of these formal requirements. According to a study by Coakley, Snelleman, Hoos, and Gundersen, published on arXiv, the proportion of papers that share both code and data increased nearly sixfold over the decade, from 11% to 64%.

Methodology and Scope

The researchers assessed all published papers from five leading AI conferences over the past decade. They identified seven reproducibility variables, which were quality-assured, and used them to analyze the 56,800 publications. The study focused on documentation practices rather than directly testing reproducibility—the reproducibility estimates were inferred from documentation practices based on empirical reproducibility rates from a prior study.

Key Findings

Metric 2014 2024
Papers sharing both code and data 11% 64%
Estimated reproducibility 28% 64%

According to the study, improvements in documentation practices predate the introduction of reproducibility checklists, suggesting these changes reflect a broader movement toward open science rather than a direct response to formal requirements. The authors noted that in the period 2014 to 2024, documentation practices have improved substantially.

Implications for AI Adoption

For enterprise technology leaders evaluating AI systems, the trend toward increased code and data sharing enhances the ability to verify and reproduce research findings. While the study does not directly assess commercial AI products, the same open-science principles that drive increased reproducibility in academic research can reduce the risk of adopting opaque or non-reproducible models. The shift from 11% to 64% code and data sharing indicates that a majority of AI research now provides the building blocks needed for independent validation.

The broader open science movement, as evidenced by this analysis, is reshaping how AI research is conducted and disseminated. Enterprise buyers of AI solutions should consider whether vendors' claims are grounded in reproducible, openly documented work—a practice that this study shows is becoming the norm rather than the exception.


Sources:

Keep Reading

Recommended Stories

Process-Verified Reinforcement Learning for Theorem Proving via Lean: A New Path to AI Reliability Technology

Process-Verified Reinforcement Learning for Theorem Proving via Lean: A New Path to AI Reliability

A new arXiv preprint presents process-verified reinforcement learning for theorem proving, using the Lean proof assistant as a symbolic process oracle. By parsing proof attempts into tactic sequences and leveraging Lean's type-theoretic feedback, the method delivers dense, verifier-grounded credit signals. Experiments with STP-Lean and DeepSeek-Prover-V1.5 show tactic-level supervision outperforms outcome-only baselines on MiniF2F and ProofNet benchmarks.

July 8, 2026
Yann LeCun's new AI startup AMI Labs raises $1bn to build flexible intelligence beyond LLMs Technology

Yann LeCun's new AI startup AMI Labs raises $1bn to build flexible intelligence beyond LLMs

Yann LeCun, former Meta chief AI scientist, has founded AMI Labs to develop a new AI architecture called JEPA, which aims to overcome the limitations of large language models (LLMs) in understanding the physical world. The startup raised over $1bn in seed funding from Nvidia and Jeff Bezos' private investment fund, marking one of Europe's largest seed rounds.

July 2, 2026
ScaleWoB Framework Synthesizes Realistic Environments to Evaluate GUI Agents at Scale Technology

ScaleWoB Framework Synthesizes Realistic Environments to Evaluate GUI Agents at Scale

ScaleWoB is a new framework that generates high-fidelity synthesized interactive environments for evaluating GUI agents across mobile, desktop, and automotive platforms. It includes 100+ environments and 1000+ verifiable tasks. Experiments on five state-of-the-art mobile GUI agents show an average success rate of only 27.92%, compared to 92.08% for humans, highlighting substantial room for improvement.

June 22, 2026
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Technology

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

A new research paper from arXiv shows that reinforcement learning with verifiable rewards (RLVR) can cause large reasoning models to forget foundational capabilities like perception and faithfulness. The authors propose RECAP, a replay strategy with dynamic objective reweighting that preserves general knowledge while maintaining reasoning gains.

June 21, 2026