iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Maharashtra’s ₹500 crore AI agriculture policy targets data, traceability and farm advisory Commercial LPG prices drop: 19-kg cylinder rate cut by ₹202 in Delhi, ₹209 in Kolkata Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Maharashtra’s ₹500 crore AI agriculture policy targets data, traceability and farm advisory Commercial LPG prices drop: 19-kg cylinder rate cut by ₹202 in Delhi, ₹209 in Kolkata Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue
Home ›› Technology ›› Software ›› Dr-DCI: New Framework Combines Retrieval and Direct Corpus Interaction for Scalable Enterprise Search

Dr-DCI: New Framework Combines Retrieval and Direct Corpus Interaction for Scalable Enterprise Search

A new research paper introduces Dr-DCI, a retriever-steered framework that scales direct corpus interaction by dynamically expanding a local workspace. Experiments show accuracy improvements up to 8.3 points over raw DCI, with stable performance from 100K to 10M documents.

iG
iGEN Editorial
June 16, 2026
Dr-DCI: New Framework Combines Retrieval and Direct Corpus Interaction for Scalable Enterprise Search

Enterprises managing large document repositories face a fundamental trade-off: retrieval systems like BM25 or ColBERT scale well but expose only ranked results, limiting granular verification and cross-document analysis. Direct Corpus Interaction (DCI) overcomes this by enabling shell-level operations on the full corpus, but becomes slow and unstable as corpus size grows. A new research paper from arXiv introduces Dr-DCI (Retriever-Steered Direct Corpus Interaction), a framework that combines the broad recall of retrieval with the precision of direct manipulation, without sacrificing scalability.

The Challenge of Large-Corpus Agentic Search

Agentic search systems—autonomous agents that retrieve and reason over documents—rely on retriever-mediated interfaces for scalable candidate discovery. According to the paper, these interfaces expose evidence only as ranked results or bounded document views, limiting an agent's ability to reorganize material and verify constraints across documents. DCI addresses this by exposing shell-executable corpus operations for flexible search, filtering, comparison, and verification. However, full-corpus terminal commands degrade in performance and efficiency as the corpus grows. The paper notes that raw DCI becomes "slow and unstable" at scale.

How Dr-DCI Works

Dr-DCI treats retrieval as an agent-callable action for expanding a local workspace. Rather than operating directly over the full corpus, the agent dynamically pulls relevant documents into an evolving workspace and conducts DCI operations within it. This design combines retriever-level recall with DCI-style precision: retrieval keeps exploration scalable, while DCI preserves the local operations needed for effective evidence resolution. The framework thus balances the strengths of both approaches.

Experimental Results

The paper reports experiments across multiple benchmarks. On Browsecomp-Plus, Dr-DCI reaches 71.2% accuracy, improving over raw DCI and ablated variants by up to 8.3 percentage points while reducing tool usage, wall time, and estimated cost. With a workspace-preserving context reset, accuracy further improves to 73.3%. The following table summarizes key results:

Method Accuracy on Browsecomp-Plus Notes
Raw DCI (baseline) Unablated variant
Dr-DCI (standard) 71.2% Up to 8.3 points improvement over raw DCI
Dr-DCI (context reset) 73.3% Workspace-preserving context reset

In corpus-scaling experiments, Dr-DCI remained effective from 100K to 10 million documents, whereas raw DCI became unstable and BM25 performed substantially worse. Dr-DCI also scaled to a 20-million-scale file-per-document setting (Wiki-18 QA), achieving an average score of 63.0 across six benchmarks, outperforming retrieval-based and trained search-agent baselines.

Key Components and Ablation Insights

Ablation analysis revealed that ranked previews and inter-document DCI are key to performance. Ranked previews provide the agent with a concise list of relevant excerpts, while inter-document DCI enables comparisons and verification across multiple documents. Removing either component significantly degraded accuracy, confirming their importance in the framework's design.

For technology leaders evaluating AI-powered search for enterprise knowledge management, Dr-DCI offers a practical path to more accurate and reliable agentic search without sacrificing efficiency. The framework demonstrates that combining retrieval-level recall with direct corpus operations can effectively scale to tens of millions of documents, a capability increasingly critical for industries dealing with large regulatory, technical, or legal repositories.


Sources:

Keep Reading

Recommended Stories

Scientists Use AI and Quantum Computing to Generate New Peptides in Spare Time Technology

Scientists Use AI and Quantum Computing to Generate New Peptides in Spare Time

Researchers at the Technical University of Denmark used a hybrid AI-quantum computing system to generate novel peptides, achieving better results than classical models especially with limited data. The work, done on weekends with leftover funds, could accelerate personalized immunotherapies and vaccines.

July 12, 2026
India and Switzerland Step Up Innovation Partnership with Focus on Startups, Research Technology

India and Switzerland Step Up Innovation Partnership with Focus on Startups, Research

India and Switzerland are enhancing their bilateral innovation partnership, leveraging Swissnex to connect startups, researchers, and industry. Key pillars include the Indo-Swiss Joint Research Programme, healthcare collaborations, and growing interest in India's digital public infrastructure. New science and technology initiatives are expected later this year.

June 26, 2026
New AI Model Lets Robots Grasp Objects Like Humans Using RGB-D Data Technology

New AI Model Lets Robots Grasp Objects Like Humans Using RGB-D Data

Researchers introduce HUG, a flow-matching AI model that generates diverse human grasps for any object from a single RGB-D image. Trained on the 1M-HUGs egocentric dataset of 1 million frames from human grasp demonstrations, HUG outperforms state-of-the-art baselines by 23% and 34% on a challenging benchmark, enabling zero-shot grasping for multi-fingered robots.

June 20, 2026
SorryDB Benchmark Tests AI Provers on Real-World Lean Theorem Completion Tasks Technology

SorryDB Benchmark Tests AI Provers on Real-World Lean Theorem Completion Tasks

Researchers present SorryDB, a benchmark of open Lean tasks from 78 GitHub projects. Evaluating a snapshot of 1000 tasks, they show current approaches are complementary, with Gemini Flash-based agentic methods leading but not outperforming all others.

June 17, 2026