Topic
deep research
DRFLOW Benchmark Targets Personalized Workflow Prediction for Enterprise AI Agents
Researchers introduce DRFLOW, a benchmark for evaluating AI agents on predicting personalized workflows from heterogeneous sources. The benchmark contains 100 tasks across five domains with 1,246 workflow steps grounded in over 3,900 sources, and defines seven diagnostic metrics. A reference agent, DRFLOW-Agent, shows improvement over baselines but highlights significant remaining challenges.
MetaResearcher AI Framework Trains Deep Research Agents via Self-Reflective Reinforcement Learning in Adversarial Environments
MetaResearcher is a novel AI framework for training deep research agents using self-reflective reinforcement learning in adversarial virtual environments. It introduces four synergistic dimensions: Evolving Virtual World, Discovery-Oriented Tasks, Self-Reflective Meta-Reward (GRPO), and Heterogeneous Multi-Agent Swarm. Built on LiteResearcher, it requires zero marginal API cost and targets improvements on GAIA and Xbench-DS benchmarks.
ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research
ScaffoldAgent, a utility-guided dynamic outline optimization framework for open-ended deep research, models outline evolution as a structured decision process with three operations: Expansion, Contraction, and Revision. It uses a utility-guided feedback mechanism to estimate the downstream value of each operation from retrieval gain, structural coherence, and trial-generation quality. Experiments on DeepResearch Bench and DeepResearch Gym show consistent improvements in long-form report generation and factual grounding over existing deep research agents.
Hybrid Open-Ended Tri-Evolution Framework Boosts Deep Research AI Performance
Researchers propose the Hybrid Open-Ended Tri-Evolution (HOTE) framework that uses hybrid-mode reinforcement learning to collaboratively evolve a proposer, solver, and judge for deep research tasks. An 8B model trained with HOTE surpasses static open 8-32B models and state-of-the-art deep research training methods while requiring less time overhead.
S1-DeepResearch: New AI Agent Combines Search and Synthesis for Long-Horizon Research Tasks
Researchers introduce S1-DeepResearch, a unified framework for training deep research agents that combine closed-ended QA with open-ended exploration. The 32B-parameter model achieves state-of-the-art among open-source models across 20 benchmarks spanning reasoning, instruction following, report generation, file understanding, and skills usage.