Topic
benchmark
NSE to launch India's first domestic benchmark-based natural gas futures on July 27
The National Stock Exchange (NSE) will launch trading in Indian Natural Gas Futures on July 27, 2026, India's first exchange-traded energy derivative linked to a domestic benchmark. The cash-settled contract, under symbol NATGASIND, is based on the Indian Gas Exchange's Gujarat (Dahej) hub price and aims to provide transparent price discovery and risk management for the domestic natural gas market.
Commodities Big Jump in Benchmark Diesel Price Is Second Largest Since Iran War Began
The DOE/EIA benchmark diesel price surged 33.8 cents/gallon to $5.134/g, marking the second-largest weekly increase since the start of the Iran war. Ultra low sulfur diesel (ULSD) on CME settled at $4.119/g, nearing a $1 gain from early July. Supply disruptions from the Strait of Hormuz and potential Houthi attacks on Saudi Arabia continue to fuel bullish sentiment.
ScholarQuest Benchmark Reveals Gaps in Agentic Academic Paper Search for Enterprise AI
A new benchmark called ScholarQuest evaluates LLM-based agents for academic paper search. Built from over 1,000 computer science topics and four research intents, it provides scalable answer construction and a shared retrieval backend. Results show agentic methods beat single-shot retrieval but the top agent only achieves 0.314 Recall@100, indicating significant room for improvement in agentic search.
Commodities Benchmark Diesel Price Falls Below $5 Per Gallon for First Time Since March 9
The DOE/EIA weekly average retail diesel price fell to $4.832 per gallon, the first sub-$5 reading since March 9, marking a seventh consecutive decline totaling 80.8 cents/g. ULSD futures on the CME dropped to $3.0530/g, while Bank of America Merrill Lynch slashed its Brent crude forecast to $82/b, citing potential deficits until Q4 2026.
Business Sensex Rises 291 Points, Nifty Settles Above 24,100 on US-Iran Talks Optimism
Benchmark Indian equity indices ended nearly 0.4% higher on Monday, with the Sensex gaining 291 points to close at 77,094 and the Nifty jumping 90 points to settle at 24,103. The rally was driven by growing optimism over progress in US-Iran talks aimed at ending the West Asia conflict. Broader indices also advanced, with the Midcap 100 adding over 0.3% and the Smallcap 100 climbing 0.6%.
DRFLOW Benchmark Targets Personalized Workflow Prediction for Enterprise AI Agents
Researchers introduce DRFLOW, a benchmark for evaluating AI agents on predicting personalized workflows from heterogeneous sources. The benchmark contains 100 tasks across five domains with 1,246 workflow steps grounded in over 3,900 sources, and defines seven diagnostic metrics. A reference agent, DRFLOW-Agent, shows improvement over baselines but highlights significant remaining challenges.
New StaminaBench Benchmark Reveals Coding Agents Fail After 5-6 Turns
Researchers introduce StaminaBench, a benchmark that measures how many consecutive interaction turns coding agents can handle. Testing six harnesses and seven open-source LLMs over 100-turn scenarios, they found all models fail within 5-6 turns. Providing test feedback improved passed turn count by up to 12x, highlighting the importance of iterative testing.
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
IHBench, a new benchmark from researchers including Salimi et al., evaluates how voice agents recover after interruptions in structured enterprise workflows. The benchmark tests 27 audio-language models from OpenAI, Google, and the open-weight community, finding that closed-weight models are consistently more robust, degrading 3.3x more slowly in long conversations.
Controlled Benchmark Finds No Quantum Advantage in Brain MRI Data Augmentation
A controlled benchmark study by Haider and Figini shows that quantum-latent GAN augmentation does not improve brain MRI classification over real-data-only training or classical GANs. The quantum and classical generators were statistically indistinguishable across all data fractions from 5% to 100%.
New PhysAssistBench Tests Medical LLMs on Interactive Doctor-Patient-EHR Coordination
Researchers introduce PhysAssistBench, a benchmark for evaluating medical LLMs on interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, it uses a scalable pipeline to create agentic patients. Experiments show leading LLMs remain unreliable, highlighting the need for coordination across knowledge, communication, and systems.
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
CADBench is a unified benchmark for multimodal CAD program generation, containing 18,000 evaluation samples across six benchmark families, five input modalities, and six metrics. The benchmark evaluates eleven AI systems, generating over 1.4 million CAD programs, and reveals key failure modes in current approaches.
MEAL Benchmark Enables Continuous Multi-Agent RL Training on 100 Tasks in Hours Using GPU Acceleration
Researchers introduced MEAL (Multi-agent Environments for Adaptive Learning), the first benchmark for continual multi-agent reinforcement learning. Using JAX and GPU acceleration, MEAL enables training on sequences of 100 tasks in hours on a single GPU, revealing failure modes not apparent at smaller scales. This addresses the limitation of previous benchmarks that only considered 3-10 sequential tasks due to CPU constraints.
New Benchmark Reveals AI Agents Leak Private Data Even When Focused on Tasks
A new benchmark called TRAP evaluates the trade-off between task accuracy and privacy leakage in AI agents handling sensitive documents. Testing 22 models, the study finds non-trivial privacy leakage across all model families, with instruction-following ability correlating with leakage rate. The authors propose structural private field isolation using hash keys to prevent leakage without sacrificing task performance.
DF3DV-1K: Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Researchers introduced DF3DV-1K, a large-scale real-world dataset with 1,048 scenes and 89,924 images for distractor-free novel view synthesis. The dataset spans 128 distractor types and 161 scene themes, enabling benchmarking of nine radiance field methods and 3D Gaussian Splatting. Fine-tuning a diffusion-based 2D enhancer on DF3DV-1K achieved average improvements of 0.96 dB PSNR and 0.057 LPIPS.
RTSGameBench Benchmark Tests Strategic Reasoning in Vision-Language Models
A new benchmark called RTSGameBench evaluates strategic reasoning in vision-language models (VLMs) using the real-time strategy game Beyond All Reason. The benchmark includes diagnostic mini-games, diverse matchup structures, and a self-evolving generation framework. Initial tests show state-of-the-art VLMs struggle with tighter coordination, multiagent tasks, and increased scale.
CRAX Benchmark Delivers 100x Speedup for Safe Reinforcement Learning Research
Researchers have introduced CRAX (Constrained RL Accelerated with JAX), a fast safe reinforcement learning benchmark that leverages hardware acceleration to achieve up to 100x speedups over CPU-based alternatives. Built on MuJoCo XLA, it includes six environment suites and three agent-specific tasks across three difficulty levels. Evaluation of six popular safe RL methods reveals trade-offs between performance and safety, with curriculum learning improving results.
FFinRED: Expert-Guided Framework Red-Teams Financial LLMs Against Regulatory Evasion and Fraud
Researchers introduce FFinRED, a red-teaming framework for financial large language models (LLMs) that uses a two-level taxonomy aligned with global standards like FATF and EU DORA. The framework converts real financial documents into behavioral prompts and includes an expert-validated rubric that reduces critical false negatives from 28 to 12. It is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox.
Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages
A new benchmark, Multi-LCB, extends the popular LiveCodeBench to 12 programming languages, revealing LLMs' struggles with multilingual code generation. Evaluation of 24 models uncovered Python overfitting and language-specific contamination.
TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement
A new benchmark called TERMS-Bench goes beyond deal rate to diagnose why LLM negotiation agents fail, evaluating 13 frontier models on surplus extraction, cue use, belief calibration, and compliance. For enterprise procurement and trade, this offers actionable insights into AI agent weaknesses.
BRITE Benchmark Reveals Critical Gaps in Text-to-Video Models' Object-Action Binding and Audio-Visual Sync
A new benchmark called BRITE provides the first unified framework for evaluating text-to-video (T2V) models on implausible prompts, audio-visual consistency, and interpretable QA-based assessment. Testing five state-of-the-art models including Sora 2 and Veo 3.1, BRITE reveals that while models excel at static object composition, they show significant degradation in object-action binding and audio-visual synchronization.
When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control
A research paper introduces RLScale-Bench, a reproducible benchmark for deep reinforcement learning on adaptive resource control. Testing six DRL algorithms and a calibrated rule-based baseline on Kubernetes autoscaling across six workload patterns, the study finds that the calibrated controller achieves the lowest cost on all workloads, though DRL agents perform better on bursty and flash traffic. Discrete-action DRL algorithms also significantly outperform continuous-action ones in constraint violations.
New EEG Benchmark Promises Standardized Evaluation of Foundation Models
A new benchmark called EEG-FM-Bench aims to standardize evaluation of electroencephalography foundation models (EEG-FMs). It integrates 14 datasets across 10 paradigms and provides tools for gradient and representation analysis. Early experiments reveal critical insights about multi-task learning, pre-training efficiency, and model scaling.
MA-ProofBench: New Benchmark Tests LLMs on Formal Theorem Proving in Mathematical Analysis
Researchers introduce MA-ProofBench, the first formal theorem-proving benchmark dedicated to mathematical analysis. It contains 200 theorems across six topics at two difficulty levels. Evaluations show that even the best model, GPT-5.5, achieves only 16% Pass@8 on undergraduate-level problems and 5% on Ph.D.-level problems, highlighting significant limitations of current LLMs in formal mathematical reasoning.
KILLBENCH: New Benchmark Tests External Kill Switches to Stop Malicious AI
Researchers propose KILLBENCH, a benchmark for evaluating external AI kill switches that stop malicious web agents without internal access. The benchmark includes four agent configurations, eight harmful scenarios, and ten jailbreak patterns. It was tested on models including GPT-5.2, Grok-4.3, Gemma4, and Qwen variants.
AuAu Benchmark Audits Authoritarian Alignment in Large Language Models from Four Regions
Researchers introduce AuAu, a benchmark to assess authoritarian alignment in LLMs using psychometric tests, vignettes, and user prompts. Testing 17 models from China, EU, Russia, and USA revealed substantial authoritarian response rates and easy manipulation via system prompts.
Snyk VulnBench JS 1.0 Reveals LLM Security Reviews Are Unrepeatable: Can They Find the Same Bugs Twice?
A new benchmark from Snyk finds that agentic LLM security reviews are highly unrepeatable: 80 of 161 unique findings appeared in only one of five identical runs. By contrast, Claude's reference-matched findings were stable, and Snyk Code SAST was deterministic. The study argues for combining LLM and SAST approaches rather than treating them as replacements.
Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection
Federated learning enables collaborative medical image segmentation without centralizing sensitive data, but real-world label noise hampers deployment. A new benchmark suite combines diverse real-world noisy datasets, client-noise scenarios, and targeted evaluation to support systematic assessment of federated noisy label learning methods, addressing the gap left by synthetic noise studies.
SkillsBench Benchmark Measures How Agent Skills Boost LLM Performance Across Diverse Tasks
Researchers introduce SkillsBench, a benchmark with 87 tasks across 8 domains to measure whether agent skills improve LLM performance. Curated skills raised average pass rate from 33.9% to 50.5%, with focused skills of at most three modules outperforming larger bundles. Smaller models with skills can match larger models without.
Cough Regression Benchmark Reveals Trade-Offs in Respiratory Acoustic Foundation Models
A new benchmark from researchers at NC State evaluates five respiratory acoustic foundation models on cough regression tasks—predicting age, BMI, and disease probability from cough audio. The study reveals that smaller MLP heads often outperform linear probes, but full-MLP heads overfit on small clinical data. HeAR and M2D+Resp achieve near-full performance with only 50 samples, while OPERA models require 400. Cross-dataset transfer is asymmetric, with large diverse datasets generalizing better to small clinical populations.
ATOM-Bench: New Benchmark Evaluates Atomic Skills and Compositional Generalization in Robotic Manipulation Policies
Researchers introduce ATOM-Bench, a real-world benchmark that factorizes tabletop manipulation into atomic skills and compositional tasks. It includes 30 atomic tasks and 24 held-out compositional tasks across single-arm and dual-arm tracks, with 3,000 human demonstrations. Through 2,700 physical rollouts, the team found that current policies struggle with fine-grained motor skills, counting, and logical filtering, and strong atomic performance does not guarantee compositional transfer.
LLM-WikiRace Benchmark Reveals Frontier AI Models Still Struggle with Planning Over Knowledge Graphs
Researchers introduced LLM-WikiRace, a benchmark to evaluate large language models on planning, reasoning, and world knowledge using Wikipedia hyperlinks. Top models like Gemini-3, GPT-5, and Claude Opus 4.5 achieve superhuman performance on easy tasks but drop sharply on hard difficulty, with Gemini-3 succeeding in only 23% of hard games. The study reveals that world knowledge helps only up to a point; beyond that, planning and long-horizon reasoning are the limiting factors.
P3B3 Benchmark Reveals Strong Brazilian Portuguese Bias in Large Language Models
According to a new research paper, a team introduced P3B3, an expert-curated benchmark for measuring bias between European and Brazilian Portuguese in large language models. Experiments show most LLMs strongly prefer Brazilian Portuguese, underscoring the need for more balanced variety representation in conversational AI.
UXBench: Measuring the Actionability of LLM-Generated UX Critiques
UXBench evaluates LLM-generated UX critiques for actionability. It uses web fixtures over ten product-surface families and measures whether repair agents can improve interfaces. Results show models vary significantly in reliability.
OmniTraffic Pipeline Enables Controlled Training of Spatio-Temporal Traffic AI for Logistics
Researchers introduce OmniTraffic, a controllable generation pipeline and benchmark for spatio-temporal traffic reasoning. Built on 12 real-world intersections and surveillance footage from two countries, it generates 8M VQA samples and a 3K human-verified test set. Evaluation of 11 frontier MLLMs shows a large human-model gap, especially in topology-grounded reasoning. Fine-tuning on OmniTraffic data improves real-world performance, offering a valuable tool for logistics and supply chain AI.
MMLongEmbed Benchmark Reveals Limitations in Long-Context Multimodal Embedding Models
MMLongEmbed is the first comprehensive benchmark for evaluating multimodal embedding models (MEMs) in long-context scenarios. It comprises four retrieval tasks covering text, document, and video modalities. The evaluation reveals that current MEMs rely heavily on superficial feature matching and struggle with deep semantic and structural dependencies, with performance degrading systematically based on context length and key information placement.
CODA-BENCH: New Benchmark Reveals Code Agents Struggle with Data-Intensive Tasks
A new benchmark called CODA-BENCH evaluates code agents on data-intensive tasks using a Kaggle-based sandbox. It comprises 1,009 tasks across 31 communities, each with an average of 980 files. Even top-performing agents achieve only a 61.1% success rate, highlighting a significant gap in integrating data discovery with code execution.
EHRNote-ChatQA: New Benchmark Tests LLMs on Multi-Turn Clinical Question Answering
Researchers introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over multiple discharge summaries. Built from MIMIC-IV data, it contains 967 patient-level samples and 16,072 QA pairs, revealing that LLMs struggle more with evidence grounding than content answering and that multi-turn errors compound.
New OSGuard Benchmark Evaluates Safety of Computer-Use Agents for Enterprise AI Deployment
Researchers introduce OSGuard, a benchmark suite for evaluating safety in computer-use agents. It includes action-level guardrail decisions and a risk-augmented execution suite to detect unsafe completions that satisfy nominal task objectives. Early tests show current multimodal guardrails perform well on isolated action judgments but reveal gaps in end-to-end safety.
New Benchmark 'AgentFairBench' Tests Whether LLM Agents Discriminate in Real Actions
Researchers introduce AgentFairBench, a reproducible benchmark for demographic disparity in LLM agent actions. Unlike traditional fairness tests that grade answers, it evaluates actions across hiring, lending, and medical triage using counterfactual matched sets. A pilot study with 864 decisions reveals that naively comparing score spreads can overstate disparity by ~2.4X; using a proper null methodology, Claude Haiku 4.5 showed no significant demographic effect.
CycliST Benchmark Reveals Video Language Models Struggle with Cyclical State Transitions
The CycliST benchmark, introduced by a team of researchers, evaluates Video Language Models on cyclical state transitions. Results show current VLMs struggle to detect and reason about periodic patterns, with no single model performing consistently across all tasks.
RSRCC Benchmark Uses Retrieval-Augmented Best-of-N Ranking for Remote Sensing Change Comprehension
RSRCC is a new benchmark for remote sensing change question-answering, containing 126k questions focused on localized, semantic changes. It uses a hierarchical semi-supervised curation pipeline with retrieval-augmented Best-of-N ranking to filter noisy candidates. The dataset is available online.
AgentLeak Benchmark Reveals Internal Channel Privacy Leaks in Multi-Agent LLM Systems
A new benchmark called AgentLeak evaluates privacy leakage in multi-agent large language model (LLM) systems, finding that inter-agent messages leak at 68.8% compared to 27.2% for final outputs. Across 1,000 scenarios and five models, total system exposure reaches 68.9%, highlighting risks invisible to standard output-only audits.
New Benchmark ARB4WM Evaluates Adversarial Robustness of World Models for Safety-Critical Control
Researchers have introduced ARB4WM, a unified benchmark for evaluating adversarial robustness of world models used in continuous control systems. The framework tests attacks across policy, value, and latent-dynamics levels, revealing that targeting value estimation and latent representations can be as harmful as direct policy disruption. Early and frequent perturbations are particularly damaging, and input-level defenses offer limited recovery.
PAL-Bench Benchmark Tests AI's Ability to Reconstruct Personal Profiles from Photo Albums
PAL-Bench, a controlled benchmark introduced in a recent paper, tests AI systems' ability to reconstruct personal profiles from longitudinal photo albums. The benchmark uses 50 synthetic users and 36,659 photo records, revealing that systems can recover some owner facts but struggle with recurring identities and evidence citation. The PAL-TRACE framework achieves the best performance but leaves hard identity resolution unsolved.
RecourseBench: Modular Framework Promises Reproducible Evaluation of AI Recourse Methods
A new framework called RecourseBench aims to standardize and validate algorithmic recourse methods—counterfactual explanations that show individuals how to reverse an AI's decision. It decomposes the evaluation pipeline into five decoupled layers and integrates 28 state-of-the-art methods, with automated tests to verify reproducibility.
Business Sensex Surges 736 Points, Nifty Jumps 231 Points on US-Iran Peace Deal; Midcap, Smallcap Indices Rise
Benchmark Indian equity indices Sensex and Nifty surged nearly 1% on June 15, 2026, following a US-Iran peace agreement that eased geopolitical tensions and triggered a decline in crude oil prices. The broader market also saw gains, with Midcap100 up 1.3% and Smallcap100 up 1.1%. Broad-based buying was observed across most sectors.
Commodities Diesel Prices Plunge Amid Geopolitical Uncertainty
Diesel prices have seen a significant drop, with the CME ULSD July contract falling to $3.4886/g. This decline is driven by geopolitical developments and potential peace talks affecting the Strait of Hormuz.
Commodities Diesel Prices Slide Amid Hormuz Strait Peace Talks
Diesel prices on the CME have dropped significantly amid potential peace talks involving the U.S., Iran, and Israel, which could lead to the reopening of the Strait of Hormuz. This decline marks the sixth drop in seven weeks.