iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
ArcBest's Q2 Results Show Operational Recovery in LTL and Brokerage Segments South Africa's Prince Edward Graving Dock Sees Major Upgrade with Troy Docking, Boosting Ship Repair Options India Salary Hikes Projected at 8.6%-10.2%; EV, Fintech, Healthcare Lead Pay Gains Carrier diversification unravels the last-mile delivery duopoly as shippers seek alternatives Ferrari’s First EV, the Luce, Blends Italian Engineering With Apple-Inspired Design Sallaum Lines Expands Car Carrier Orderbook with Fresh China Deal, Boosting Ro-Ro Capacity Old Dominion nearly breaks 70% operating ratio in Q2 despite lower volumes BGN Launches US Gulf Bunkering Arm, Expanding into Direct Physical Supply of Marine Fuels Grip taps industry veteran John Hummel to lead cold chain fulfillment expansion India Retains Global Dairy Lead as USDA Forecasts Milk Output Rise to 105.4 MT in 2026 ArcBest's Q2 Results Show Operational Recovery in LTL and Brokerage Segments South Africa's Prince Edward Graving Dock Sees Major Upgrade with Troy Docking, Boosting Ship Repair Options India Salary Hikes Projected at 8.6%-10.2%; EV, Fintech, Healthcare Lead Pay Gains Carrier diversification unravels the last-mile delivery duopoly as shippers seek alternatives Ferrari’s First EV, the Luce, Blends Italian Engineering With Apple-Inspired Design Sallaum Lines Expands Car Carrier Orderbook with Fresh China Deal, Boosting Ro-Ro Capacity Old Dominion nearly breaks 70% operating ratio in Q2 despite lower volumes BGN Launches US Gulf Bunkering Arm, Expanding into Direct Physical Supply of Marine Fuels Grip taps industry veteran John Hummel to lead cold chain fulfillment expansion India Retains Global Dairy Lead as USDA Forecasts Milk Output Rise to 105.4 MT in 2026
Home ›› Topics ›› benchmark

Topic

benchmark

48 stories
NSE to launch India's first domestic benchmark-based natural gas futures on July 27 Commodities
Energy & Petrochemicals #nse#natural gas

NSE to launch India's first domestic benchmark-based natural gas futures on July 27

The National Stock Exchange (NSE) will launch trading in Indian Natural Gas Futures on July 27, 2026, India's first exchange-traded energy derivative linked to a domestic benchmark. The cash-settled contract, under symbol NATGASIND, is based on the Indian Gas Exchange's Gujarat (Dahej) hub price and aims to provide transparent price discovery and risk management for the domestic natural gas market.

Jul 24, 2026 1 source
Big Jump in Benchmark Diesel Price Is Second Largest Since Iran War Began Commodities
Energy & Petrochemicals #diesel#energy

Big Jump in Benchmark Diesel Price Is Second Largest Since Iran War Began

The DOE/EIA benchmark diesel price surged 33.8 cents/gallon to $5.134/g, marking the second-largest weekly increase since the start of the Iran war. Ultra low sulfur diesel (ULSD) on CME settled at $4.119/g, nearing a $1 gain from early July. Supply disruptions from the Strait of Hormuz and potential Houthi attacks on Saudi Arabia continue to fuel bullish sentiment.

Jul 21, 2026 1 source
ScholarQuest Benchmark Reveals Gaps in Agentic Academic Paper Search for Enterprise AI Technology
Artificial Intelligence #benchmark#academic search

ScholarQuest Benchmark Reveals Gaps in Agentic Academic Paper Search for Enterprise AI

A new benchmark called ScholarQuest evaluates LLM-based agents for academic paper search. Built from over 1,000 computer science topics and four research intents, it provides scalable answer construction and a shared retrieval backend. Results show agentic methods beat single-shot retrieval but the top agent only achieves 0.314 Recall@100, indicating significant room for improvement in agentic search.

Jul 8, 2026 1 source
Benchmark Diesel Price Falls Below $5 Per Gallon for First Time Since March 9 Commodities
Energy & Petrochemicals #diesel#fuel prices

Benchmark Diesel Price Falls Below $5 Per Gallon for First Time Since March 9

The DOE/EIA weekly average retail diesel price fell to $4.832 per gallon, the first sub-$5 reading since March 9, marking a seventh consecutive decline totaling 80.8 cents/g. ULSD futures on the CME dropped to $3.0530/g, while Bank of America Merrill Lynch slashed its Brent crude forecast to $82/b, citing potential deficits until Q4 2026.

Jun 23, 2026 1 source
Sensex Rises 291 Points, Nifty Settles Above 24,100 on US-Iran Talks Optimism Business
Markets #sensex#nifty

Sensex Rises 291 Points, Nifty Settles Above 24,100 on US-Iran Talks Optimism

Benchmark Indian equity indices ended nearly 0.4% higher on Monday, with the Sensex gaining 291 points to close at 77,094 and the Nifty jumping 90 points to settle at 24,103. The rally was driven by growing optimism over progress in US-Iran talks aimed at ending the West Asia conflict. Broader indices also advanced, with the Midcap 100 adding over 0.3% and the Smallcap 100 climbing 0.6%.

Jun 23, 2026 1 source
DRFLOW Benchmark Targets Personalized Workflow Prediction for Enterprise AI Agents Technology
Artificial Intelligence #deep research#workflow prediction

DRFLOW Benchmark Targets Personalized Workflow Prediction for Enterprise AI Agents

Researchers introduce DRFLOW, a benchmark for evaluating AI agents on predicting personalized workflows from heterogeneous sources. The benchmark contains 100 tasks across five domains with 1,246 workflow steps grounded in over 3,900 sources, and defines seven diagnostic metrics. A reference agent, DRFLOW-Agent, shows improvement over baselines but highlights significant remaining challenges.

Jun 22, 2026 1 source
New StaminaBench Benchmark Reveals Coding Agents Fail After 5-6 Turns Technology
Artificial Intelligence #staminabench#coding agents

New StaminaBench Benchmark Reveals Coding Agents Fail After 5-6 Turns

Researchers introduce StaminaBench, a benchmark that measures how many consecutive interaction turns coding agents can handle. Testing six harnesses and seven open-source LLMs over 100-turn scenarios, they found all models fail within 5-6 turns. Providing test feedback improved passed turn count by up to 12x, highlighting the importance of iterative testing.

Jun 22, 2026 2 sources
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows Technology
Artificial Intelligence #voice agents#post-interruption recovery

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

IHBench, a new benchmark from researchers including Salimi et al., evaluates how voice agents recover after interruptions in structured enterprise workflows. The benchmark tests 27 audio-language models from OpenAI, Google, and the open-weight community, finding that closed-weight models are consistently more robust, degrading 3.3x more slowly in long conversations.

Jun 22, 2026 1 source
Controlled Benchmark Finds No Quantum Advantage in Brain MRI Data Augmentation Technology
Artificial Intelligence #quantum-latent gan#gan

Controlled Benchmark Finds No Quantum Advantage in Brain MRI Data Augmentation

A controlled benchmark study by Haider and Figini shows that quantum-latent GAN augmentation does not improve brain MRI classification over real-data-only training or classical GANs. The quantum and classical generators were statistically indistinguishable across all data fractions from 5% to 100%.

Jun 21, 2026 1 source
New PhysAssistBench Tests Medical LLMs on Interactive Doctor-Patient-EHR Coordination Technology
Artificial Intelligence #llms#artificial intelligence

New PhysAssistBench Tests Medical LLMs on Interactive Doctor-Patient-EHR Coordination

Researchers introduce PhysAssistBench, a benchmark for evaluating medical LLMs on interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, it uses a scalable pipeline to create agentic patients. Experiments show leading LLMs remain unreliable, highlighting the need for coordination across knowledge, communication, and systems.

Jun 21, 2026 1 source
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation Technology
Artificial Intelligence #ai#cad

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

CADBench is a unified benchmark for multimodal CAD program generation, containing 18,000 evaluation samples across six benchmark families, five input modalities, and six metrics. The benchmark evaluates eleven AI systems, generating over 1.4 million CAD programs, and reveals key failure modes in current approaches.

Jun 21, 2026 1 source
MEAL Benchmark Enables Continuous Multi-Agent RL Training on 100 Tasks in Hours Using GPU Acceleration Technology
Artificial Intelligence #ai#reinforcement learning

MEAL Benchmark Enables Continuous Multi-Agent RL Training on 100 Tasks in Hours Using GPU Acceleration

Researchers introduced MEAL (Multi-agent Environments for Adaptive Learning), the first benchmark for continual multi-agent reinforcement learning. Using JAX and GPU acceleration, MEAL enables training on sequences of 100 tasks in hours on a single GPU, revealing failure modes not apparent at smaller scales. This addresses the limitation of previous benchmarks that only considered 3-10 sequential tasks due to CPU constraints.

Jun 21, 2026 1 source
New Benchmark Reveals AI Agents Leak Private Data Even When Focused on Tasks Technology
Artificial Intelligence #benchmark#privacy

New Benchmark Reveals AI Agents Leak Private Data Even When Focused on Tasks

A new benchmark called TRAP evaluates the trade-off between task accuracy and privacy leakage in AI agents handling sensitive documents. Testing 22 models, the study finds non-trivial privacy leakage across all model families, with instruction-following ability correlating with leakage rate. The authors propose structural private field isolation using hash keys to prevent leakage without sacrificing task performance.

Jun 21, 2026 1 source
DF3DV-1K: Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis Technology
Artificial Intelligence #dataset#benchmark

DF3DV-1K: Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Researchers introduced DF3DV-1K, a large-scale real-world dataset with 1,048 scenes and 89,924 images for distractor-free novel view synthesis. The dataset spans 128 distractor types and 161 scene themes, enabling benchmarking of nine radiance field methods and 3D Gaussian Splatting. Fine-tuning a diffusion-based 2D enhancer on DF3DV-1K achieved average improvements of 0.96 dB PSNR and 0.057 LPIPS.

Jun 21, 2026 1 source
RTSGameBench Benchmark Tests Strategic Reasoning in Vision-Language Models Technology
Artificial Intelligence #rtsgamebench#benchmark

RTSGameBench Benchmark Tests Strategic Reasoning in Vision-Language Models

A new benchmark called RTSGameBench evaluates strategic reasoning in vision-language models (VLMs) using the real-time strategy game Beyond All Reason. The benchmark includes diagnostic mini-games, diverse matchup structures, and a self-evolving generation framework. Initial tests show state-of-the-art VLMs struggle with tighter coordination, multiagent tasks, and increased scale.

Jun 21, 2026 1 source
CRAX Benchmark Delivers 100x Speedup for Safe Reinforcement Learning Research Technology
Artificial Intelligence #reinforcement learning#safe rl

CRAX Benchmark Delivers 100x Speedup for Safe Reinforcement Learning Research

Researchers have introduced CRAX (Constrained RL Accelerated with JAX), a fast safe reinforcement learning benchmark that leverages hardware acceleration to achieve up to 100x speedups over CPU-based alternatives. Built on MuJoCo XLA, it includes six environment suites and three agent-specific tasks across three difficulty levels. Evaluation of six popular safe RL methods reveals trade-offs between performance and safety, with curriculum learning improving results.

Jun 20, 2026 1 source
FFinRED: Expert-Guided Framework Red-Teams Financial LLMs Against Regulatory Evasion and Fraud Technology
Artificial Intelligence #artificial intelligence#llms

FFinRED: Expert-Guided Framework Red-Teams Financial LLMs Against Regulatory Evasion and Fraud

Researchers introduce FFinRED, a red-teaming framework for financial large language models (LLMs) that uses a two-level taxonomy aligned with global standards like FATF and EU DORA. The framework converts real financial documents into behavioral prompts and includes an expert-validated rubric that reduces critical false negatives from 28 to 12. It is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox.

Jun 20, 2026 1 source
Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages Technology
Artificial Intelligence #multi-lcb#livecodebench

Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages

A new benchmark, Multi-LCB, extends the popular LiveCodeBench to 12 programming languages, revealing LLMs' struggles with multilingual code generation. Evaluation of 24 models uncovered Python overfitting and language-specific contamination.

Jun 20, 2026 1 source
TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement Technology
Artificial Intelligence #llm#negotiation

TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement

A new benchmark called TERMS-Bench goes beyond deal rate to diagnose why LLM negotiation agents fail, evaluating 13 frontier models on surplus extraction, cue use, belief calibration, and compliance. For enterprise procurement and trade, this offers actionable insights into AI agent weaknesses.

Jun 17, 2026 1 source
BRITE Benchmark Reveals Critical Gaps in Text-to-Video Models' Object-Action Binding and Audio-Visual Sync Technology
Artificial Intelligence #benchmark#text-to-video

BRITE Benchmark Reveals Critical Gaps in Text-to-Video Models' Object-Action Binding and Audio-Visual Sync

A new benchmark called BRITE provides the first unified framework for evaluating text-to-video (T2V) models on implausible prompts, audio-visual consistency, and interpretable QA-based assessment. Testing five state-of-the-art models including Sora 2 and Veo 3.1, BRITE reveals that while models excel at static object composition, they show significant degradation in object-action binding and audio-visual synchronization.

Jun 16, 2026 1 source
When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control Technology
Artificial Intelligence #deep reinforcement learning#adaptive resource control

When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control

A research paper introduces RLScale-Bench, a reproducible benchmark for deep reinforcement learning on adaptive resource control. Testing six DRL algorithms and a calibrated rule-based baseline on Kubernetes autoscaling across six workload patterns, the study finds that the calibrated controller achieves the lowest cost on all workloads, though DRL agents perform better on bursty and flash traffic. Discrete-action DRL algorithms also significantly outperform continuous-action ones in constraint violations.

Jun 16, 2026 1 source
New EEG Benchmark Promises Standardized Evaluation of Foundation Models Technology
Artificial Intelligence #eeg#foundation models

New EEG Benchmark Promises Standardized Evaluation of Foundation Models

A new benchmark called EEG-FM-Bench aims to standardize evaluation of electroencephalography foundation models (EEG-FMs). It integrates 14 datasets across 10 paradigms and provides tools for gradient and representation analysis. Early experiments reveal critical insights about multi-task learning, pre-training efficiency, and model scaling.

Jun 16, 2026 1 source
MA-ProofBench: New Benchmark Tests LLMs on Formal Theorem Proving in Mathematical Analysis Technology
Artificial Intelligence #large language models#theorem proving

MA-ProofBench: New Benchmark Tests LLMs on Formal Theorem Proving in Mathematical Analysis

Researchers introduce MA-ProofBench, the first formal theorem-proving benchmark dedicated to mathematical analysis. It contains 200 theorems across six topics at two difficulty levels. Evaluations show that even the best model, GPT-5.5, achieves only 16% Pass@8 on undergraduate-level problems and 5% on Ph.D.-level problems, highlighting significant limitations of current LLMs in formal mathematical reasoning.

Jun 16, 2026 1 source
KILLBENCH: New Benchmark Tests External Kill Switches to Stop Malicious AI Technology
Artificial Intelligence #ai#artificial intelligence

KILLBENCH: New Benchmark Tests External Kill Switches to Stop Malicious AI

Researchers propose KILLBENCH, a benchmark for evaluating external AI kill switches that stop malicious web agents without internal access. The benchmark includes four agent configurations, eight harmful scenarios, and ten jailbreak patterns. It was tested on models including GPT-5.2, Grok-4.3, Gemma4, and Qwen variants.

Jun 16, 2026 1 source
AuAu Benchmark Audits Authoritarian Alignment in Large Language Models from Four Regions Technology
Artificial Intelligence #benchmark#auditing

AuAu Benchmark Audits Authoritarian Alignment in Large Language Models from Four Regions

Researchers introduce AuAu, a benchmark to assess authoritarian alignment in LLMs using psychometric tests, vignettes, and user prompts. Testing 17 models from China, EU, Russia, and USA revealed substantial authoritarian response rates and easy manipulation via system prompts.

Jun 16, 2026 1 source
Snyk VulnBench JS 1.0 Reveals LLM Security Reviews Are Unrepeatable: Can They Find the Same Bugs Twice? Technology
Artificial Intelligence #snyk#vulnbench

Snyk VulnBench JS 1.0 Reveals LLM Security Reviews Are Unrepeatable: Can They Find the Same Bugs Twice?

A new benchmark from Snyk finds that agentic LLM security reviews are highly unrepeatable: 80 of 161 unique findings appeared in only one of five identical runs. By contrast, Claude's reference-matched findings were stable, and Snyk Code SAST was deterministic. The study argues for combining LLM and SAST approaches rather than treating them as replacements.

Jun 16, 2026 1 source
Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection Technology
Artificial Intelligence #federated learning#medical image segmentation

Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection

Federated learning enables collaborative medical image segmentation without centralizing sensitive data, but real-world label noise hampers deployment. A new benchmark suite combines diverse real-world noisy datasets, client-noise scenarios, and targeted evaluation to support systematic assessment of federated noisy label learning methods, addressing the gap left by synthetic noise studies.

Jun 16, 2026 1 source
SkillsBench Benchmark Measures How Agent Skills Boost LLM Performance Across Diverse Tasks Technology
Artificial Intelligence #skillsbench#benchmarking

SkillsBench Benchmark Measures How Agent Skills Boost LLM Performance Across Diverse Tasks

Researchers introduce SkillsBench, a benchmark with 87 tasks across 8 domains to measure whether agent skills improve LLM performance. Curated skills raised average pass rate from 33.9% to 50.5%, with focused skills of at most three modules outperforming larger bundles. Smaller models with skills can match larger models without.

Jun 16, 2026 1 source
Cough Regression Benchmark Reveals Trade-Offs in Respiratory Acoustic Foundation Models Technology
Artificial Intelligence #benchmark#cough regression

Cough Regression Benchmark Reveals Trade-Offs in Respiratory Acoustic Foundation Models

A new benchmark from researchers at NC State evaluates five respiratory acoustic foundation models on cough regression tasks—predicting age, BMI, and disease probability from cough audio. The study reveals that smaller MLP heads often outperform linear probes, but full-MLP heads overfit on small clinical data. HeAR and M2D+Resp achieve near-full performance with only 50 samples, while OPERA models require 400. Cross-dataset transfer is asymmetric, with large diverse datasets generalizing better to small clinical populations.

Jun 16, 2026 1 source
ATOM-Bench: New Benchmark Evaluates Atomic Skills and Compositional Generalization in Robotic Manipulation Policies Technology
Artificial Intelligence #benchmark#atomic skills

ATOM-Bench: New Benchmark Evaluates Atomic Skills and Compositional Generalization in Robotic Manipulation Policies

Researchers introduce ATOM-Bench, a real-world benchmark that factorizes tabletop manipulation into atomic skills and compositional tasks. It includes 30 atomic tasks and 24 held-out compositional tasks across single-arm and dual-arm tracks, with 3,000 human demonstrations. Through 2,700 physical rollouts, the team found that current policies struggle with fine-grained motor skills, counting, and logical filtering, and strong atomic performance does not guarantee compositional transfer.

Jun 16, 2026 1 source
LLM-WikiRace Benchmark Reveals Frontier AI Models Still Struggle with Planning Over Knowledge Graphs Technology
Artificial Intelligence #llm#benchmark

LLM-WikiRace Benchmark Reveals Frontier AI Models Still Struggle with Planning Over Knowledge Graphs

Researchers introduced LLM-WikiRace, a benchmark to evaluate large language models on planning, reasoning, and world knowledge using Wikipedia hyperlinks. Top models like Gemini-3, GPT-5, and Claude Opus 4.5 achieve superhuman performance on easy tasks but drop sharply on hard difficulty, with Gemini-3 succeeding in only 23% of hard games. The study reveals that world knowledge helps only up to a point; beyond that, planning and long-horizon reasoning are the limiting factors.

Jun 16, 2026 1 source
P3B3 Benchmark Reveals Strong Brazilian Portuguese Bias in Large Language Models Technology
Artificial Intelligence #llm#benchmark

P3B3 Benchmark Reveals Strong Brazilian Portuguese Bias in Large Language Models

According to a new research paper, a team introduced P3B3, an expert-curated benchmark for measuring bias between European and Brazilian Portuguese in large language models. Experiments show most LLMs strongly prefer Brazilian Portuguese, underscoring the need for more balanced variety representation in conversational AI.

Jun 16, 2026 1 source
UXBench: Measuring the Actionability of LLM-Generated UX Critiques Technology
Artificial Intelligence #llms#ux

UXBench: Measuring the Actionability of LLM-Generated UX Critiques

UXBench evaluates LLM-generated UX critiques for actionability. It uses web fixtures over ten product-surface families and measures whether repair agents can improve interfaces. Results show models vary significantly in reliability.

Jun 16, 2026 1 source
OmniTraffic Pipeline Enables Controlled Training of Spatio-Temporal Traffic AI for Logistics Technology
Artificial Intelligence #omnitraffic#controllable generation

OmniTraffic Pipeline Enables Controlled Training of Spatio-Temporal Traffic AI for Logistics

Researchers introduce OmniTraffic, a controllable generation pipeline and benchmark for spatio-temporal traffic reasoning. Built on 12 real-world intersections and surveillance footage from two countries, it generates 8M VQA samples and a 3K human-verified test set. Evaluation of 11 frontier MLLMs shows a large human-model gap, especially in topology-grounded reasoning. Fine-tuning on OmniTraffic data improves real-world performance, offering a valuable tool for logistics and supply chain AI.

Jun 16, 2026 1 source
MMLongEmbed Benchmark Reveals Limitations in Long-Context Multimodal Embedding Models Technology
Artificial Intelligence #multimodal#embedding

MMLongEmbed Benchmark Reveals Limitations in Long-Context Multimodal Embedding Models

MMLongEmbed is the first comprehensive benchmark for evaluating multimodal embedding models (MEMs) in long-context scenarios. It comprises four retrieval tasks covering text, document, and video modalities. The evaluation reveals that current MEMs rely heavily on superficial feature matching and struggle with deep semantic and structural dependencies, with performance degrading systematically based on context length and key information placement.

Jun 16, 2026 1 source
CODA-BENCH: New Benchmark Reveals Code Agents Struggle with Data-Intensive Tasks Technology
Artificial Intelligence #code agents#benchmark

CODA-BENCH: New Benchmark Reveals Code Agents Struggle with Data-Intensive Tasks

A new benchmark called CODA-BENCH evaluates code agents on data-intensive tasks using a Kaggle-based sandbox. It comprises 1,009 tasks across 31 communities, each with an average of 980 files. Even top-performing agents achieve only a 61.1% success rate, highlighting a significant gap in integrating data discovery with code execution.

Jun 16, 2026 1 source
EHRNote-ChatQA: New Benchmark Tests LLMs on Multi-Turn Clinical Question Answering Technology
Artificial Intelligence #ehr#clinical

EHRNote-ChatQA: New Benchmark Tests LLMs on Multi-Turn Clinical Question Answering

Researchers introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over multiple discharge summaries. Built from MIMIC-IV data, it contains 967 patient-level samples and 16,072 QA pairs, revealing that LLMs struggle more with evidence grounding than content answering and that multi-turn errors compound.

Jun 16, 2026 1 source
New OSGuard Benchmark Evaluates Safety of Computer-Use Agents for Enterprise AI Deployment Technology
Artificial Intelligence #ai safety#benchmark

New OSGuard Benchmark Evaluates Safety of Computer-Use Agents for Enterprise AI Deployment

Researchers introduce OSGuard, a benchmark suite for evaluating safety in computer-use agents. It includes action-level guardrail decisions and a risk-augmented execution suite to detect unsafe completions that satisfy nominal task objectives. Early tests show current multimodal guardrails perform well on isolated action judgments but reveal gaps in end-to-end safety.

Jun 16, 2026 1 source
New Benchmark 'AgentFairBench' Tests Whether LLM Agents Discriminate in Real Actions Technology
Artificial Intelligence #llm#ai agents

New Benchmark 'AgentFairBench' Tests Whether LLM Agents Discriminate in Real Actions

Researchers introduce AgentFairBench, a reproducible benchmark for demographic disparity in LLM agent actions. Unlike traditional fairness tests that grade answers, it evaluates actions across hiring, lending, and medical triage using counterfactual matched sets. A pilot study with 864 decisions reveals that naively comparing score spreads can overstate disparity by ~2.4X; using a proper null methodology, Claude Haiku 4.5 showed no significant demographic effect.

Jun 16, 2026 1 source
CycliST Benchmark Reveals Video Language Models Struggle with Cyclical State Transitions Technology
Artificial Intelligence #cyclist#video language model

CycliST Benchmark Reveals Video Language Models Struggle with Cyclical State Transitions

The CycliST benchmark, introduced by a team of researchers, evaluates Video Language Models on cyclical state transitions. Results show current VLMs struggle to detect and reason about periodic patterns, with no single model performing consistently across all tasks.

Jun 16, 2026 1 source
RSRCC Benchmark Uses Retrieval-Augmented Best-of-N Ranking for Remote Sensing Change Comprehension Technology
Artificial Intelligence #remote sensing#benchmark

RSRCC Benchmark Uses Retrieval-Augmented Best-of-N Ranking for Remote Sensing Change Comprehension

RSRCC is a new benchmark for remote sensing change question-answering, containing 126k questions focused on localized, semantic changes. It uses a hierarchical semi-supervised curation pipeline with retrieval-augmented Best-of-N ranking to filter noisy candidates. The dataset is available online.

Jun 16, 2026 1 source
AgentLeak Benchmark Reveals Internal Channel Privacy Leaks in Multi-Agent LLM Systems Technology
Artificial Intelligence #ai#llm

AgentLeak Benchmark Reveals Internal Channel Privacy Leaks in Multi-Agent LLM Systems

A new benchmark called AgentLeak evaluates privacy leakage in multi-agent large language model (LLM) systems, finding that inter-agent messages leak at 68.8% compared to 27.2% for final outputs. Across 1,000 scenarios and five models, total system exposure reaches 68.9%, highlighting risks invisible to standard output-only audits.

Jun 16, 2026 1 source
New Benchmark ARB4WM Evaluates Adversarial Robustness of World Models for Safety-Critical Control Technology
Artificial Intelligence #ai#adversarial robustness

New Benchmark ARB4WM Evaluates Adversarial Robustness of World Models for Safety-Critical Control

Researchers have introduced ARB4WM, a unified benchmark for evaluating adversarial robustness of world models used in continuous control systems. The framework tests attacks across policy, value, and latent-dynamics levels, revealing that targeting value estimation and latent representations can be as harmful as direct policy disruption. Early and frequent perturbations are particularly damaging, and input-level defenses offer limited recovery.

Jun 16, 2026 1 source
PAL-Bench Benchmark Tests AI's Ability to Reconstruct Personal Profiles from Photo Albums Technology
Artificial Intelligence #pal-bench#profile reconstruction

PAL-Bench Benchmark Tests AI's Ability to Reconstruct Personal Profiles from Photo Albums

PAL-Bench, a controlled benchmark introduced in a recent paper, tests AI systems' ability to reconstruct personal profiles from longitudinal photo albums. The benchmark uses 50 synthetic users and 36,659 photo records, revealing that systems can recover some owner facts but struggle with recurring identities and evidence citation. The PAL-TRACE framework achieves the best performance but leaves hard identity resolution unsolved.

Jun 16, 2026 1 source
RecourseBench: Modular Framework Promises Reproducible Evaluation of AI Recourse Methods Technology
Artificial Intelligence #algorithmic recourse#machine learning

RecourseBench: Modular Framework Promises Reproducible Evaluation of AI Recourse Methods

A new framework called RecourseBench aims to standardize and validate algorithmic recourse methods—counterfactual explanations that show individuals how to reverse an AI's decision. It decomposes the evaluation pipeline into five decoupled layers and integrates 28 state-of-the-art methods, with automated tests to verify reproducibility.

Jun 16, 2026 1 source
Sensex Surges 736 Points, Nifty Jumps 231 Points on US-Iran Peace Deal; Midcap, Smallcap Indices Rise Business
Markets #stock market#equity indices

Sensex Surges 736 Points, Nifty Jumps 231 Points on US-Iran Peace Deal; Midcap, Smallcap Indices Rise

Benchmark Indian equity indices Sensex and Nifty surged nearly 1% on June 15, 2026, following a US-Iran peace agreement that eased geopolitical tensions and triggered a decline in crude oil prices. The broader market also saw gains, with Midcap100 up 1.3% and Smallcap100 up 1.1%. Broad-based buying was observed across most sectors.

Jun 15, 2026 1 source
Diesel Prices Plunge Amid Geopolitical Uncertainty Commodities
Energy & Petrochemicals #diesel#energy

Diesel Prices Plunge Amid Geopolitical Uncertainty

Diesel prices have seen a significant drop, with the CME ULSD July contract falling to $3.4886/g. This decline is driven by geopolitical developments and potential peace talks affecting the Strait of Hormuz.

Jun 4, 2026 1 source
Diesel Prices Slide Amid Hormuz Strait Peace Talks Commodities
Energy & Petrochemicals #diesel#fuel

Diesel Prices Slide Amid Hormuz Strait Peace Talks

Diesel prices on the CME have dropped significantly amid potential peace talks involving the U.S., Iran, and Israel, which could lead to the reopening of the Strait of Hormuz. This decline marks the sixth drop in seven weeks.

May 30, 2026 1 source