iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition
Home ›› Technology ›› Ai ›› Llms ›› Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages

Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages

A new benchmark, Multi-LCB, extends the popular LiveCodeBench to 12 programming languages, revealing LLMs' struggles with multilingual code generation. Evaluation of 24 models uncovered Python overfitting and language-specific contamination.

iG
iGEN Editorial
June 20, 2026
Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages

LiveCodeBench (LCB) has become a widely adopted benchmark for evaluating large language models (LLMs) on code-generation tasks. However, according to a paper published on arxiv.org, LCB remains restricted to Python, leaving open the question of whether LLMs can generalize across the diverse programming languages required in real-world software engineering. To address this, researchers introduced Multi-LCB, a benchmark that evaluates LLMs across 12 programming languages, including Python, while preserving LCB's contamination controls and evaluation protocol.

Benchmarks and Methodology

Multi-LCB transforms Python tasks from the LCB dataset into equivalent tasks in other languages. Because it is fully compatible with the original LCB format, Multi-LCB will automatically track future LCB updates, enabling systematic assessment of cross-language code generation competence. The paper evaluated 24 LLMs for instruction and reasoning on Multi-LCB, uncovering evidence of Python overfitting, language-specific contamination, and substantial disparities in multilingual performance.

Key Findings

The researchers reported that the results establish Multi-LCB as a rigorous new benchmark for multi-programming-language code evaluation, directly addressing LCB's primary limitation and exposing critical gaps in current LLM capabilities. The study found that models that perform well on Python often degrade significantly when asked to generate code in languages such as Java, C++, or JavaScript. Language-specific contamination, where training data for one language leaks into others, also skewed performance.

Aspect Details
Original Benchmark LiveCodeBench (Python-only)
Extended Benchmark Multi-LCB (12 languages including Python)
Number of Models Evaluated 24
Key Issues Uncovered Python overfitting, language-specific contamination, performance disparities
Compatibility Fully compatible with LCB; tracks future updates

Our results establish Multi-LCB as a rigorous new benchmark for multi-programming-language code evaluation, directly addressing LCB's primary limitation and exposing critical gaps in current LLM capabilities.

Implications for Enterprise AI Adoption

For enterprise technology leaders evaluating LLMs for code generation across multiple programming languages—common in legacy system modernization, polyglot microservices, and full-stack development—the findings suggest that relying solely on Python benchmarks can give a misleading picture of a model's true coding ability. Multi-LCB provides a contamination-aware evaluation that may help procurement teams select models that generalize better across languages, reducing the risk of vendor lock-in or hidden performance cliffs. As organizations increasingly integrate AI-assisted coding into their software supply chains, tools like Multi-LCB can inform more realistic assessments.

The paper is available on arxiv.org, authored by Ivanova; Maria; Zadorozhny; Pavel; Levichev; Rodion; Petrov; Adamenko; Lopatin; Kutalev; Alexey; Babaev; Dmitrii.


Sources:

Keep Reading

Recommended Stories

TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement Technology

TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement

A new benchmark called TERMS-Bench goes beyond deal rate to diagnose why LLM negotiation agents fail, evaluating 13 frontier models on surplus extraction, cue use, belief calibration, and compliance. For enterprise procurement and trade, this offers actionable insights into AI agent weaknesses.

June 17, 2026
SPARK Method Activates Latent Security Knowledge in LLMs for Secure Code Generation Technology

SPARK Method Activates Latent Security Knowledge in LLMs for Secure Code Generation

SPARK (Security Knowledge Priming and Representation-Guided Knowledge Activation) is a new inference-time method that improves the security of code generated by large language models without requiring retraining. The researchers argue that pretraining data already contains sufficient security material; the bottleneck is activation. Evaluated on 9 open-source and 7 proprietary models, SPARK matches or improves secure code generation baselines while preserving code utility.

June 16, 2026
Z.ai GLM 5.3 open-weight model arrives with near-frontier hacking skills Technology

Z.ai GLM 5.3 open-weight model arrives with near-frontier hacking skills

Chinese AI company Z.ai announced GLM 5.3, an open-weight model it says automates coding and cybersecurity tasks almost as well as Anthropic and OpenAI's best models. It also launched OpenVuln for code scanning. Z.ai is staging access to security partners before full release in two weeks.

August 18, 2026
Mistral Seizes Opening as US AI Restrictions Push Europe Toward Open Source Technology

Mistral Seizes Opening as US AI Restrictions Push Europe Toward Open Source

Mistral, a French AI lab, is capitalizing on US restrictions on rival AI models and safety incidents at OpenAI and Anthropic to position itself as Europe's open-source alternative. The company raised nearly $2 billion at a $13.5 billion valuation and reports 20x revenue growth, with deals from Microsoft, HSBC, and the French government.

August 4, 2026