iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout
Home ›› Technology ›› Ai ›› Llms ›› DeepSeek-V4 Unveils Million-Token Context Models with Major Efficiency Gains

DeepSeek-V4 Unveils Million-Token Context Models with Major Efficiency Gains

DeepSeek-AI released the preview of DeepSeek-V4 series, including two MoE language models supporting one-million-token contexts. The V4-Pro achieves a 73% reduction in inference FLOPs and 90% lower KV cache compared to its predecessor, making long-context tasks more feasible.

iG
iGEN Editorial
June 20, 2026
DeepSeek-V4 Unveils Million-Token Context Models with Major Efficiency Gains

DeepSeek-AI has released a preview version of its DeepSeek-V4 series, introducing two Mixture-of-Experts (MoE) language models capable of processing context lengths of up to one million tokens. The new models — DeepSeek-V4-Pro and DeepSeek-V4-Flash — incorporate architectural innovations that dramatically reduce computational costs for long-context inference, according to the team's paper published on arXiv.

Models and Parameters

The series includes two variants:

  • DeepSeek-V4-Pro: 1.6 trillion total parameters, with 49 billion activated per token.
  • DeepSeek-V4-Flash: 284 billion total parameters, with 13 billion activated per token.

Both models support a context length of one million tokens. The larger model, V4-Pro, also has a maximum reasoning effort mode called DeepSeek-V4-Pro-Max, which sets a new state-of-the-art among open models on core tasks, the authors state.

Architectural Innovations

Three key upgrades drive the efficiency improvements:

  • Hybrid Attention Architecture: Combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency.
  • Manifold-Constrained Hyper-Connections (mHC): Enhances traditional residual connections, improving gradient flow in deep networks.
  • Muon Optimizer: Accelerates convergence and improves training stability.

Both models were pre-trained on more than 32 trillion diverse and high-quality tokens, followed by a comprehensive post-training pipeline that further unlocks capabilities.

Efficiency Benchmarks for Million-Token Contexts

Metric DeepSeek-V4-Pro vs. DeepSeek-V3.2
Single-token inference FLOPs 27% (73% reduction)
KV cache size 10% (90% reduction)

The paper notes that these gains make it routine to support million-token contexts, enabling long-horizon tasks and further test-time scaling. The efficiency improvements come from the hybrid attention mechanisms, which reduce the memory and compute required for very long sequences.

Availability and Open-Source Release

Model checkpoints for both DeepSeek-V4-Pro and DeepSeek-V4-Flash are publicly available at the linked repository in the paper. The open-source release allows enterprise technology teams to deploy and fine-tune the models for custom long-context applications, from document analysis to code generation over millions of tokens.

For enterprise technology decision-makers evaluating AI for large-scale document processing, supply chain analytics, or trade compliance, the ability to process million-token contexts without proportional compute costs could unlock new automation possibilities. The models' open nature also enables on-premise deployment, addressing data sovereignty concerns common in international trade environments.


Sources:

Keep Reading

Recommended Stories

Google Limits Meta’s Use of Its Gemini AI Models Due to Compute Constraints Technology

Google Limits Meta’s Use of Its Gemini AI Models Due to Compute Constraints

Google has placed limits on Meta’s use of its Gemini AI models after the social media company sought more computing capacity than Google could provide. The shortfall disrupted and delayed some of Meta’s internal AI projects, according to the Financial Times. The incident underscores the broader industry struggle to secure enough computing power for AI workloads.

June 28, 2026
China's Z.ai Emerges as Low-Cost Challenger to OpenAI and Anthropic with GLM-5.2 Technology

China's Z.ai Emerges as Low-Cost Challenger to OpenAI and Anthropic with GLM-5.2

Chinese AI startup Z.ai is gaining traction with its latest flagship model GLM-5.2, which offers advanced coding and AI agent capabilities at significantly lower cost than OpenAI and Anthropic. The model has climbed developer rankings and sparked comparisons to DeepSeek, while US export restrictions fuel interest in alternatives. Pricing in India starts at about Rs 1,410 per month, undercutting ChatGPT Plus and Claude Pro.

July 6, 2026
AI enters cost-conscious era as enterprises chase returns on investment Technology

AI enters cost-conscious era as enterprises chase returns on investment

After two years of rapid AI deployment, enterprises are demanding measurable returns. Companies like Uber and Meta have introduced usage caps, while Indian firms expect a 45% increase in AI investment despite keeping budgets below 20% of IT spend. Experts warn that without proper measurement, AI spending risks becoming noise.

June 29, 2026
IndiGo Trials AI-Powered OptiClimb by SITA to Cut Fuel Burn During Take-Offs Technology

IndiGo Trials AI-Powered OptiClimb by SITA to Cut Fuel Burn During Take-Offs

IndiGo begins trials of SITA's AI-powered OptiClimb solution to reduce fuel consumption during the climb phase. The airline aims to save 60-65 kg per takeoff, with potential daily savings of tens of tonnes across its 2,000-odd flights. The trial runs on its Airbus fleet starting Thursday.

June 24, 2026