iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition
Home ›› Technology ›› Hardware ›› FP8 Debunks FP64 as HPC Holy Grail in New Paper from Satoshi Matsuoka

FP8 Debunks FP64 as HPC Holy Grail in New Paper from Satoshi Matsuoka

A new arXiv preprint by Satoshi Matsuoka challenges the long-held belief that native FP64 hardware is essential for high-performance scientific computing. The paper proposes that FP8 tensor-core operations, combined with the Ozaki Scheme II, can deliver equivalent double-precision accuracy, reducing FP64 from a hardware requirement to a derived guarantee. The analytical framework is tested across a five-layer hierarchy, projecting improved performance on upcoming NVIDIA GPUs.

iG
iGEN Editorial
June 16, 2026
FP8 Debunks FP64 as HPC Holy Grail in New Paper from Satoshi Matsuoka

The assumption that native hardware FP64 is the irreducible foundation of scientific computing is being directly challenged in a new paper on arXiv. Authored by Satoshi Matsuoka, the preprint titled "FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (June 13th version)" argues that on AI-optimized GPUs of the NVIDIA B300 generation and beyond, FP8 tensor throughput has grown to multiple PFLOPS while native FP64 throughput has collapsed to approximately 1.3 TFLOPS. The paper claims this shift is not just survivable but preferable: the FP8 tensor-core matrix-multiply can serve as the sole computational primitive for double-precision scientific computing.

The Ozaki Scheme II and FP8 Composition

According to the paper, every canonical kernel in scientific computing—dense and sparse linear algebra, spectral transforms, stencils—along with every application composing them, can be reduced via the Ozaki Scheme II to sequences of FP8 matrix operations. The Ozaki Scheme II relies on the Chinese Remainder Theorem to maintain accuracy. The only non-FP8 arithmetic involved is a bounded, fixed-width integer accumulation at reconstruction. This approach demotes native FP64 from a hardware requirement to a derived accuracy guarantee obtained by composition over the FP8 primitive.

Five-Layer Hierarchy and the TME Model

Matsuoka organizes the claim as a five-layer hierarchy: the FP8 op, Ozaki II, the basic kernels or Berkeley "dwarfs", composite solvers, and full applications. Because the dwarf taxonomy already spans scientific computing, the paper establishes the claim by exhibiting the reduction for every dwarf rather than a sample. The claim is falsifiable, and the paper builds an instrument to test it: a Tensor-Memory Equilibrium (TME) model that extends the Roofline model with emulation parameters designated as alpha, beta, and gamma.

The TME model identifies register-level fusion as the mechanism that keeps emulation memory-bound. It projects recovered FP64 performance across NVIDIA's B300 and Rubin architectures against an H100 baseline. The model could have returned a negative verdict, but according to the paper, it passes across the dwarfs and their compositions. This is the analytical half of a two-part program, with a follow-on implementation to validate the thesis on real silicon.

Implications for Enterprise HPC

For enterprise technology decision-makers evaluating high-performance computing investments, the paper suggests that hardware roadmaps favoring FP8 over FP64 may not compromise scientific accuracy. The breakdown of FP64 throughput on AI-optimized GPUs (as low as ~1.3 TFLOPS on B300) compared to FP8 throughput (multiple PFLOPS) means that relying on native FP64 could become a bottleneck. The Ozaki Scheme II offers a mathematical guarantee that FP8-based computation can match double-precision results, assuming the proposed composition is implemented efficiently.

Metric NVIDIA B300 FP64 NVIDIA B300 FP8
Throughput ~1.3 TFLOPS Multiple PFLOPS
Role in HPC Traditional requirement Proposed primitive via Ozaki II

While the paper is analytical and awaits hardware validation, it provides a framework for CTOs to reassess the necessity of FP64-capable hardware in their HPC clusters. The TME model's ability to project performance across architectures (B300, Rubin, H100 baseline) offers a tool for procurement planning.

Matsuoka's work is part of a broader trend in which AI-driven hardware optimizations are reshaping scientific computing. The paper's five-layer hierarchy and the TME model are designed to be extensible, and the author invites the community to test the claims once the follow-on implementation is released. For now, the preprint serves as a provocation: native FP64 may no longer be the holy grail of HPC, and FP8, with the right algorithmic scaffolding, could be all you need.


Sources:

Keep Reading

Recommended Stories

Former Intel CEO Pat Gelsinger Joins VC Firm to Revive Moore's Law Using Light-Based Chips Technology

Former Intel CEO Pat Gelsinger Joins VC Firm to Revive Moore's Law Using Light-Based Chips

Former Intel CEO Pat Gelsinger has joined venture capital firm Playground Global as a general partner, taking a board seat at xLight, a startup developing novel lithography techniques using light. Gelsinger believes that advances in nanometer-scale light etching can revive Moore's Law, which has stalled due to physical limits of transistor shrinking.

July 21, 2026
China Defies US Restrictions to Build World's Fastest Supercomputer LineShine Technology

China Defies US Restrictions to Build World's Fastest Supercomputer LineShine

China's new supercomputer LineShine has claimed the top spot in the TOP500 ranking, outperforming the US system El Capitan by over 20%. Built entirely with domestic hardware and software, it relies exclusively on CPUs rather than GPUs, showcasing China's ability to innovate despite US export restrictions.

June 28, 2026
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Technology

Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop

Relay, a London startup founded by two former Nothing employees, is building the Relay Q, a portable AI microphone for high-fidelity voice dictation, with software debuting now and hardware due in early 2027. The macOS-first app, powered by Google's Gemini models, adds contextual Skills that automate Slack messages and calendar entries. WIRED's hands-on found the transcription workable but less polished than Google's Pixel 11 Rambler feature, and flagged privacy trade-offs from screen-access permissions.

August 27, 2026
Nvidia sales soar above $96bn as AI data centre buildout accelerates Technology

Nvidia sales soar above $96bn as AI data centre buildout accelerates

Nvidia reported Q2 revenue of $96bn, more than double year-on-year, with data centre revenue up 117% to $89bn. CEO Jensen Huang said AI has reached its inflection point, while the company guided to $108bn in next-quarter revenue.

August 26, 2026