iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Old Dominion nearly breaks 70% operating ratio in Q2 despite lower volumes BGN Launches US Gulf Bunkering Arm, Expanding into Direct Physical Supply of Marine Fuels Grip taps industry veteran John Hummel to lead cold chain fulfillment expansion India Retains Global Dairy Lead as USDA Forecasts Milk Output Rise to 105.4 MT in 2026 Months After Apple Warned of Low Supply, Mac Mini Shortage Persists with Long Lead Times and Price Hikes Data centres could pay hundreds of millions in deposits for power demands under Ofgem proposals Dry Bulk Volatility Is No Longer the Risk but the Business Model, Says Sagitta Marine CEO One of These Ethernet Switches Will Give Your Router the Ports You Need Zhenghe Mainline Orders Six 4,600 TEU Boxships at Hengli Shipbuilding for Baltic Service Zanskar Revives Failing Geothermal Well, Sets US Productivity Record Old Dominion nearly breaks 70% operating ratio in Q2 despite lower volumes BGN Launches US Gulf Bunkering Arm, Expanding into Direct Physical Supply of Marine Fuels Grip taps industry veteran John Hummel to lead cold chain fulfillment expansion India Retains Global Dairy Lead as USDA Forecasts Milk Output Rise to 105.4 MT in 2026 Months After Apple Warned of Low Supply, Mac Mini Shortage Persists with Long Lead Times and Price Hikes Data centres could pay hundreds of millions in deposits for power demands under Ofgem proposals Dry Bulk Volatility Is No Longer the Risk but the Business Model, Says Sagitta Marine CEO One of These Ethernet Switches Will Give Your Router the Ports You Need Zhenghe Mainline Orders Six 4,600 TEU Boxships at Hengli Shipbuilding for Baltic Service Zanskar Revives Failing Geothermal Well, Sets US Productivity Record
Home ›› Technology ›› Ai ›› Llms ›› LoRDO Algorithm Cuts Communication by 10x for Distributed AI Model Training

LoRDO Algorithm Cuts Communication by 10x for Distributed AI Model Training

LoRDO (Low-Rank Distributed Optimization) unifies low-rank optimization with infrequent synchronization to reduce communication overhead in distributed training of foundation models. According to an arXiv paper, it achieves near-parity with low-rank DDP at scales 125M–720M parameters while cutting communication by approximately 10x, and shows further gains in very low-memory settings.

iG
iGEN Editorial
July 8, 2026
LoRDO Algorithm Cuts Communication by 10x for Distributed AI Model Training

Distributed training of large AI models faces a fundamental bottleneck: interconnect bandwidth limits how fast workers can synchronize gradients. Infrequent communication strategies reduce synchronization frequency, but optimizer states still consume significant memory and bandwidth. A new algorithm, LoRDO, addresses this by unifying low-rank optimization with infrequent synchronization, offering a practical path to more efficient distributed training.

According to the arXiv paper by Jovanović et al., LoRDO (short for Low-Rank Distributed Optimization) is designed to overcome the limitations of low-rank optimizers in the local-update regime. When workers train locally, they lack access to the full-batch gradients required to compute low-rank projections, degrading performance. The authors first demonstrate that global projections based on pseudo-gradients are theoretically superior but permanently restrict the optimization trajectory to a low-rank subspace.

How LoRDO Restores Subspace Exploration

To restore the ability to explore beyond this subspace, LoRDO introduces a full-rank quasi-hyperbolic update. This step allows the optimization path to escape the constrained low-rank subspace while still benefiting from the memory and communication savings of low-rank methods. The result is a principled framework that balances efficiency with model quality.

Experimental Results: Near-Parity with 10x Less Communication

The paper reports results on language modeling and downstream tasks at model scales ranging from 125 million to 720 million parameters. LoRDO achieves near-parity with low-rank Distributed Data Parallel (DDP) training, while reducing communication by approximately 10x. The authors also note that LoRDO improves performance even more in very low-memory settings where rank and batch size are small.

Metric Low-Rank DDP LoRDO
Model scales tested 125M–720M 125M–720M
Communication overhead Baseline ~10× reduction
Performance vs. low-rank DDP Baseline Near-parity
Additional gains in low-memory N/A Yes

Implications for Enterprise AI Training

For enterprises training foundation models or large language models, communication efficiency directly translates to lower infrastructure costs and faster iteration cycles. LoRDO’s ability to maintain model accuracy while slashing bandwidth requirements makes it attractive for organizations with limited interconnect resources or those operating across distributed data centers. The algorithm is particularly promising for scenarios where memory constraints are severe, such as on-device or edge training.

The authors—Jovanović, Andrej, Iacob, Alex, Safaryan, Mher, Modoranu, Ionut-Vlad, Sani, Lorenzo, Shen, William F, Qiu, Xinchi, Alistarh, Dan, and Lane, Nicholas D—have published the work on arXiv under a Creative Commons license, allowing the research community to build on these ideas. While LoRDO is still an algorithmic advance rather than a commercial product, it points to a future where distributed training can scale more efficiently without sacrificing model quality.


Sources:

Keep Reading

Recommended Stories

Beijing Accuses US AI Firms of Using Chinese Models for Training Technology

Beijing Accuses US AI Firms of Using Chinese Models for Training

The Chinese commerce ministry accused US artificial intelligence firms of using Chinese models to train their own AI systems through a process called distillation. This comes after US Treasury Secretary Scott Bessent threatened sanctions against China over alleged technology theft. China defended distillation as a widely used industry practice and vowed to take all necessary measures to safeguard its interests.

July 28, 2026
project44 CEO: AI Agents Without Context Are Just Guessing Faster Technology

project44 CEO: AI Agents Without Context Are Just Guessing Faster

project44 CEO Jett McCandless argues that AI agents require rich contextual data to be effective. The company's Agentic Workflow Manager layers first- and third-party agents on top of shipment-level data to automate tasks like LTL dispatch reconciliation, processing 75,000 dispatches daily and matching over 2,000 that would otherwise require manual intervention.

July 13, 2026
Self-Improving AI Isn't Just for Frontier Labs: How Enterprises Can Build Their Own Technology

Self-Improving AI Isn't Just for Frontier Labs: How Enterprises Can Build Their Own

A journalist demonstrates building a self-improving AI using tools from Andrej Karpathy's AutoResearch and startup Prime Intellect. The experiment shows that recursive self-improvement is accessible beyond big labs, with implications for enterprises seeking specialized models.

July 8, 2026
Bi-Anchor Interpolation Solver Cuts Generative Modeling Steps from 100 to 10, Researchers Show Technology

Bi-Anchor Interpolation Solver Cuts Generative Modeling Steps from 100 to 10, Researchers Show

Researchers introduce the Bi-Anchor Interpolation Solver (BA-solver) for accelerating flow matching generative models. It achieves quality comparable to 100+ step solvers in just 10 steps, using a small SideNet (1-2% of backbone size) and novel bidirectional temporal perception. The method is plug-and-play with existing pipelines.

July 8, 2026