iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout
Home ›› Technology ›› Ai ›› Computer Vision ›› SelectStream: A Selective Memory Framework for Streaming Video Understanding

SelectStream: A Selective Memory Framework for Streaming Video Understanding

Researchers propose SelectStream, a selective latent-memory framework for streaming video models that uses surprise-driven adaptive windowing and query-conditioned graph reasoning to allocate memory efficiently. It achieves 82.67% on StreamingBench and 74.4% on offline benchmarks.

iG
iGEN Editorial
June 17, 2026
SelectStream: A Selective Memory Framework for Streaming Video Understanding

Streaming video understanding models face a fundamental challenge: they must answer queries at any point during an ongoing stream, using only what they have observed so far, while operating under fixed memory and computation budgets. Existing approaches add memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines have shown that indiscriminate history injection can dilute current-scene perception. According to a research paper published on arXiv, the key challenge is not whether to use memory, but how to allocate it selectively.

The Problem: Budgeted Online Latent Evidence Allocation

The authors — Ge, Haonan, Wang, Yiwei, Wu, Hang, and Cai, Yujun — formulate this as budgeted online latent evidence allocation. They propose a framework called SelectStream, a selective latent-memory framework that keeps the current observation directly visible to a frozen Vision-Language Model (VLM) while exposing historical information only through a compact, query-conditioned evidence budget. This design avoids replaying frames or growing the context with stream length, making it suitable for real-time applications.

SelectStream's Three Coordinated Mechanisms

SelectStream employs three coordinated mechanisms to govern when to write, what to preserve, and how to retrieve historical information:

  • Surprise-driven adaptive windowing: Determines which moments in the stream are worth remembering based on novelty or deviation from expected patterns.
  • Priority-preserving consolidation: Ensures that important historical details are retained in a fixed-capacity latent memory graph without being overwritten by less relevant information.
  • Query-conditioned graph reasoning: Retrieves relevant evidence from the memory graph based on the current query, injecting it as latent tokens for answer generation.

Retrieved evidence is calibrated and then injected as latent tokens into the frozen VLM for answer generation. The system does not replay frames or expand the context window as the stream lengthens, keeping computation constant.

Experimental Results

The paper reports that SelectStream achieves strong online streaming performance while preserving general video understanding. The specific results are:

Benchmark Score
StreamingBench 82.67%
OVO-Bench 67.03%
Offline video benchmarks (average) 74.4%

SelectStream outperforms strong recent-window baselines and prior streaming memory methods on these metrics.

Implications for Streaming Video Applications

Because SelectStream is designed to operate under fixed memory and computation budgets while answering queries at any moment during an ongoing stream, it is directly applicable to use cases such as live surveillance, autonomous driving, and real-time video analytics. The framework demonstrates that selective memory allocation, rather than exhaustive history retention, can improve both efficiency and accuracy. The paper is available on arXiv under a Creative Commons Attribution 4.0 International license.


Sources:

Keep Reading

Recommended Stories

X-Tokenizer: Semantic Action Tokenizer Boosts Robot Control by 13.5% Over FAST Technology

X-Tokenizer: Semantic Action Tokenizer Boosts Robot Control by 13.5% Over FAST

Researchers propose X-Tokenizer, a new action tokenizer that treats tokenization as semantic interface learning rather than mere compression. Using a lightweight encoder-Semantic Residual Quantization (SRQ)-decoder architecture, it improves multimodal grounding by 13.5% and long-horizon task performance by 8.25 points over existing methods like FAST.

June 16, 2026
Beijing Accuses US AI Firms of Using Chinese Models for Training Technology

Beijing Accuses US AI Firms of Using Chinese Models for Training

The Chinese commerce ministry accused US artificial intelligence firms of using Chinese models to train their own AI systems through a process called distillation. This comes after US Treasury Secretary Scott Bessent threatened sanctions against China over alleged technology theft. China defended distillation as a widely used industry practice and vowed to take all necessary measures to safeguard its interests.

July 28, 2026
project44 CEO: AI Agents Without Context Are Just Guessing Faster Technology

project44 CEO: AI Agents Without Context Are Just Guessing Faster

project44 CEO Jett McCandless argues that AI agents require rich contextual data to be effective. The company's Agentic Workflow Manager layers first- and third-party agents on top of shipment-level data to automate tasks like LTL dispatch reconciliation, processing 75,000 dispatches daily and matching over 2,000 that would otherwise require manual intervention.

July 13, 2026
Self-Improving AI Isn't Just for Frontier Labs: How Enterprises Can Build Their Own Technology

Self-Improving AI Isn't Just for Frontier Labs: How Enterprises Can Build Their Own

A journalist demonstrates building a self-improving AI using tools from Andrej Karpathy's AutoResearch and startup Prime Intellect. The experiment shows that recursive self-improvement is accessible beyond big labs, with implications for enterprises seeking specialized models.

July 8, 2026