iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing
Home ›› Technology ›› Ai ›› Llms ›› Scribby Multi-Level LLM Framework Promises Fine-Grained Semantic Analysis of Long-Form Video

Scribby Multi-Level LLM Framework Promises Fine-Grained Semantic Analysis of Long-Form Video

Researchers propose Scribby, an LLM-based framework for semantic video analysis that balances macro-level comprehension with micro-level semantic indexing. The approach analyzes full transcripts, individual sentences, and groups sentences by semantic similarity using an LLM as a judge, enabling more detailed understanding of video structure and thematic progression.

iG
iGEN Editorial
June 16, 2026
Scribby Multi-Level LLM Framework Promises Fine-Grained Semantic Analysis of Long-Form Video

As video content continues to expand across educational platforms, recorded lectures, and live-streamed entertainment, the need for efficient and structured analysis of long-form footage has increased, according to a new arXiv preprint. However, many existing AI programs provide only high-level video summaries based on AI-generated transcripts, which are often limited to coarse overviews and lack detailed analysis of a video's structure, thematic progression, and semantic relationships.

Scribby: A Multi-Level LLM Framework aims to address this gap by proposing an LLM-based video summarization framework that balances macro-level comprehension with micro-level semantic analysis. The framework, detailed in the paper by Abelarde, Julian, Belinchon, and Hugo Garrido-Lestache, establishes a foundation for video analysis tools that visualize semantic chunking and semantic matching through relevance-based heatmaps.

Technical Approach: Micro-Level Indexing with LLM as Judge

The first stage of the Scribby process indexes the video at a micro level through three steps:

  1. Analyzing the full transcript at a global level
  2. Analyzing individual transcript sentences
  3. Grouping these sentences by semantic similarity using an LLM as a judge

Contextual continuity is retained during sentence-level processing by incorporating both the global transcript analysis and adjacent sentence information into each evaluation prompt. This approach ensures that micro-level understanding is grounded in the broader narrative of the video.

Step Description
1 Full transcript analysis (macro-level context)
2 Individual sentence analysis
3 Sentence grouping by semantic similarity via LLM-as-judge

The use of an LLM as a judge — where the language model evaluates semantic similarity — is a key innovation, allowing the framework to capture nuanced relationships between segments without requiring pre-defined categories.

Potential Applications and Limitations

The framework is designed for comprehensive video analysis, particularly for content where structural and thematic details matter, such as educational lectures or recorded presentations. The paper discusses limitations and future expansions of the framework, though specific application domains beyond general semantic analysis are not detailed in the source.

As the authors note, the work establishes a foundation for tools that can visualize semantic chunking and relevance-based heatmaps, pointing toward future interactive analytical interfaces. The paper is available on arXiv under a Creative Commons Zero license.


Sources:

Keep Reading

Recommended Stories

New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs Technology

New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs

Multimodal Large Language Models (MLLMs) traditionally lack intrinsic 3D awareness. Researchers present GeoVR, a framework that learns geometric representations from 2D video sequences, restructuring the semantic latent space to unlock spatial intelligence. GeoVR uses four complementary geometric targets from pre-trained 3D foundation models, achieving state-of-the-art performance on spatial reasoning benchmarks.

July 8, 2026
M*: A Modular, Extensible Serving System for Efficient Multimodal AI Inference Technology

M*: A Modular, Extensible Serving System for Efficient Multimodal AI Inference

Researchers have developed M*, a universal serving system for composite AI models that integrates diverse components like vision encoders and language backbones. Using a novel 'Walk Graph' abstraction, M* achieves significant performance improvements: 20% lower latency for text-to-image, up to 2.7x higher throughput for text-to-speech, and 12.5x faster robotic planning rollouts compared to existing baselines.

June 16, 2026
Wasserstein Equilibrium Decoding Boosts Reliability in Medical Visual Question Answering Technology

Wasserstein Equilibrium Decoding Boosts Reliability in Medical Visual Question Answering

Researchers have extended game-theoretic decoding to vision-language models for medical visual question answering, introducing a Wasserstein stopping criterion that improves accuracy by up to 3.5 percentage points and reduces inference iterations by 20% while maintaining reliability.

June 16, 2026
Language-Guided AI Framework CLARITY Boosts Road Scene Segmentation for Autonomous Logistics Technology

Language-Guided AI Framework CLARITY Boosts Road Scene Segmentation for Autonomous Logistics

Researchers propose CLARITY, a language-guided framework for RGB-Thermal semantic segmentation that dynamically adapts fusion strategies based on scene illumination. On the MFNet dataset, it achieves 62.3% mIoU and 77.5% mAcc, setting a new state-of-the-art for robust road scene understanding in autonomous driving, critical for logistics automation.

June 16, 2026