iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing
Home ›› Technology ›› Ai ›› Computer Vision ›› Modality-Aware Novelty Detection Framework MAND Improves Open-World Egocentric Activity Recognition

Modality-Aware Novelty Detection Framework MAND Improves Open-World Egocentric Activity Recognition

A new research paper introduces MAND, a modality-aware framework for multimodal egocentric open-world continual learning. MAND addresses limitations of existing methods that underutilize IMU cues and suffer from catastrophic forgetting, leading to improved novelty detection and known-class accuracy on a public benchmark.

iG
iGEN Editorial
June 16, 2026
Modality-Aware Novelty Detection Framework MAND Improves Open-World Egocentric Activity Recognition

Multimodal egocentric activity recognition, which combines visual and inertial cues to understand first-person behavior, faces significant hurdles when deployed in open-world environments. According to a paper on arXiv, existing methods struggle to detect activities never seen before while continuously learning from non-stationary data streams. The authors propose MAND (Modality-Aware Novelty Detection), a framework that adaptively leverages complementary evidence from multiple modalities to improve reliability.

The Problem with Existing Approaches

Traditional multimodal systems rely on the main fused logits for novelty scoring, according to the paper. This approach fails to fully exploit the complementary evidence available from individual modalities. Because these logits are often dominated by RGB, cues from other modalities—particularly IMU (inertial measurement unit)—remain underutilized. The paper notes that this imbalance worsens as catastrophic forgetting accumulates, where neural networks overwrite previously learned knowledge when integrating new tasks.

MAND: Dual Mechanism for Adaptive Learning

MAND introduces two key components. At inference, the Modality-aware Adaptive Scoring (MoAS) mechanism adaptively adjusts modality contributions using sample-wise reliability. It refines novelty scoring with deviation and disagreement penalties, ensuring that less reliable modalities are downweighted. During training, Modality-aware Representation Stabilization Training (MoRST) preserves the discriminative capacity of each modality across tasks. This is achieved through modality-specific heads and modality-wise logit distillation, preventing catastrophic forgetting.

Experimental Results

The authors tested MAND on a public multimodal egocentric benchmark. The results show that MAND consistently improves novel activity detection and known-class accuracy while substantially reducing FPR95 (false positive rate at 95% recall). This indicates more reliable open-world recognition compared to existing methods. The source code is publicly available at the link in the paper.

Metric Existing Methods MAND
Novel activity detection Baseline Improved
Known-class accuracy Baseline Improved
FPR95 Higher Substantially reduced

The research was conducted by Im, Hyejeong; Lim, Wonseon; and Kim, Dae-Won. The paper is titled "MAND: Modality-Aware Novelty Detection for Open-World Egocentric Activity Recognition."

Implications for Enterprise AI

While the research is academic, the ability to detect novel activities in first-person video with multimodal data has relevance for enterprise systems that require anomaly detection, such as monitoring worker actions in manufacturing or logistics. The MAND framework's focus on robustness and adaptability aligns with the needs of open-world deployments where unseen events must be detected reliably without manual retraining.

The publication on arXiv and the availability of source code enable further exploration and adoption by the research community.


Sources:

Keep Reading

Recommended Stories

Akasha 2 Achieves 4x Faster Visual Synthesis with Hamiltonian-Inspired AI Architecture Technology

Akasha 2 Achieves 4x Faster Visual Synthesis with Hamiltonian-Inspired AI Architecture

Akasha 2 introduces Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architecture, achieving state-of-the-art video prediction with 4x faster synthesis than diffusion models and 3-18x speedup over transformers. The system enforces physical conservation laws for spatiotemporal coherence.

June 16, 2026
New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models Technology

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models

A new computer vision paper from arXiv investigates how visual tokens are integrated into large language models (LLMs) under two paradigms: in-context prompting and layer-wise injection. The authors find that visual tokens enter the LLM as 'disguised visual context' lacking linguistic structure, then evolve differently depending on the integration architecture. They show that attention allocation alone is insufficient, and performance depends on the quality of visual representations at each layer.

July 8, 2026
New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs Technology

New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs

Multimodal Large Language Models (MLLMs) traditionally lack intrinsic 3D awareness. Researchers present GeoVR, a framework that learns geometric representations from 2D video sequences, restructuring the semantic latent space to unlock spatial intelligence. GeoVR uses four complementary geometric targets from pre-trained 3D foundation models, achieving state-of-the-art performance on spatial reasoning benchmarks.

July 8, 2026
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation Technology

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

CADBench is a unified benchmark for multimodal CAD program generation, containing 18,000 evaluation samples across six benchmark families, five input modalities, and six metrics. The benchmark evaluates eleven AI systems, generating over 1.4 million CAD programs, and reveals key failure modes in current approaches.

June 21, 2026