Topic
vision-language
REVEAL++: Continuous Phenotypic Grouping Improves Vision-Language Retinal Model for Alzheimer's Risk
Researchers propose REVEAL++, a vision-language model that models phenotypic similarity as a continuous signal rather than discrete clusters, improving Alzheimer's disease risk prediction from retinal fundus images. Evaluated on UK Biobank data, it outperforms prior baselines by using a soft-target contrastive objective.
See-and-Reach: Researchers Propose 3DG-VLN for Precise UAV Vision-Language Navigation Within Field of View
Researchers introduce UAV-VLN-FOV, a target-visible navigation task that isolates the see-and-reach stage for UAVs, and propose 3DG-VLN, a vision-language waypoint prediction framework that uses dynamic 3D direction cues. The framework achieves a 13.82% improvement in success rate over baselines on a new benchmark of 2,717 trajectories.
MapDream: Task-Driven Map Learning Achieves State-of-the-Art Vision-Language Navigation
Researchers propose MapDream, a framework that learns bird's-eye-view maps directly from navigation objectives rather than hand-crafted reconstruction. The approach achieves state-of-the-art monocular performance on the R2R-CE and RxR-CE benchmarks.
FusionRS Dataset Advances Dual-Modal Vision-Language AI for Remote Sensing
Researchers introduced FusionRS, the first large-scale RGB-infrared-text dataset for dual-modal vision-language learning in remote sensing. The dataset pairs RGB and infrared images with scene and IR-aware captions, enabling models to achieve better alignment and retrieval than RGB-only approaches.
New Multi-Scale Two-Stream Framework Aims to Decouple Semantics from Distortions in AI-Generated Image Quality Assessment
Researchers introduce MST-CLIPIQA, a multi-scale two-stream vision-language framework that decouples semantic understanding from distortion detection to improve AI-generated image quality assessment. The method uses dual CLIP encoders and an information bottleneck gated fusion mechanism, achieving state-of-the-art results on five benchmarks with only 0.8 million trainable parameters.
JoyAI-VL-Interaction Model Brings Real-Time Vision-Language AI to Enterprise Applications
JoyAI-VL-Interaction is an open-source, 8B-scale vision-language model that continuously monitors video streams and decides in real time whether to stay silent, speak, or delegate to a background model. Human raters preferred it over Doubao and Gemini in six real-world scenarios. The system includes pluggable ASR/TTS, memory, and API integration.
MINT Demo 2 Framework Detects Training Data in Vision-Language Models With 90% Accuracy
Researchers introduced MINT Demo 2, a framework to determine if specific data was used to train vision-language models. The system achieves up to 90% accuracy and includes a web platform for auditing multiple model types, aiming to improve AI transparency and regulatory compliance.
RoboPIN: New AI Method Pins Chain-of-Thought to Visual Evidence for Embodied Reasoning
Researchers propose Pinned Chain-of-Thought (PINCoT), a structured reasoning paradigm that binds each reasoning step to visual evidence via reasoning anchors. The method trains a 4B parameter model that outperforms 7B open-source embodied models by 12% on 14 benchmarks, addressing issues of entity drift and decoupling in vision-language models.