Topic
deep learning
New Graph Neural Network Learns Protein Representations with Secondary Structure and Energy-Filtered Hydrogen Bonds
Researchers propose a secondary-structure-aware graph neural network for protein representation learning. The model augments residue-level node representations with secondary structure assignments and constructs edges from hydrogen-bond interactions filtered by energetic strength. It achieves consistent improvements over existing methods on standard protein benchmarks and offers enhanced biological interpretability.
New Research Shows Pretraining Data Composition Can Engineer Neural Scaling Laws for Particle Physics
A new arXiv paper demonstrates that neural scaling laws in particle physics can be engineered by adjusting pretraining data composition. The study shows that including more diverse and task-aligned synthetic data can shift scaling behavior to require more data rather than larger models, offering insights for efficient AI training.
Bi-Anchor Interpolation Solver Cuts Generative Modeling Steps from 100 to 10, Researchers Show
Researchers introduce the Bi-Anchor Interpolation Solver (BA-solver) for accelerating flow matching generative models. It achieves quality comparable to 100+ step solvers in just 10 steps, using a small SideNet (1-2% of backbone size) and novel bidirectional temporal perception. The method is plug-and-play with existing pipelines.
EEG Foundation Models Show Promise for Burst-Suppression Detection in ICU Without Patient-Specific Calibration
A new study on arXiv evaluates three EEG foundation models—REVE-base, LUNA-large, and LuMamba-Tiny—for automatic burst-suppression detection in ICU patients, finding REVE-base achieves the highest event-based F1-score (0.868) and reduces burst-per-minute error by 52.1% compared to a task-specific EEGNet baseline.
Emyx: New AI Model Generates All-Atom Proteins Faster and More Efficiently
Researchers have developed Emyx, a 140M-parameter conditional flow matching model for all-atom protein generation. Despite being the smallest model, Emyx outperforms both Proteína-Complexa and RFdiffusion3 on the AME enzyme design benchmark across success rate, structural novelty, scaffold diversity, and geometric validity, while training in just 682 GPU-hours—roughly 4× less than RFdiffusion3.
New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models
A new computer vision paper from arXiv investigates how visual tokens are integrated into large language models (LLMs) under two paradigms: in-context prompting and layer-wise injection. The authors find that visual tokens enter the LLM as 'disguised visual context' lacking linguistic structure, then evolve differently depending on the integration architecture. They show that attention allocation alone is insufficient, and performance depends on the quality of visual representations at each layer.
FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching
FlowMaps, a latent flow matching model, predicts multimodal distributions of future object locations in 3D space by learning from past human interactions. Tested in over 600 episodes, it outperforms state-of-the-art approaches for dynamic Object Navigation tasks in simulated and real environments. The research, published on arXiv, has potential applications for robotics in changing environments.
DiverseDistill: New Knowledge Distillation Method Recovers Over 70% of Performance Gap Using Teacher Committees
Researchers propose DiverseDistill, a knowledge distillation framework that combines a large foundation model with domain-specific experts as a diverse committee. The method recovers 73–114% of the teacher-student performance gap on recommendation and vision tasks while requiring no parameter updates or architectural changes.
New Framework for Class-Incremental Motion Forecasting Enables Autonomous Vehicles to Adapt to Novel Objects
Researchers introduce class-incremental motion forecasting, a setting where autonomous vehicles learn new object classes over time. They propose the first end-to-end framework that adapts to novel classes while mitigating catastrophic forgetting, using pseudo-labels and open-vocabulary segmentation. Evaluations on nuScenes and Argoverse 2 show preserved performance on known classes and effective adaptation to new ones.
Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
Researchers propose an object-centric OOD detection framework that leverages object co-occurrence patterns to overcome simplicity bias, achieving competitive results on near-OOD and full-spectrum settings.
New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs
Multimodal Large Language Models (MLLMs) traditionally lack intrinsic 3D awareness. Researchers present GeoVR, a framework that learns geometric representations from 2D video sequences, restructuring the semantic latent space to unlock spatial intelligence. GeoVR uses four complementary geometric targets from pre-trained 3D foundation models, achieving state-of-the-art performance on spatial reasoning benchmarks.
VCG: Multimodal Retrieval Framework Solves Extreme Cold-Start Problem for E-Commerce Video Feeds
E-commerce platforms are shifting to video feeds but face extreme cold-start problems because new videos lack interaction history. The VCG system, described in a recent arXiv paper, uses a CLIP-based multimodal retrieval engine to map users and videos into a shared semantic space, enabling zero-shot retrieval. Online A/B testing showed a 50% uplift in deep video completion, demonstrating effective mitigation of engagement biases.
Technology Yann LeCun's new AI startup AMI Labs raises $1bn to build flexible intelligence beyond LLMs
Yann LeCun, former Meta chief AI scientist, has founded AMI Labs to develop a new AI architecture called JEPA, which aims to overcome the limitations of large language models (LLMs) in understanding the physical world. The startup raised over $1bn in seed funding from Nvidia and Jeff Bezos' private investment fund, marking one of Europe's largest seed rounds.
Controlled Benchmark Finds No Quantum Advantage in Brain MRI Data Augmentation
A controlled benchmark study by Haider and Figini shows that quantum-latent GAN augmentation does not improve brain MRI classification over real-data-only training or classical GANs. The quantum and classical generators were statistically indistinguishable across all data fractions from 5% to 100%.
QC-GAN: Parameter-Efficient Speech Enhancement Model Delivers High Fidelity with 0.89M Parameters
A new speech enhancement framework, QC-GAN, combines a Quaternion Conformer generator with MetricGAN-based training to deliver state-of-the-art perceptual quality using remarkably few parameters. The model achieves a PESQ score of 3.48 with only 0.89M parameters, and a 35K-parameter variant reaches 3.23, outperforming conventional methods at a fraction of the size. This parameter efficiency makes it suitable for edge deployment in voice-controlled systems, including logistics and supply chain applications.
LLM Paraphrase Augmentation Boosts Sign Language Translation Performance
A new study proposes using a large language model (GPT-4o) to generate controlled paraphrase variants of training targets for sign language translation (SLT). Evaluated on three datasets, the method yields a modest BLEU-4 improvement on PHOENIX14T and reveals gains in semantic fidelity not captured by lexical metrics.
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models
A new research paper from arXiv shows that reinforcement learning with verifiable rewards (RLVR) can cause large reasoning models to forget foundational capabilities like perception and faithfulness. The authors propose RECAP, a replay strategy with dynamic objective reweighting that preserves general knowledge while maintaining reasoning gains.
FreeStyle: Scalable Style-Content Dual-Reference Generation via Community LoRA Mining
FreeStyle is a scalable dual-reference generation framework that leverages community LoRAs as compositional anchors for style and content. It introduces a two-stage curriculum with attention-level enrichment and frequency-aware RoPE modulation to suppress leakage from style references. The framework is evaluated on a new benchmark covering style similarity, content preservation, and leakage rejection, achieving a strong balance among these objectives.
Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation
Researchers Makarov and Gerkmann propose a method to repurpose a conventionally trained speech classifier as the backbone for diffusion-based speech generation. By attaching a lightweight subnetwork and training only that under a Denoising Score Matching objective, they achieve high-quality speech synthesis with reduced memory footprint and computational cost compared to traditional classifier guidance that requires two separately trained models.
New AI Research Shows Vision-Language Models Think Better with Visual Grounding
Researchers introduce visually grounded thinking, a reasoning process that interleaves natural-language thoughts with explicit point or box groundings to image regions. The method, using a scalable synthesis pipeline and grounding-aware reinforcement learning, consistently improves performance of Gemma3-4B-IT on counting and spatial reasoning benchmarks, with the 4B model matching or surpassing the 27B variant.
HilDA: Hierarchical Distillation with Diffusion Advances Self-Supervised LiDAR Pre-Training
Researchers propose HilDA, a self-supervised pretraining framework for LiDAR backbones that uses hierarchical distillation and temporal occupancy diffusion. The method achieves state-of-the-art results on cross-modal distillation benchmarks for 3D object detection, scene flow, and semantic occupancy prediction.
MakeupMirror Model Boosts Facial Attribute Preservation in Diffusion-Based Makeup Transfer
Researchers propose MakeupMirror, a diffusion-based makeup transfer model that preserves facial identity and skin tone better than previous solutions. It achieves 60% higher facial recognition similarity, 50% lower skin tone difference, and 0.7s latency, with 94% expert acceptance, advancing virtual try-on for e-commerce.
DF3DV-1K: Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Researchers introduced DF3DV-1K, a large-scale real-world dataset with 1,048 scenes and 89,924 images for distractor-free novel view synthesis. The dataset spans 128 distractor types and 161 scene themes, enabling benchmarking of nine radiance field methods and 3D Gaussian Splatting. Fine-tuning a diffusion-based 2D enhancer on DF3DV-1K achieved average improvements of 0.96 dB PSNR and 0.057 LPIPS.
StreamKL Delivers up to 43× Speedup in Memory-Efficient Attention Distillation
Researchers propose StreamKL, a fused GPU primitive for Kullback-Leibler divergence in attention distillation. It eliminates quadratic memory materialization, enabling up to 43× and 14× speedups in forward and backward passes, and reduces extra HBM footprint to O(1).
SL-S4Wave: Self-Supervised Learning Framework Improves ECG and EEG Analysis with State Space Models
Researchers propose SL-S4Wave, a self-supervised learning framework combining contrastive learning with structured state space models (S4) to analyze long-sequence physiological waveforms. The model outperforms state-of-the-art baselines in arrhythmia detection and EEG tasks, demonstrates strong label efficiency, and generalizes to unseen arrhythmia types.
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods
A team of researchers has introduced PAVE (Policy-Aware Value-field Equalization), a critic-centric regularization framework that stabilizes the Q-gradient field in continuous actor-critic reinforcement learning. The method addresses erratic high-frequency oscillations in learned policies without modifying the actor, achieving smoothness comparable to policy-side regularization while maintaining task performance.
REST-GAN: A Deep Generative Model for Resting-State EEG Synthesis and Transferable Representation Learning
Researchers introduce REST-GAN, a generative adversarial network for resting-state EEG that both synthesizes realistic neural signals and learns transferable representations. The model achieves high precision and recall in band-power features and shows competitive performance in demographic classification tasks, requiring substantially less training data and computational resources than existing methods.
Hierarchical BART strategy achieves state-of-the-art Vietnamese multi-document summarization
A research team presents a novel hierarchical BART-based strategy for Vietnamese multi-document abstractive summarization, achieving a ROUGE2-F1 score of 0.2468 on the VLSP 2022 public test set. The approach condenses documents guided by a golden summary, producing fluent and concise outputs, and releases additional training data to the community.
Transformer Feed-Forward Block Linearity: Learned, Not Architectural, According to New Research
A new study introduces R^2_lin, a measure of linearity for transformer feed-forward blocks. Across models like GPT-2 and Pythia-160m, R^2_lin varies widely and is not determined by activation function. The findings offer targeted compression signals and reveal pitfalls in training linear baselines.
Boundary Embedding Shaping with Adaptive Contrastive Learning Boosts GNN Classification by 3.3%
Graph neural networks suffer from structural entanglement, especially near class boundaries. A new plug-in module called Boundary Embedding Shaping (BES) uses adaptive contrastive learning to selectively suppress spurious correlations, boosting GCN node classification by an average of 3.3% (up to 5% on WikiCS) and improving link prediction accuracy.
New Tokenization Method Merges Tokens to Improve Diffusion Transformer Efficiency
A research paper introduces a variable-length tokenizer that merges tokens instead of truncating them, enabling adaptive compression for diffusion transformers. The method, called learnable global merging, addresses representational alignment issues across token lengths and achieves a superior trade-off between image quality (gFID) and computational cost.
CSWinUNETR: Deep Learning Model Segments Thin Anatomical Structures with Cross-Shaped Self-Attention
Researchers propose CSWinUNETR, a deep learning backbone for 2D and 3D segmentation of thin anatomical structures such as retinal vessels, cerebral vasculature, and facial wrinkles. The model employs cross-shaped stripe self-attention, cyclic shifts, and sparse-control dynamic snake convolution to improve segmentation accuracy. It outperforms state-of-the-art methods on four benchmarks without task-specific post-processing.
Interpretable Sperm Morphology Classification via Attention-Guided Deep Learning
A study proposes an interpretable deep learning framework combining EfficientNet-B0 with a Convolutional Block Attention Module for sperm morphology classification, achieving 90.2% and 93.9% accuracy on SMIDS and HuSHem datasets respectively.
Large Language Models Can Read Compressed Text That Humans Cannot, Researchers Find
A new research paper introduces BabelTele, a compact, non-human-readable text format that large language models can still interpret with high semantic fidelity. The approach compresses text to 27.9% of its original length while preserving 99.5% of meaning, potentially reducing context overhead and costs in enterprise AI deployments.
Adaptive Binning Boosts Self-Supervised Learning on Medical Tabular Data, Researchers Report
Researchers propose Adaptive Binning, a training-adaptive discretization pretext for tabular self-supervised learning. The method progressively refines discretization per feature and uses a heterogeneity-aware objective. Experiments on public medical tabular datasets show consistent gains over fixed binning approaches.
Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices
A new paper introduces four complementary techniques to reduce peak memory during LoRA fine-tuning of large language models on edge devices. Experiments on Llama-3.2 3B and Qwen-2.5 3B demonstrate up to 26x and 28x memory reduction, respectively, without sacrificing model quality.
AutoPass: Evidence-Guided LLM Agents Achieve Compiler Speedups of 1.117x on ARM64
Researchers present AutoPass, a multi-agent LLM framework that uses compiler and runtime evidence to guide compiler optimization decisions. Without training, it outperforms expert-tuned heuristics and classical autotuning, achieving geometric-mean speedups of 1.043x on x86-64 and 1.117x on ARM64 over LLVM -O3.
New Training-Free Method Compresses Vision-Language-Action Models by 50% Without Performance Loss
A research team led by Gia-Binh Ho et al. discovered that Vision-Language-Action (VLA) models exhibit severe layer-wise redundancy. They introduced a training-free compression pipeline using Centered Kernel Alignment to remove twin layers, achieving up to 50% depth reduction, 40-50% faster fine-tuning, and 30% faster inference while matching or exceeding full-scale performance.
LLM-Driven Stepwise Refinement Framework Promises Verifiable Hardware Generation
A new framework from researchers Li et al. combines large language models with formal methods to generate verifiable hardware designs. By applying stepwise transformation rules, the LLM agent produces correct register-transfer level (RTL) programs, addressing the reluctance of engineers to trust AI in high-stakes chip design.
IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources
Researchers present IHUBERT, a monolingual Persian language model pretrained on a 45GB curated subset of the Sepahr-Danesh collection using a multi-stage pipeline that includes vector-database-based semantic deduplication and domain-balanced pretraining. IHUBERT achieves top scores on extractive QA benchmarks PQuAD and ParsiNLU-RC, and best results on FarsTail NLI, while remaining competitive on NER and topic classification.
Triangular Consistency Constraint Offers Universal Plug-and-Play Component for Optical Flow Learning
Researchers propose triangular consistency, a first-principled constraint for optical flow that is agnostic to network architecture, supervision type, and dataset. The constraint composes two flows to induce a third and enforces consistency, showing consistent improvement across supervised, unsupervised, and transfer learning with negligible computational overhead.
BrainG3N Tokenizer Enables Controllable 3D Brain MRI Generation with Clinical-Grade Embeddings
BrainG3N, a novel tokenizer for 3D brain MRI latent diffusion, decouples encoder and decoder to preserve clinical information while enabling high-quality reconstruction. Pretrained on 35,309 volumes, it outperforms SOTA models on 21 of 23 clinical tasks and supports controllable generation for disease simulation and privacy-preserving data sharing.
Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models
Token Factory is a framework that converts diverse traditional signals into soft tokens for large recommendation models (LRMs), addressing challenges of long prompts, memory footprint, and computational overhead. The approach has been validated in a production-scale environment, promising enhanced performance and efficiency.
First Billion-Parameter Generative Foundation Model for Chest Radiography Achieves Expert-Level Synthesis Fidelity
Ribeiro et al. present the largest specialist generative foundation model for chest radiographs, with over 1.3 billion parameters. Trained on 1.2 million radiographs, the model supports controllable generation across demographics, views, and pathologies, advancing synthesis fidelity to clinical indistinguishability.
ITNet: A Learnable Integral Transform That Unifies Convolution, Attention, and Recurrence in One Architecture
Researchers introduce ITNet, a neural architecture built on a learnable integral transform that subsumes convolution, self-attention, and recurrence as special cases. A single ITNet with a shared operator matches or exceeds specialized models on ImageNet-1K, GLUE, ModelNet40, VQA v2, and NLVR2, enabled by efficient techniques such as tiled kernel fusion and Monte Carlo integration.
Diffusion Language Models Show Promise but Demand Careful Inference Tuning, Study Finds
A new systematic study from researchers analyzes eight state-of-the-art Diffusion Language Models (DLMs) across eight benchmarks covering reasoning, coding, translation, and more. The research highlights how inference-time choices like denoising steps and context length create trade-offs between generation quality and computational efficiency, offering guidance for enterprise deployment.
New AI Framework PSCT-Net Reduces Radiation Risk in Pediatric Skull CT Imaging
PSCT-Net is a novel deep learning framework for reconstructing 3D CT scans of pediatric skulls from only two X-ray images, significantly reducing radiation exposure. The method uses differentiable back-projection and attention-guided refinement to overcome depth ambiguity. It was evaluated on a private dataset called PedSkull-CT.
Concept Flow Models Anchor AI Reasoning with Hierarchical Bottlenecks to Reduce Information Leakage
Researchers Wang and Paschke propose Concept Flow Models (CFMs) that replace the flat bottleneck in Concept Bottleneck Models (CBMs) with a hierarchical, concept-driven decision tree. CFMs mitigate information leakage by reducing effective concept usage, matching predictive performance of flat CBMs while providing stepwise decision flows for transparent and auditable model reasoning.
Researchers Identify Shrinkage Bias in LLM FP4 Pretraining, Propose UFP4 Recipe for Stability
A new study from researchers on arXiv identifies 'Shrinkage Bias' in E2M1-based FP4 pretraining for large language models, a systematic error that accumulates across layers. The proposed UFP4 recipe, using uniform grids like E1M2/INT4, demonstrates lower BF16-relative loss degradation on models up to 124B parameters, urging hardware support for uniform 4-bit formats.
FastMix: Gradient-Based Data Mixture Optimization Reduces Search Cost in AI Training
FastMix is a novel framework that automates data mixture discovery by training only a single proxy model and jointly optimizing mixture coefficients and model parameters via gradient descent. It reformulates mixture selection as a bilevel optimization problem, enabling efficient, scalable optimization that outperforms baselines.
New Temporal Pyramid Model Enhances Spoofed Speech Detection for Voice Security Systems
Researchers introduced a Temporal Pyramid Adapter for spoofed speech detection that uses parallel temporal convolutions with varying receptive fields to capture multi-scale cues. The model achieved a 99.24% AUC and 3.87% EER on the PartialSpoof dataset, significantly outperforming existing methods like LCNN-BLSTM (9.87% EER) and TRACE (8.08% EER). The work highlights the potential for improving voice authentication security but notes performance degradation under domain and language shifts.
New Study Challenges Prior Claims on Scaling Context Length in Imitation Learning
Researchers evaluated diffusion policies for robotic imitation learning across varying context lengths, challenging prior claims that long-context scaling is fragile. They propose a training algorithm that jointly trains policies at multiple context lengths, reducing sample complexity.
New AI Framework Synthesizes Fluorescein Angiography from Fundus and Sparse OCT Scans
A research team led by Ma introduced a novel deep learning framework that synthesizes fluorescein angiography (FFA) from color fundus photography (CFP) using structural guidance from sparse optical coherence tomography (OCT) scans. The method uses a tri-modally aligned dataset of 3,676 patient eyes and achieves superior synthesis and downstream diagnosis performance compared to existing methods.
OmniMouse Brain Model Trained on 150 Billion Neural Tokens Reveals Unusual Scaling Laws
Researchers trained OmniMouse, a multi-modal, multi-task brain model, on 150 billion neural tokens from 3.1 million mouse visual cortex neurons. The model achieves state-of-the-art performance across neural prediction, behavioral decoding, and neural forecasting. Scaling analysis shows performance improves with more data but gains from increasing model size saturate, contrasting with language and vision AI.
New AI Research Analyzes When Score-Based Models Outperform Traditional Channel Estimation
A new paper from Skocaj, Eller, and Boban provides a theoretically grounded analysis of score-based generative models for channel estimation in wireless communications. The study uses the perception-distortion tradeoff to reveal when score-matching offers advantages over traditional discriminative learning, with numerical results showing benefits under high predictive uncertainty but recommending simpler approaches otherwise.
Norm-Agnostic Residual Networks Offer Path to Scaling Adaptive Depth in Deep Learning
Researchers introduce NAG, a norm-agnostic residual architecture that prevents later layers from being suppressed by norm growth. This enables training of much deeper models and introduces an interpretable Mixture-of-Depths mechanism that can serve as a pretraining scaling strategy, with 20-25% sparsity matching full-depth baseline under equal compute.
S-SPPO: Semantic Calibration Boosts LLM Preference Alignment Without Human Data
S-SPPO, a dual-space semantic calibration framework, fixes instability in Self-Play Preference Optimization (SPPO) for large language models. By annealing win targets and enforcing geometric diversity, it achieves superior alignment results on AlpacaEval 2.0 without extra human preferences.
Lightweight Attention Mechanism Boosts Robust Multimodal Integration in Global Workspace Architecture
A new arXiv paper introduces a lightweight attention mechanism for multimodal integration in a global workspace architecture. The method improves robustness against corrupted modalities while using far fewer trainable parameters than end-to-end attention baselines. Tests on Simple Shapes and MM-IMDb 1.0 show transferable selection strategies across tasks and unseen modalities.
From Tokens to Regions: CUDA-Sensitive Instruction Tuning for GPU Kernel Generation
A new method called CuSeT (CUDA-Sensitive Instruction Tuning) addresses the struggle of large language models to generate correct CUDA kernels. By combining adaptive token-level masking with region-aware sample reweighting, CuSeT improves functional correctness across multiple model families and scales, achieving competitive performance against frontier CUDA kernel generation models at lower inference cost.
FOUNDv2: Unified Quantized Tokenizers Transform User Representation Learning
FOUNDv2, a novel user representation framework, uses quantized tokenizers to transform heterogeneous data into discrete tokens, achieving superior performance and reduced storage costs. The model outperforms task-specific baselines and has been deployed at scale on Alipay, demonstrating practical efficiency.