Skip to yearly menu bar
Skip to main content
Main Navigation
Select Year: (2026)
2026
2024
2022
My Stuff
Create Profile
Reset Password
Login
Getting Started
Schedule
Tutorials
Workshops
Main Conference
Keynotes and Panels
Orals
Spotlights
Papers
Paper Awards
Sponsors
Organizers
Help
Layout:
mini
compact
topic
detail
×
No topics available
No sessions available
title
author
topic
session
shuffle
by
serendipity
bookmarked first
visited first
not visited first
bookmarked but not visited
Loading...
Enable Javascript in your browser to see the papers page.
AVQ-Attention: Adaptive Vector-Quantized Attention
From Local to Global: A Progressive Reconstruction Network for Diffractive Snapshot Spectral Imaging
CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
Sim, Yet Same: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling under Extreme Motion Blur
Sticking Information in Plain Sight: Encoding and Detecting Hidden Stickers in the Real World
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization
MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots
Boosting 3D Foundation Models with Featureless Pose Optimization
GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
Ada-VNNs: Adaptive Equivariance for Vector Neural Networks
HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy
EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage
Disentangling Rotation and Translation from SE(3)-Equivariant Features for Shape Assembly
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
Generative Relightable Avatars
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
SFM: Taming State Space Models for Text-to-Motion via Spatial-Frequency Modeling
Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
GR-GRPO: Graph-Diffused Credit for Autoregressive Image RL Alignment
HandSCS: Structural Coordinate Space for Animatable Hand Gaussian Splatting
EVEE: Event-Based Online Adaptation for Matching on Unknown Targets
Less is More: Reducing Complexity in Vision-Language-Action Systems
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
Bayesian Uncertainty Attribution-Guided Fine-Tuning for Open-Set Action Recognition
PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment
Stylized Video Generation via Decoupled Data Synthesis and Gated Style Token Injection
Director: Instance-aware Gaussian Splatting for Dynamic Scene Modeling and Understanding
FUSE: A Flow-based Mapping Between Shapes
Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation
DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering
Comprehensive language–image pre-training for 3D medical image understanding
EGM: Efficient Visual Grounding Language Models
DARL: Efficient Document-to-Markup Generation via Look-Ahead Diffusion Trajectory Sampling
Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination
Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation
Beyond the Embedding Bottleneck: Adaptive Retrieval-Augmented 3D CT Report Generation
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
Gradient sparsity regularization for training unlearning-compatible models
Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics
GKDT: General Keypoint Detection Transformer
TurboMPLE: Joint Infrared Turbulence Mitigation and Physical Fields Estimation via Mutual Progressive Layered Extraction
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
R3RECON: Radiance-Field-Free Active Reconstruction via Renderability
InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics
Distribution Matching Distillation Meets Reinforcement Learning
FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables
From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
The Sterkfontein Caves Dataset: A Novel View Rendering Challenge from the Cradle of Humankind
ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control
Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration
On Locality and Length-Generalization in Visual Reasoning
STVFocus: Query-guided Spatio-Temporal Visual Focusing for Video LLMs
Puppet-CNN: Continuous Parameter Dynamics for Input-Adaptive Convolutional Networks
EgoMAN: Interaction-Structured Reasoning for Egocentric 3D Hand Trajectory Prediction
Objects as Audio-Visual Modal Sound Fields
Towards Spatial Supersensing in the Wild
SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations
DriveVA: Video Action Models are Zero-Shot Drivers
On the real-world generalisability of Optical Flow models
DriveFine: Refining-Augmented Masked Diffusion VLA for Accurate and Robust Driving
KATANA: Knowledge-Aligned Topology-Aware Neural Agents for RL-Driven Vision-Language Model Compression
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
CoCo-IR: Conversational Composed Image Retrieval
CausalDrive: Real-time Causal World Models for Autonomous Driving
Learning Generatable Mutual Distance for Scene-Aware Human Motion Generation
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
If It's Not Efficient, It's Not Usable: Real-Time OOD Detection with Latent De-Biasing and High-Quality Negative Samples
2D Features Are All You Need for 3D Shape Understanding
SONIC: Spectral Optimization of Noise for Inpainting with Consistency
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
Unified and Efficient Point-Line Local Features
Predictive Photometric Uncertainty in Gaussian Splatting for Novel View Synthesis
Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability
Dense Video Understanding with Inter-tokenization Acceleration
RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing
MixCompress: Mixture of Experts for Variable Rate Learned Image Compression
Improved Immiscible Diffusion: Accelerating Diffusion Training by Reducing Miscibility
Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution
WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild
SCORE: SubDistribution-aware Collaborative Knowledge Reinforcing for Cloth-Hybrid Lifelong Person Re-Identification
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
Infinite Gaze Generation for Videos with Autoregressive Diffusion
Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models
ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
OCTA-SOT: Online Cross-Modal Trajectory Adjustment for RGBT Anti-UAV Single Object Tracking under Spatio-Temporal Misalignment
CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation
ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement
Face Anything: 4D Face Reconstruction from Any Image Sequence
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation
DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
WildProp: Visual Estimation of Wildlife Body Proportions at Scale
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
CMDer: Controllable Mode Decomposition-Based Single Motion Synthesis with Diffusion
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Compositional Non-Face Re-Identification Pressure under Cumulative Vision Releases
Grounding World Simulation Models in a Real-World Metropolis
Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding
Incentive Noise and Structural Prior Infusion for Multi-Modal Object Re-Identification
BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular priors
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
Low-Rank Ternary Adaptation for Fine-Tuning Transformers
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
MotionAtlas: A High-Quality Dataset and Benchmark for Dense Motion Captioning
Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
iMED: A Multi-Endoscope Dataset for Surgical 3D Perception
AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
MotionChain: Fine-Grained Video Motion Understanding via Structured Decomposition
Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation
Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time
DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents
A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP
Event Stream-based Sign Language Translation: A High-Definition Benchmark Dataset and A Novel Baseline
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
Pose Anything Anywhere: Model-free Object Poses from Arbitrary References
TaskTok: Delving into Task Tokens for Task-driven Image Restoration
K-Mask: Kinematic-Aware Masked Modeling for Controllable Text-to-Motion Synthesis
FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
XSurfer: Reconstructing surface meshes of cerebral and cerebellar cortex from diverse MRI data using untrained neural networks
Early Estimation of Language to Latent Alignment in Diffusion Models
Tri-Efficient Transfer Learning for Point Cloud Videos
Parametric SDF for Dynamic Surface Reconstruction
Control-DINO: Feature Space Conditioning for Controllable Video Diffusion
CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion
Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation
Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses
Detect by Track: Making Detector-Free Matcher Trackable
Unsupervised Point Cloud Registration via Training-Time Semantic Guidance
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
SEMIR: Topology-Preserving Graph Minors for Thin-Structure Segmentation
Long-term Traffic Simulation via Structured Autoregressive Modeling
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
DASAM3D: A Unified Foundation Model for Enhanced 3D Scene Reconstruction and Segmentation
One Demonstration Is Enough for Real-World Robotic Reinforcement Learning
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
MMAgent-R2: Learning to Rerank and Reject for Agentic mRAG
Plug-and-Play Attention Linearization for Pretrained Transformers
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
ReliefSAM: A Geometry-Augmented Multi-Prior Adapter for Bas-Relief Segmentation
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
STEP: Spatial Thinking and Egocentric Pointing for Embodied Instruction Following
FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
TriNLOS: Triplane Representations for Neural Non-Line-of-Sight Imaging
Kirin: Animal Motion Generation from In-the-Wild Video
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
Harnessing SSL for Segmentation in 3D Microscopy with Noisy Labels and Hard Patches
XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity
TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging
When 3D Gaussian Splatting Recovers Real Surfaces
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution
SLER-IR: Spherical Layer-wise Expert Routing for All-in-One Image Restoration
Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
ProtoMappingNet: Interpretable Hierarchical Prototypes through Relational Prototype Mappings
HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations
SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction
Towards Robustness against Typographic Attack with Training-free Concept Localization
SAFE-EQA: Semantic-Aware Efficient Exploration for Embodied Question Answering
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval
Uncertainty-Weighted Fusion of Image and Synthetic Event for Video Anomaly Detection
X2SAM: Any Segmentation in Images and Videos
Spatiotemporal Flux Probing for Single-Photon Videography
BIP: Bi-level Information Transfer and Completion Prompting for Visual Recognition with Missing Modalities
Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
Spectral Gradient Orthogonalization Improves Differentially Private Training at Scale
Test-time Counterfactual Calibration for Hallucination-Resistant Temporal Grounding
Adapting MLLMs for Nuanced Video Retrieval
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Single-Query Person-Centric Bimanual Hand-Object Interaction Detection
Geodesic Flow Matching on a Riemannian Degradation Manifold for Blind Image Restoration
Interact3D: Compositional 3D Generation of Interactive Objects
Bridge-UniPS: Bridging Calibrated Photometric Stereo toward Universal Photometric Stereo
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation
DRPO: Disentangling Demographic Bias from Rewards for Fair Diffusion Alignment
CS-TTA: Preserving Concept Sensitivity in Test-Time Adaptation
Unsupervised Source-Free Ranking of Biomedical Segmentation Models Under Distribution Shift
SiPhy: Single-Image Physical Property Reasoning
LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
Solving Diffusion Inverse Problems with Restart Posterior Sampling
DeMuS: Learning Decoupled Matching and Scoring for Batch Zero-Shot Industrial Anomaly Detection
Physically Grounded Dual-Opacity Gaussian Splatting for Joint RGB-TIR Reconstruction
LangLoc: “Tell Me What You See”
SupIR-GS: Thermal Infrared Super-Resolution Novel View Synthesis with Imaging-Calibrated 3D Gaussian Splatting
Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion
AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation
Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning
Manifold-Aware Spectral Compaction: A Graph Signal Processing Perspective on Online Gaussian Reduction for 3DGS SLAM
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Implicit Neural Representation Facilitates Unified Universal Vision Encoding
StereoEdit: A Diffusion-Based Framework for Stereo-Consistent Image Editing
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
Conversational Human Audio-visual Talking Dialogue Generation
Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
Rethinking Adversary in Semantic Segmentation: An Out-of-Distribution Perspective
Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation
Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models
Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
Minute4D: Training High-Fidelity 4D Gaussian Splatting in One Minute
DPGS: A Diffusion-Prior Guided Framework for Large-Scale 3D Gaussian Splatting Reconstruction
Improving Knowledge Distillation Under Unknown Covariate Shift Through Confidence-Guided Data Augmentation
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach
Beyond Linear Shortcuts: Rectifying Diffusion Preference Optimization with Intrinsic Generative Geometry
On the Faithfulness of Post-Hoc Concept Bottleneck Models
WaterGen: Decoupling Scene and Medium in Underwater Image Generation
Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models
OmniPoser: Flexible Human Motion Recovery in the Wild with Masked Flow Matching
Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space
GradingBench: Evaluating End-to-End Compositional Reasoning of MLLMs for Automated Exam Grading
EgoTraj: Real-World Egocentric Human Trajectory
PriorMaskMap: Robust Online Vectorized Map Construction with Biased Priors
SGC-Lane: Monocular 3D Lane Detection with Standard-Definition Map Guidance and Lane Completion
Attention-based Vision-Language Memory for Spatial Reasoning
LUCE: Constrained Curve-Domain Guidance for Training-Free Low-Light Enhancement with Hue-Preserving Decoupling
TopoGS: Planar Reconstruction via Topology-Aware 3D Gaussian Splatting
Geometry-Preserving Image Generation for 6D Object Pose Estimation
ARGENT: Adaptive Hierarchical Image-Text Representations
OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration
StAR: Segment Anything Reasoner
Ex-Sim(3)-Reg: 2D-3D Correspondence Pruning via Extended Sim(3) Registration
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
Recurrent Cross-View Object Geo-Localization
FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility
Mind2Cloud: EEG-to-Point Cloud Generation with Two-Granularity Diffusion Decoding
TDSR-VLA: Transition-aware Denoising Sequence Representations for Vision-Language-Action
Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching
AlphaRad: Grounded Zero-Shot Classification in Chest Radiology via α-Corrected Binary Cross Entropy and Factorized Latent Supervision
OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning
Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
World Reconstruction From Inconsistent Views
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers
Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation
EgoEverything: A Benchmark for Human Behavior–Inspired Long-Context Egocentric Video Understanding in AR Environment
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
A Simple Baseline with Placement Prior for Point-Supervised Oriented Object Detection
∂DIBR: Differentiable Depth Image-based Rendering for Fast Novel View Synthesis
DIGS: Differentiable, Incremental, Global, Scalable Pruning for Language Models
DiffuPrompt: Adapting Video Foundation Models to 3D Medical Volumes via Latent Trajectory Priors
Low-Level Dataset Distillation for Medical Image Enhancement
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
Noise-Robust Face Recognition via Non-target Similarity Distribution Guided Sample Selection
GeoDetect: Geometric Adversarial Detection for VLPs
Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction
LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
PMGC-SimVP: Parametric Multi-scale Gated Convolution for Global Ionospheric TEC Prediction
Where and What: Long-Term Object Tracking in Egocentric Videos
Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition
Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
Let ViT Speak: Generative Language-Image Pre-training
Prototype-Conditioned Imagination for Compositional Zero-Shot Learning
MambaRaw: Selective State Space Modeling for Efficient 4K RAW Image Reconstruction
SwiftWA: An Efficient Action-Centered World-Action Model
VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On
Robust onion: Peeling Open Vocab Object Detectors Under Noise
Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset
UniCSG: Unified High-Fidelity content-constrained style-driven generation via Staged Semantic and Frequency Disentanglement
RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
Setting the Stage: Text-Driven Scene-Consistent Image Generation
Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding
ESC: Emotional Self-Correction for Reliable Vision-Language Models
Kilometer-Vision: A New Frontier for Large-Scale Spatial Awareness in VLMs
ReSWD: ReSTIR‘d, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction.
From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting
Enhanced Neural Video Representation Compression with High Scalability
FUSE: Filter-Free Unified Spatiotemporal Estimation of SpO2 via Wave-Transport Modeling
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
Safe Responses Matter: Output-Aware Safety Guardrail Mitigate Over-Refusal in MLLMs
Exploiting Local Flatness for Efficient Out-of-Distribution Detection
REDistill: Robust Estimator Distillation for Balancing Robustness and Efficiency
Event-based Gaze Control Systems for Real-time Spin Estimation in Professional Ball Games
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings
Rethink Backdoor Robustness in Vision Transformers
FlowInOne: Unifying Multimodal Generation as Image-in, Image-out Flow Matching
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
Frames2Residual: Spatiotemporal Decoupling for Self-Supervised Video Denoising
Do Multimodal LLMs Understand Intraoral Dental Data? Dataset, Platform, and Baselines
SyncFix: Multi-View Consistent Diffusion Refinement of 3D Reconstructions
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
EgoExoMoCap: Distributed Human Motion Capture via Ego- and Exocentric Body Tracking from Head-Mounted Devices
SDUM: A Scalable Deep Unrolled Model for Universal Cardiac MRI Reconstruction
From Blobs to Spokes: High-Fidelity Surface Reconstruction via Oriented Gaussians
SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization
A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention
SWSL: Semantic-aware Weakly Supervised Learning for 3D Motion Generation using 2D Motion Data
Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes
LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering
Multi-Head Normalization for Wide Vision Transformers
Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion
VIGS-SLAM: Visual Inertial Gaussian Splatting SLAM
FMS2: Unified Flow Matching for Segmentation and Synthesis of Thin Structures
Local Spacing-Aware Hungarian Matching for Stable Point-Supervised Crowd Counting
LDC-MTL: Balancing Multi-Task Learning through Scalable Loss Discrepancy Control
Task-Agnostic Incremental Vision-Language Object Detection via Prompt Augmentation and Distribution-Aware Fusion
OCTOPUS: Multi‑Agentic Universal Compositional Visual Retrieval
VLA Knows Its Limits
Masked Depth Modeling for Spatial Perception
D-VLAM: Differential Vision and Language Mixing for Rehearsal Free Continual Learning
SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising
Beyond Common Sense: Grounding Logical Anomaly Detection in Inspection Criteria
Virtual Category-Guided Continual Generalized Category Discovery
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
Unpaired Geometry-Guided Sim2Real Translation for Autonomous Driving
QST-SAM: Leveraging Cross-modal Instructions for Few-shot Referring Video Object Segmentation
FedNASP: Federated Vision-Language Navigation with Adaptive Step-wise Personalization
ICLAgent: Integrated Circuit Footprint Geometry Labeling via LMM-empowered Multi-Agent Framework
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Learning Video Dynamics with Predictive Differentiable Rendering
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Histopathology Multi-modal Embedding for Pathology Composed Retrieval
Implicit Neural Representation for Spherical Harmonics Reconstruction of Motion-Corrupted Fetal Diffusion MRI
Unsupervised Semantic Segmentation Facilitates Model Understanding
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction
ReflectCAP: Detailed Image Captioning with Reflective Memory
Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks
Multi-view Multi-vehicle Driving Dataset for Novel View Synthesis
Obliviate: Erasing Concepts from Autoregressive Image Generation Models
ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID
Inference-time Motion Calibration for Video Generation
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
Recognizing Co-Speech Gestures in-the-Wild
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
Direct Preference Optimization for Perceptual Alignment via Vision-Language Consistency
Preventing Expert Collapse in MoE-dVLMs via Modality-Wise Norm Alignment
Explicit Semantic–Spatial Alignment for Open-Vocabulary Object Detection
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
Context-Interactive Reasoning for Group Activity Detection
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
SDSA: Shallow-Deep Squeezing Adapter for Vision-Language Models
FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
CtrlCoMo: Controllable Co-Speech Motion Generation with Gesture–Action Disentanglement
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
Revisiting Deepfake Detection: BCNet for Robust Generalization Beyond Semantic Dependence
PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification
Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder
Semantic Generative Tuning for Unified Multimodal Models
InstanceControl: Controllable Complex Image Generation without Instance Labeling
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Optimizing Mesh Animation from Video via Shape Flow Guidance
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
Rethinking IRSTD: Single-Point Supervision Guided Encoder-only Framework is Enough for Infrared Small Target Detection
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
Fast Spatial Memory with Scalable Elastic Test-Time Training
ComplexMimic: Human–Scene Interaction Imitation in Complex 3D Environments
CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition
Debiased Textual Prompt Tuning for Enhancing Unknown Class Discovery
Inference-Time Scaling of Diffusion Models via Progressive Pruning Search
QualiTeacher: Quality-Conditioned Pseudo-Labeling for Real-World Image Restoration
DiffVP:Differential Visual Semantic Prompting for LLM-Based CT Report Generation
Prefill-Time Interventions against Adversarial Attacks on Large Vision-Language Models
PriSM: Parsing and Style-Mixed Consistency for Unsupervised Domain Adaptation in Facial Landmark Detection
DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion
PercepTax: Benchmarking Cross-Property Reasoning in Vision-Language Models
Geometric Probing for Isotropic Optimization Manifold in Sparse-View 3D Gaussian Splatting
PPTArena: A Benchmark for PowerPoint Editing
MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction
DeRA: Decoupled Representation Alignment for Video Tokenization
LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos
Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability
MindBlock: Probing Spatial Assembly and Structure in Unified Multimodal Models
BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure
FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs
LiteGS: a high-performance framework to train 3dgs in subminutes via system and algorithm codesign
AffoGato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets
VoxAnchor: Explicit Voxel-Semantic Grounding for Spatial Understanding in Videos
VNC: A Scale-Space Foundation for Learnable 3D Surface Evolution
MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization
RhymeFlow: Training Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling
GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
Comprehensive Robustness Analysis of LiDAR-based 3D Object Detection in Autonomous Driving
CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection
Going Deep: Deep Visual Prompting with LoTeP
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
CURE: Contextual Debiasing and Unbiased Refinement for Training-Free Open-Vocabulary Semantic Segmentation
EvDiff: High Quality Video with an Event Camera
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents
MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer
Straight-Path Flow Matching for Incomplete Multi-View Clustering
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
GTR: Guide-Then-Refine Token Compression for Training-Free Acceleration of Video-LLMs
Driving is a Game: Combining Planning and Prediction with Bayesian Iterative Best Response
Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models
340 FPS Reflection-free Video from Spikes Modulated by a Rapidly Rotating Polarizer
Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation
UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation
Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
FusionTrack: Collaborative Multi-Object Tracking with Arbitrary Multi-UAVs
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
Clue Matters: Empower Video Reasoning with Brain-Inspired Latent Clue Learning
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
FD²: A Dedicated Framework for Fine-Grained Dataset Distillation
Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning
MobileSAM2: Lightweight Segment Anything in Images and Videos via Hypergraphical Knowledge Distillation
TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs
Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training
Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
Triangular Consistency as a Universal Constraint for Learning Optical Flow
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction
OccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy Space
D²R²OSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution
ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation
Nexus-Vid: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting
Topology-Weighted Effective Rank: A Zero-Cost Proxy for Training Dynamics Stability in Neural Architecture Search
RBE-Flow:Recurrent Bayesian Estimation on Feature Manifolds for Cross-Modal Registration
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
Lifting Ego World Models for Planning and Control
GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures
ReInGS: Re-Initializing 3D Gaussians against Sparsity Discrepancy in Few-Shot Novel View Synthesis
Diffusion Model as a Generalized Segmentation Learner
WebEyeTrack: Scalable Eye-Tracking for the Browser via On-Device Few-Shot Personalization
REAL-OW: Rehearsal-free Open World Object Detection with Low-Rank Adaptation and Dual-Stage Objectness Modeling
In-context Region-based Drag: Drag Any Region to Any Shape
Unmasking-Time Visual Calibration for Hallucination Mitigation in Multimodal Discrete Diffusion Language Models
GridVQA-X: A Diagnostic Framework for Evaluating Multimodal Explainability Methods
MoE-KD: Your Teacher Model is Worth Mixture-of-Experts for Knowledge Distillation
MotionDreamer: Universal Skeletal Motion Generation for 3D Rigged Shapes
ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
SSBP: Stage-Specialized Block Pruning for Video Diffusion Models
Stand Up and Move: Benchmarking Interactive Spatial Intelligence in WalkerBench
Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective
EAGS: Error-Aware Gaussian Splatting with Dual-Confidence-Guided Modeling for Uncalibrated Driving Scenes
RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory
SynHMR: Synergistic Joint-Mesh Modeling for LiDAR-based Human Mesh Reconstruction
Boba: Batched Simulation for Physics-Based Gaussian Digital Twins
Weather-Conditioned Depth Anything
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
LEO-Fuse: A Modality- and Task-Agnostic Universal Framework for Multimodal Human Sensing
Twin-DAgger: Synergizing Digital Twins and Human Corrections for Efficient Robot Manipulation
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions
VPA-WM: Vision-Priors-Aligned World Models for Robust Visual Reinforcement Learning
Trajectory-aware Cross-view Geo-Localization with Sequential Observations
FlatLands: Generative Floormap Completion From a Single Egocentric View
Don’t Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
Agentic Collaborative Cognition for Zero-Shot 3D Understanding
RGB-Pointmap Pretraining for Unified 3D Scene Understanding
Multi-THuMBS: Multi-person Tracking of 3D Human Meshes Beyond Video Shots
HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models
IMMoE: Incomplete Multi-View Anomaly Detection via Mixture of View Experts Fusion
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
DriftScope: Measuring The Hidden Effects of Diffusion Model Fine-Tuning
LlamaSeg: Image Segmentation via Autoregressive Mask Generation
Don’t Teach Instability, Teach Robustness: Selective Sensitivity Gating for Adversarial Robust Distillation
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
Learning Accurate Segmentation Purely from Self-Supervision
MG2-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
UniStitch: Unifying Semantic and Geometric Features for Image Stitching
3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
Articulated Object Reconstruction from Rest-State Observation
SPEAR: A Simulator for Photorealistic Embodied AI Research
Bounding-Box Trajectories Matter for Video Anomaly Detection
Defect-aware Hybrid Prompt Optimization for Zero-Shot Multi-type Anomaly Detection and Segmentation
MED-LCDS: Multi-Expert-Domain CLIP Classification via Logit Calibration
VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human–Scene Interactions
FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Arbitrary Images
RawGen: Learning Camera Raw Image Generation
Generalized Biomedicine Discovery
Demystifing Video Reasoning
Calibrated Harmonic Overlaid Implicit Neural Representations for Multi-Dimensional Data
CustomX: Unified Character, Action, and Scene Customization in Video World Models
Vero: Open Reinforcement Learning Recipes for Visual Reasoning
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
Towards Geometry-Grounded Dense Semantic Matching with VGGT Priors
PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models
Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Multi-Block-Attention-based Color Constancy
Activation Quantization of Vision Encoders Needs Prefixing Registers
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
Capacity Overflow: A Blind Spot for Backdoor Attacks in Vision MoE
FedDO: Dynamic Client Optimization for Adaptive Federated Learning
A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR
Sparse auto-regressive modeling for scene generation from multi-view images
Swap the Right Identity: Spatio-Temporal Preference Optimization for Identity Swapping
VKSR: Scalable Kernel Surface Reconstruction Using Vecchia's Approximation
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
Mimicking Radiologists: A Coarse-to-Fine Framework with Structural Sparse Tokens for Dual-LLM Computed Tomography Report Generation
Training-free Discriminative Patch Mining for Robust Few-Shot Recognition with CLIP
Learning 1-Bit LiDAR-based Localization with Auxiliary Objective
Decoding Children’s Gait Behavior
AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting
BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Analyzing and Improving Training-Free Fast Sampling of Text-to-Image Diffusion Models
Thinking from the Robot’s View: The CoT-HRC Benchmark for Human Intent Reasoning in Embodied Collaboration
GridFlow: Structured Latent Flow for Seamless City-Scale 3D Point Cloud Generation
Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
PyraE2E: Enhancing End-to-End WSI Analysis via Cross-Scale Super-Resolution
RefracGS: Novel View Synthesis Through Refractive Water Surfaces with 3D Gaussian Ray Tracing
SGMatch: Semantic-Guided Non-Rigid Shape Matching with Flow Regularization
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
Rethinking Real-World MRI Denoising: Learning from Physical Noise
LEAP-VLA: Latent-Enhanced Action Prototyping via Continuous Residual Latent Spaces for Vision-Language-Action Models
BRepFacetGen: Reverse Engineering B-Reps By Generative Face Segmentation
Controlling Embedding Spaces with Text-Conditioned Transformations
Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
DiNBV-Grasp: Real-Time Distance-Aware Two-Stage Next-Best-View for Robotic Grasping
Boosting Correspondence Learning with Structure-Aware Estimator
StreamSpatial: A Benchmark and Framework for Streaming 3D Visual-Spatial Reasoning
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
NavWM: A Unified Navigation World Model for Foresight-Driven Planning
GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery
SLAIR: Structured Latent Flow Matching for All-in-One Image Restoration
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
General Incomplete Multimodal Learning via Dynamic Quality Perception
Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification
PASDiff: Physics-Aware Semantic Guidance for Joint Real-world Low-Light Face Enhancement and Restoration
HIDA: A Human-Intuition-Guided Depth-Aware Framework for Zero-Shot Amodal Segmentation
Fast Dynamic Prototypes for Unsupervised Anomaly Detection and Localization
SP-TransientBench: A Real-Captured Single Photon Perception Benchmark
3DGS3: Joint Super Sampling and Frame Interpolation for Real-Time Large-Scale 3DGS Rendering
BWAFDA: Block-wise Weighted Attention Fusion with Detail-aware for No-Reference Image Quality Assessment
SCALE: Semantic-Calibrated Guidance Enhancement for Prompt-Faithful Diffusion
Causal Yet Future-Aware: Dual-Path Temporal Modeling for Online Action Segmentation
CameraAnything: Refilming Videos with Arbitrary Camera Control
Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding
Attention is Case-Sensitive
Wavelet-Guided Semantic Signal Compensation for Inversion-Free Image Editing
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Small Vision-Language Models are Smart Compressors for Long Video Understanding
PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design
Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
ATP-Bench: Towards the Agentic Tool Planning for MLLM Interleaved Generation
UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
HVA-Fusion:Hierarchical Velocity-Aware 4D Radar-LiDAR Fusion for Robust 3D Object Detection
From Predictions to Embeddings: Dual Knowledge Distillation for Instance-Dependent Partial Label Learning
One-Shot Feed-Forward 360° Animatable Avatar via Inpainted UV-Space Gaussian Modeling
AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation
Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting
Purify then Guide: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing
SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models
EGVLR: Evidence-Grounded Vision–Language Reinforcement for Anomaly Reasoning
Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation
3D-Aware VLMs with Implicit and Explicit Geometries
λSplit: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion
Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding
TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger
PointSplat: Compact Gaussian Splatting via Human-Centric Prediction
Decoupling Complexity from Scale in Latent Diffusion Model
Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
SCDL: Synergistic Confidence-Dispersion Learning for Semi-Supervised Video Polyp Segmentation
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
Stabilizing Real-World Visual Active Tracking with Action-Smooth Test-Time Adaptation
ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation
Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning
SAND: Stage-Aware Noise Decomposition for Training-Free Diffusion Guidance
Fast Sam 3D Body: Accelerating SAM 3D Body for Real-Time Full-Body Human Mesh Recovery
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
Cross-token Guidance Transformer for Weakly Supervised Object Localization
OrthoTailor: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation
From smooth to sharp: Frequency-Decoupled Latent Optimization for Realistic Image Generation
TimeWalker: Personalized Neural Space for Lifelong Head Avatars
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression
Frozen CLIP Priors for Robust Self-Supervised Poisson Inverse Problems
XDen-1K: A Density Field Dataset of Real-World Objects
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications
An Inverse-Adversarial and Difficulty-Adaptive Robust Vision-Language Model
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute
Tricam-rPPG: A Multimodal Multispectral Dataset for remote Photoplethysmography
RCEdit-500K: Reference Completion for Image-Conditioned Image Editing
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal
Denoising the Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging
Noise-Robust Facial Expression Recognition via Mamba-driven Neighbor Weight Refinement
Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
PaD-GS: Leveraging Distortion Map for Panoramic Gaussian Splatting
Beyond the Boundary: RL-Driven Solution Space Exploration for Blind Face Restoration
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
Every Dog Has Its Day, Probably: A Balanced Synthetic Benchmark and Probabilistic Modeling for 3D Dog Pose Estimation
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
NanoVSR: Towards Real-Time Video Super-Resolution on Edge Devices
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
Accurate Zero-shot Quantization via Hierarchical Teacher-Assistant Distillation
MapDreamer: Aerial Imagery Conditioned Latent Diffusion For Lane Level Map Generation
CLDefocus: Physically Grounded Compound-Lens Defocus Blur Synthesis
Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies
FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation
OmniNWM: Unifying the State-Action-Reward Triad for Closed-Loop Panoramic Driving Navigation World Models
G2FM: A Geodesic Flow Matching Framework with Geometric Prior for Category-Level 9-DoF Pose Estimation
CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction
TriMotion: Modality-Agnostic Camera Control for Video Generation
Unified Video Dense Prediction from Disjoint Data
NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation
SHINE-PPG: Non-Lambertian Intrinsic Decomposition for Illumination-Robust rPPG
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
Score-Based Matching with Target Guidance for Cryo-EM Denoising
MATCH: Flow Matching for Multi-View Anomaly Detection
AIMold: An Autonomous AI-based Pipeline for Complex Mold Design
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
ProSR: Semantic-Prototype-Guided Discrete Modeling for Physically Consistent SAR Super-Resolution
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
Bayesian Self-Attention with Local Pixel Correlations for Lightweight Denoising Transformers
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
NeLU3D: Neural Inverse Structured Light without Modeling the Projector
X-Stream: Benchmarking MLLMs as Multiplexers for Multi-Stream Understanding
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
GuideMe: Benchmarking Multi-Domain Task Guidance and Intervention in Streaming Video
MCVL: Multi-Space Cross-View Learning for Aerial-Ground Person Re-Identification
A Comprehensive Analysis about Unsupervised Outlier Detection for Images
Generative Refinement Network for Visual Synthesis
SkySplat-OV: Generalizable Language Gaussian Splatting for Open-Vocabulary Scene Understanding from Sparse Satellite Views
EvoVLA: Self-Evolving Vision-Language-Action Model
ChronusOmni: Improving Time Awareness of Omni-Modal Large Language Models
LumiDepth: Stable Monocular Depth in Multi-Illumination Scenes
Neuromorphic X-ray Computed Tomography
Why Feature Magnitude Deceives OOD Detectors: An Angular Separation Perspective
Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors
SkyLume: A Large-Scale Multi-Illumination Aerial Benchmark for Urban Scene Reconstruction and Beyond
Geometry-Aware Style Transfer in 3D Gaussian Splatting
Learning to Mask: Cross-Modal Noise Modulation for Hallucination Mitigation in Multi-modal Large Language Models
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation
Orthogonal Knowledge Refreshing for Domain-Incremental Object Detection
SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
UniTranslator: A Unified Multi-modal framework for End-to-end In-Image Machine Translation
Gaussian Belief Propagation Network for Depth Completion
StrucTab: A Structured Optimization Framework for Table Parsing
Beyond Script Family Boundaries: Towards Unified Open-Set Scene Text Recognition
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
Mapping Dark-Matter Clusters via Physics-Guided Diffusion Models
Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration
Pix2NPHM: Learning to Regress NPHM Reconstructions From a Single Image
VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching
Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation
Attention-Logit Steering to Compositional Generalization for Continual VQA
WildWorld: A Large-Scale Dataset for Action-Conditioned World Modeling with Explicit State Annotations
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
VideoSfM: Exploiting Temporal Structure for Video-Based Structure-from-Motion
BrepLLM: Enabling Large Language Models to Understand Boundary Representations
SGQA: Semantic-Geometric Quality Alignment for Training-Free Few-Shot Instance Segmentation
Schroedinger’s Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics
Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization
Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection
GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure
Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization
When the Teacher Has More Bits: Self-Teacher Latent Distillation for Learned Image Compression
CaRe: Critical Parameter Rectification for Efficient Visual Modeling
Beyond Artifacts: Real-Centric Envelope Modeling for Reliable AI-Generated Image Detection
Temporally Aware Densification for Dynamic 3D Gaussian Splatting
SegFly: A 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale
RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement
Gaussians on Fire: High-Frequency Reconstruction of Flames
Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models
PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
Fabric Image Demoiréing Benchmark from Synthesis to Restoration
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
RUTaL: Residual Upcycling with Task Ladder for Efficient Multi-Task Learning
SPDA: Efficient Online Test-Time Adaptation for Promptable Medical Segmentation
Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training
Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation
One Trap to Block Them All: Defending Encoder Stealing via Isotropic Uniformity
TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts
Interference-Aware Continual Vision–Language Learning via Instance-Level Expert Routing
Identifiable Gated Residual Personalization for Federated Parameter-Efficient Fine-Tuning
DiscoVL: Unveiling Disentangled Cross-Modal Representation Learning via Orthogonal Adversarial Regularization for Vision-Language Models
Curvature-Guided Mixing for MLLM Adaptation
PACO: Stabilizing Vision Embeddings along Local Paths for Robust Vision-Language Models
Aggregating Cross-Domain Knowledge via Learnable Tokens for Multi-Teacher Distillation
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
R2M: Real-Aware Residual Model Merging for Robust and Generalizable Deepfake Detection
BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming
Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
SiGMA: Sign-Guided Merging and Adaptation framework for Multimodal Continual Instruction Tuning
COVERT: Privacy-Preserving Covariant Obfuscation for VLMaaS via Exact Reparameterization and Tailored Tuning
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
VD-LoRA: Adaptive Reuse of Low-Rank Directions for Continual Learning
Probe, Anchor, and Amend: Active Test-Time Adaptation of Vision-Language Models
Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent
Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification
COLA: Continual Orthogonal Low-Rank Adaptation for Class-Incremental Learning
Stealthy Multi-task Adversarial Attacks
SAMPLe: A Sharpness Aware Minimization based Optimizer for Prompt Learning in Vision-Language Models
GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models
Locality-Aware Continual Unlearning for Diffusion Models
Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift
Indelible Backdoors: On the Limits of Post-Training Defenses
Condensing Large-Scale Datasets Directly with Minimal Information Loss
Isotropic Embedding Perturbations for Robust Vision Language Encoders
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
Stabilizing Ultra-Low-Bit Quantization of Multimodal LLMs via Global Bit Allocation
Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation
Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
Rethinking Training and Inference for Trajectory Forecasting: Linking Winner-Take-All back to GMMs
SPARC: Scalable Path-Specific Counterfactual Fairness via Causal Conditional Independence
Breaking Rigidity in Adversarial Patch Attacks
Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning
AnyGround3D: Towards Grounding Any 3D Object in the Wild via 2D-to-3D Lifting
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound
Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
Taming LLMs for Codematic Indoor Scene Generation
TORA: Topological Representation Alignment for 3D Shape Assembly
LiFlow: Flow Matching for 3D LiDAR Scene Completion
Event-LiDAR: 3D Eventification for Efficient Point Cloud Processing
CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose
SuperFlex: Deformable Superquadrics for Point Cloud Decomposition
GraphCPD: Coherent Point Drift for Point Cloud Registration via Graph Signal Processing
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
DepWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
PGCR: Pose–Geometry Coupled Reasoning for Image-to-Point Cloud Registration
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
Geometric Context Transformer for Streaming 3D Reconstruction
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
Vulnerability of Privacy-Preserving Visual Localization against Diffusion-based Attacks
HHA: Hierarchical Hyperbolic Constraints for Imperceptible Point Cloud Attacks
SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
Learning Physics-based Forward Model Corrections in Unrolled Networks for Diffuser-based Imaging
RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes
Ranked Activation Shift for Post-hoc Out-of-Distribution Detection
WiFlow: Estimating Optical Flow using WiFi Channel State Information
Stabilizing Deep Reconstruction Operators with Contractive Anchoring
Provable and Robust Wavefront Sensing via Self-Reference Interferometry
Estimating Individual Tree Height and Species from UAV Imagery
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization
Compact Low-Cost Hyperspectral Imaging via Angular-to-Spectral Diversity Conversion
SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification
High-speed Imaging through Turbulence with Event-based Light Fields
Any to Full: Prompting Depth Anything for Depth Completion in One Stage
Filterless Snapshot Hyperspectral Imaging using Guided Patch Diffusion
Modeling and Compensating Phase Error in High-speed 3D Reconstruction
Learning to Suppress SPAD-based LiDAR Flare
Broadband Wide Field of View Imaging with Computational Mirrors
Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding
ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Under One Sun: Multi-Object Generative Perception of Materials and Illumination
TPCNet: A Low-Light Image Enhancement Network Inspired by Triple Physical Constraints
AiSCREAM: Absolute Target Localization with Language-Conditioned Cross-View Alignment for Autonomous Vehicles
Video Generative Models as Geometry Learner
Parallax Portrait Matting
SynLF: Zero-Shot Metric Depth from Light Field Cameras via Physics-Grounded Synthesis
OneHSI: A Unified Hyperspectral Foundation Model with Physical Consistency
DP-BOA: Dirichlet-Process Birth-or-Assign for On-the-Fly Category Discovery
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Leveraging Phase Information to Boost Unrolled Network Learning for Image Deblurring
Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation
Consistent Monocular Depth Estimation with Contact Region Boundary-Aware Refinement
EventSpecPS: Photometric Stereo with Multispectral Reflectance Using an Event Camera
Color Pass-Through via Camera-Display Coupling
A Mechanism-Driven Theory of Phase Transitions in Active Learning
Stable and Scalable Bundle Adjustment of Holistic 3D Structures
FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation
FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness
Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning
A second-order theory of texture for depth from focus
Depth-guided Multi-view Exposure Bracketing for HDR Robot Vision
The Devil Is in the Dark Pixels: Toward Brightness Bias-Robust Denoising
SOMA: From Surface Observations to Muscle Anatomy
CSS-BA: Gate Guided Column Space Search for Bundle Adjustment
Asymmetric Anchoring: Opening the Black Box of MLLMs for Forgery Detection
Geometry-Aware Visual Representation for Remaining Useful Life Prediction
Semantic Line Diffusion: Character-Consistent Line Art from text-annotated Storyboards
Dynamic Inverse Rendering for Enhanced Material-Lighting Decomposition
Video Can Teach PAN-Sharpening: PSF-Aware Cross-Domain Supervision
OmniPoint: Universal Monocular Metric Pointcloud from Any Camera
Don’t Mask Out the Background! Natural-Light Photometric Stereo via Illumination Reconstruction
mmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC Synthesis
ESNE: Efficient Surface Normal Estimation for LiDAR Point Clouds with Sequential Modeling and Variability Guidance
SFD-Net: Sharp Feature Detection Network Based on Local Geometric Features
Pixel-wise Planarity for High-Precision Monocular Plane Segmentation
Boosting 6D Object Pose Estimation via Monocular Depth Cues
Synthetic Sub-Aperture Phase Augmentation for Demosaicing 2×2 Shared Microlens Sensors
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
PixVOD: Pixel-Distributed Direct Visual Odometry and Depth Estimation
Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space
Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction
The 3D Mirage: Probing and Taming 3D Hallucinations
Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting
OmniRen: Neural Rendering wih Heterogeneous Scene Primitives
OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization
Real-Time LiDAR Gaussian Splatting SLAM via Geometry-Aware Covariance Coupling
VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment
Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation
RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction
Incremental Online Scene Reconstruction by 3D Gaussian Triangulation
Matryoshka Gaussian Splatting
Identity-Preserving Human Reconstruction from a Single Image via 3D Token Inference
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
Decoupled Illumination Priors for Spatially Controllable Multi-View Indoor Scene Relighting
OmniX: Any-view and Any-time 4D reconstruction via Feed-forward Trajectory Fields
NanoGS: Training-Free and Lightweight Gaussian Splat Simplification
Wat3R: Underwater 3D Geometry Learning without Underwater Annotations
StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors
Holo-Captioning: A Comprehensive Textual View of 3D Scenes
Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space
StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views
IRIS: Intersection-aware Ray-based Implicit Editable Scenes
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
FaCT-GS: Fast and Scalable CT Reconstruction with Gaussian Splatting
GAP-Track: Bridging the Resolution Gap for Cross-Resolution RGBT Tracking
GaINeR: Geometry-Aware Implicit Neural Representation for Image Editing
Free-Range Gaussians: Non-Grid-Aligned Generative 3D Gaussian Reconstruction
Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth
PanoLess: Environment Reconstruction from Partial Reflective Views
CARA: Collision-Aware Resolution Adaptation for Multiresolution Hash Encoding Based Image Fitting
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
FillGS: Filling Observation Gaps in 4D Gaussian Splatting via Viewpoint-Time Selection and Generative Refinement
Confidence-Based Mesh Extraction from 3D Gaussians
InstantHDR: Single-forward Gaussian Splatting for High Dynamic Range 3D Reconstruction
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision
Do Flat Minima Improve Sparse Novel View Synthesis?
ViewSplat: View-Adaptive Dynamic Gaussian Splatting for Feed-Forward Synthesis
Fast and Compact 3D Gaussian Splatting with Polarized Opacity Prior
DINO-SLAM: DINO-Informed RGB-D SLAM for Neural Implicit and Explicit Representations
COSY: Compositional 3DGS Synthesis for Disentangled Human Head Editing
DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling
Active View Selection with Perturbed Gaussian Ensemble for Tomographic Reconstruction
REFINE: Super-efficient Pruning for 3D Gaussian Splatting via Rendering-Free Primitive Importance
UniGeo: Unifying Geometric Constraints for Camera-Controllable Image Editing via Video Priors
HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits
Nexels: Neurally-Textured Surfels for Real-Time Novel View Synthesis with Sparse Primitives
Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-View Videos
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
FreeGen: Feed-Forward Reconstruction–Generation Co-Training for Free-Viewpoint Driving Scene Synthesis
MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing
AVSplat:Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning
Enhancing Embodied Reasoning and Grounding by Novel View Synthesis
Reflection-aware generative novel view synthesis
Instant Expressive Gaussian Head Avatars at Over 100 FPS
DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting
Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors
PriSplat: Propagating Reliable Multi-view Information for Distractor-Free 3DGS
City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion
AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing
Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering
DR-GS: Physically-Based Deformable and Relightable 2D Gaussians
StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction
Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation
3D Gaussian Splatting Compression with Object Scalability
InstGS: Shared-Template Gaussian Instancing for Object-Redundancy-Free Rendering
CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance
Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning
Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization
ReSplat: Learning Recurrent Gaussian Splatting
AnchorSplat: Fast and Structure Consistent Detail Synthesis for Gaussian Splatting
AdaptiveSplat: Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction
Denoising-GS: Gaussian Splatting with Spatial-aware Denoising
Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering
DINOv3D: 2D-3D Joint Optimization for Unified Spatial Understanding
MVGS: Multi-view Regulated Gaussian Splatting for Novel View Synthesis
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
SharpGS: Sharpness-Preserving 3D Gaussian Splatting with Differentiable Blur-Driven Density Control
GlassGS: Geometry and Concept-Aware 3D Gaussian Splatting for Reflective Enclosures
ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views
SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction
BEV-GS: Feed-forward Gaussian Splatting in Bird’s-Eye-View for Road Reconstruction
CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
Racing in Volume with Flow Ensembles
LiteMatch: Lightweight Zero-Shot Stereo Matching via Cost Volume Stabilization
ReGen3D: Generalizable Unified Representation Learning for 3D Understanding
PointGT: Simultaneous Geometric and Textural Editing for Point-Based Representations
LiDAR-EVS: Enhance Extrapolated View Synthesis for 3D Gaussian Splatting with Pseudo-LiDAR Supervision
UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment
Relaxed Rigidity with Ray-based Grouping for Dynamic Gaussian Splatting
WildSplat: Feedforward Gaussian Splatting from Unposed In-the-Wild Images
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
Geometry-Preserving in 3D Gaussian Splatting for LiDAR-Camera Extrinsic Calibration
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
LVSPM: Long Sequence View Synthesis and Pose Estimation Model
FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction
UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras
DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
PDF-Omni: Poincaré Dual Disk Distortion Field-based Recurrent Update for Omnidirectional Stereo Matching
UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization
Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
KISS-GS: 3D Gaussian Splatting Compression Kept Simple
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
Drop-In Perceptual Optimization for 3D Gaussian Splatting
SkipGS: Post-Densification Backward Skipping for Efficient 3DGS Training
Structure Gaussian Splatting SLAM
F⁴Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling
TRiGS: Temporal Rigid-Body Motion for Scalable 4D Gaussian Splatting
Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving
RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering
ActiveStructure: Plane Scene Graph-Guided Active 3D Gaussian Splatting
3D Gaussian Texture for Real-time Mesoscale Appearance Synthesis and Rendering
Geometry Grounding: Elevating Blind Distortion Correction with 3D Structural Priors
GeoV2V: Geometry-Grounded Video Diffusion Model for Driving Scene Generation
AGE: Agentic Gaussian Editing in 3D Scenarios
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views
GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
Proto-Gaussian: MRI Modality Translation Based on Learnable Structural Prototypes and 2D Gaussian Splatting
TriSplat: Adaptive Triplane for Sparse-View Large-Scale Scene Reconstruction
IndoorSplat: Enhanced Indoor Scene Reconstruction with Structured 2D Gaussian Splatting
MVFusion-GS: Motion-Variance Guided Temporal Attention for High-Quality Dynamic Gaussian Splatting
HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video
PIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond Decoding
Geometry-Propagated Gaussian Splatting for Aerial Sparse Novel View Synthesis
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
Graph-GSReg: Leveraging 3D Scene Graphs for Gaussian Splatting Registration
LEGO: Leveled Language Gaussian Splatting
NoPA: Non-Parametric Online 3D Scene Graph Generation
Triangle Splatting SLAM
Towards Alias-Free 4D Gaussian Representations with Motion-Aware Filtering
Latent Fusion: Decoding Consolidated 3D Geometry from Feed-forward Geometry Transformer Latents
Novel View Synthesis as Video Completion
Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
MLP Splatting: Object-Centric Neural Fields
RaPTGS: Render-Agnostic Post-Training Compression of 3D Gaussian Splatting
Fourier Splatting: Generalized Fourier encoded primitives for scalable radiance fields
CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
CAM3R: Camera-Agnostic Model for 3D Reconstruction
FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors
3D-ReGen: A Unified 3D Geometry Regeneration Framework
CoDePose: Multi-View 3D Human Pose Estimation via Coupled 2D-3D Denoising Diffusion
SMP-UWGS: Coupled Physics-Geometry Optimization for Scalable Multi-Partition Underwater 3D Reconstruction
Diversity-Aware View Partitioning for Scalable VGGT
Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation
StratoSplat: Taming Layered Regularities for Sparse Aerial 3D Gaussian Splatting
Fixed Reality, Diffused Possibility: Disentangling Stochastic and Deterministic Latent for Cluttered Grasping
Steering 3D Generations: Preference Alignment via Direct Reward and Preference Optimization
EGGS: Explicitly Granular 3D Gaussian Splatting via Luma-Aware and Volume-Preserving Attribute Factorization
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation
DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing
X-SG2S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks
Online Segment 3D Gaussians via Launching Virtual Drones
Dynamic-Robust Photometric–Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding
BioMTBee: Biologically Constrained Multi-View Template-Based 3D Reconstruction of Bumblebee
RoomPlanner: Reachability-Aware View Sampling for Text-to-Room 3D Gaussian Splatting
MedGSSR: Generalizable Medical Image Super-Resolution 3D Reconstruction via Hierarchical Feed-forward Gaussian Splatting
SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation
TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid
Same Person, Different Depiction: Counterfactual Evaluation of Vision-Language Models on Individuals with Limb Deficiencies
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
Token-Based Affordance Grounding with Large Vision-Language Models
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
GEM: Generative Supervision Helps Embodied Intelligence
CFM: Language-aligned Concept Foundation Model for Vision
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming
RAU: Reference-based Anatomical Understanding with Vision-Language Models
Trustworthy Image Authentication using Forensic Knowledge Graphs
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
AdaBoosting Text Prompts for Vision-Language Models
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
C3-Bench: A Context-Aware Change Captioning Benchmark
To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
Domain Generalization via Text-Anchored Information Bottleneck
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
See Only When Needed: Context-Aware Attention Intervention for Hallucination-Free LVLMs
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
PhenoLIP: Phenotype Guided Medical Vision–Language Pretraining
Exploring Efficient Reasoning Segmentation with Small Language Models
Molmo-Point: Better Pointing for VLMs with Grounding Tokens
Benchmarking Vision-Language Models for Microscopic Plant Image Understanding
HIVE: Understanding Post Hallucination Reasoning in Vision Language Models
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating
NaVLM-PVC: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning
Gripper-aware Vision Language Action Models
Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
Are Video Reasoning Models Ready to Go Outside?
Towards Reliable Medical Large Vision-Language Models via Counterfactual Preference Optimization
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
VisWordBench: Bridging the Gap in Cross-modal Reasoning for Multimodal Large Language Models
The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models
3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing
Anchored, Not Graded: How Vision-Language Models Fail at Slant-from-Texture Perception
HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks
Towards More Efficient Decoding for Autoregressive Vision-language-action Models
Visual Prompt Discovery via Semantic Exploration
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark and Metrics for VLM Image Tampering
Learning to Deny: Action Denial in Multimodal Large Language Models
DIVA: Instruction-Aware Vision Token Pruning via Dual-Probe Attention Discrepancy
CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models
Learning from Primitive: Probing Visual Reasoning of LVLMs via Counting
Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection
Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models
PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
ScenarioControl: Vision-language Controllable Vectorized Latent Scenario Generation
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision–Language Models
CURE: Cumulative Knowledge Reuse for Efficient Device-Server Hybrid Inference in Vision-Language Models
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
Same Pool, Different Answer: Stable Best-of-N Selection for Vision-Language Models
SAFE-Pruner: Semantic Attention–Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
Do Vision Language Models Recognize Visual Ambiguity?
ATOMIC: A Domain-Specific Vision-Language Model for Transmission Electron Microscopy
Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
On Test-Time Scaling for Vision-Language Models
Overlap-Consistent View Decomposition for Adapting Vision--Language Models to 360° Panoramas
SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
Why Do Vision Language Models Struggle To Recognize Human Emotions?
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production–Living Simulations with Stardew Valley
BeTTER: Diagnose the Illusion of Embodied Reasoning in Vision-Language-Action Models
3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
PARL-VLA: Pruning-Aware On-Policy Reinforcement Learning for Vision-Language-Action Model
Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
URoPE: Universal Relative Position Embedding across Geometric Spaces
CulinaryCut: A Physics-aware Vision-Language-Action Benchmark for Food Cutting via Material Point Method
EoS-FM: Can an Ensemble of Specialist Models act as a Generalist Feature Extractor?
VERITAS: A Multi-agent Co-scientist for Verifiable Image-Derived Hypothesis Testing
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection
From Gaze to Meaning: An AI Agent for Unified Zero-Shot Grounding and Explanation
Natural Image Pretraining Improves Abstract Reasoning
Reinforcing Vision-Language Models for Image Quality Assessment with Grounding Process Rewards
MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving
Seeing Isn't Orienting: A Cognitively Grounded Hierarchical Benchmark for Object Orientation in MLLMs
Ceptor: Vision-Language Model-Infused Diverse Guidance for Detecting Anything
Gender Bias in Vision-Language In-Context Learning
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
Personalize Your Large Vision-language Models With In-context Prompt Tuning
TecoPrompt: Temporal-Conservative Prompt Learning for Vision-Language Models
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
DICE: Disentangled Instance-Class knowlEdge prompt tuning via SAE for Vision-Language Models
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
EmbedCopilot: Evaluating Vision-Language Models for Hardware-Aware Embedded System Development
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
Delineating Knowledge Boundaries for Honest Large Vision-Language Models
Teaching Vision-Language-Action Models What to See and Where to Look
ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision–Language Models
Open Your Eyes: Benchmarking the Detection of Fabricated Realities and Weaponized Ethics in VLMs
Evaluating and Understanding Model Editing for Medical Vision Language Models
CrossView: Can Vision-Language Models Reason Across Cameras?
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
VISOR++ : VISUAL INPUT BASED STEERING FOR LARGE VISION LANGUAGE MODELS
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
MomentSeg: Moment-Centric Sampling for Enhanced Referring Video Object Segmentation
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
Linear Scaling Video VLMs for Long Video Understanding
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
Online Reasoning Video Object Segmentation
InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
Accelerating Text-to-Video Generation with Calibrated Sparse Attention
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
SALT: Self-Consistent Distribution Matching with Cache-Aware Training for Few-Step Video Generation
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
Surprise Forcing: What to Remember, When to Skip in Long Video Generation
Consistent Video-to-Video Translation via Explicit Correspondences
InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing
SV-TAD: Native Sparse Convolutions for Efficient Temporal Action Detection
Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility
PhysDrape: Learning Explicit Forces and Collision Constraints for Physically Realistic Garment Draping
MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning
Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning
SpecEyes: Accelerating Agentic Multimodal LLM via Speculative Planning and Perception
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
S2Gest: Split-Scan State Space Models for Dynamic Hand Gesture Recognition
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
Latent Visual Diffusion Reasoning with Monte Carlo Tree Search
Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
DART: Deformable Adaptive Reasoning with Temporal Queries for Online Skeleton-Based Action Recognition
FingerCap: Fine-grained Finger-level Hand Motion Captioning
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
LogFA: Efficient Feature-Space Data Augmentation for Egocentric Temporal Action Segmentation
OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
SPAR: A Sequential Primacy and Attribution Ranking Framework for Skill Determination
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
FeatTracker: Short- and Long-Range Temporal Feature Consistency for Robust Underwater Object Tracking
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition
SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning
LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning
DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding
Audio-Visual Continual Test-Time Adaptation without Forgetting
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
Cambrian-P: Pose-Grounded Video Understanding
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
SMART: When is it Actually Worth Expanding a Speculative Tree?
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
VoCa: Unified Autoregressive Modeling for Talking Audio-Video Generation
Video-Holmes: Can MLLM Think like Holmes for Complex Video Reasoning?
Guiding the Blind: Generalizing GUI Agents to Unseen Websites via Multimodal Tutorials
Test Time Training for Long Videos via Frame Forgetting Network
Trajectory-Level Continuous Action Representation for Robotic Manipulation
Temporal and Cross-modal Alignment for Enhanced Audiovisual Video Captioning
SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning
Towards Long-Form Spatio-Temporal Video Grounding
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Pathwise Test-Time Correction for Autoregressive Long Video Generation
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding
DG-Force: Disentangling and Gathering Forensic Cues is Needed for Image Manipulation Localization
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
C3ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection
PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Egocentric Procedure Parsing
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
AnaPFL: When Closed-Form Solutions Meet Generalizationand Personalization in Personalized Federated Learning
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
Counterfactual World Models via Digital Twin-conditioned Video Diffusion
ProAct: Agentic Lookahead in Interactive Environments
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA
SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking
HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers
Q-TriM: Question-Guided Tri-Modal Attention for Audio–Visual Question Answering
VC-VAE: Leveraging Video Codecs for Training-Efficient and High-Fidelity Video VAE
LINA: Learning INterventions Adaptively for Physical Alignment and Counterfactual Generation in Diffusion Models
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO
Table-MCR2TR: Merged-Cell-Aware Table Recognition via Reinforced Multimodal Language Models
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
Open-Vocabulary Long Term Action Anticipation
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video-LLMs
Thinking in Streaming Video
Break Visual-Linguistic Asymmetry: Unleashing VLM's Cross-Modal Potential for General Face Forgery Detection
TaxoGrasp: Taxonomy-Guided Human Grasp Synthesis with Sparse Contact Constraint
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-visual Language Models
SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation
Reward Modeling for Computer-Using Agent from Video Execution
SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation
EgoCogNav: Cognition-aware Human Egocentric Navigation
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
LARY: A Latent Action Representation Yielding Benchmark
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization
Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
From Script to Shot: A Benchmark for Grounding Screenplays in Movies
RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction
S3-Prune: Stability-Aware Token Budgeting for Long-Form Video-Language Models
Video-Oasis: Rethinking Evaluation of Video Understanding
Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
H2SVC: Head-aware Heterogeneous Streaming Video Cache for Online Video Understanding
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
EgoPHI: Estimating 3D Hand-Object Contact and Force from Egocentric Vision
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
Accelerating Diffusion Models via Equal-Risk Caching
Continuous Heart Rate Variability Estimation from Egocentric Systems for Skill Assessment
AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine
Natural Language Camera Movement Understanding
UniTemp: Unlocking Video Generation in Any Temporal Order via Autoregressive Distillation
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
Towards Effective Long Video Understanding: Dynamic MAS Construction via Meta-Agent
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
FeVOS: Foresight Expression Video Object Segmentation
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation
FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Perceptual Projection Pruning: Diversity-Aware Video Token Pruning for Multimodal Large Language Models
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
Unifying CNNs and ViTs for Learning-Efficient and Scalable Variational AutoEncoder
RotateAttention : RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation
CTEPM: Continuous-Time Event Process Memory for Long-Video Language Models
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
Remembering Across Blocks: Topology-Conditioned Block-Progressive Memory for Skeleton-Based Action Recognition
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
LogicIR: Logic Gate Networks for Image Restoration
FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration
Freqformer: Image-Demoiréing Transformer via Effective Frequency Decomposition
NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction
Posterior Samplings are Missing Modalities Generators for Medical Image Translation
TRAM: Finetuning-Free Test-Time Adaptation for Generalized Face Anti-Spoofing with Only a Few Bonafide Samples
MegaFlow: Zero-Shot Large Displacement Optical Flow
From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution
ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion
WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching
IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
Synesthesia via Direct Latent Augmentation: Bypassing the Decode-Encode Loop for Cross-Modal Distillation
Flexible Control of 3D CT Generation via Text and Semantically-Defined Segmentation Prompts
Resolution-Agnostic Neural Operators for Multi-Rate Sparse-View CT
MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations
SeekFlow: Synergizing Radiology and Pathology Foundation Models for Precision Oncology via Knowledge-Guided Evidence Flow
MArFE: Multi-Contrast MRI Arbitrary Scale Super-Resolution with Fourier Enhancement
Kiroshi: An Agentic Perception System for High-Accuracy Image Parsing
EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics
Hybrid Event–Frame Sensors: Modeling, Calibration, and Simulation
Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution
DRIFT: Difficulty-aware Rectified Flows for Through-plane MRI Super-Resolution
VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
RePer-360: Releasing Perspective Priors for 360° Depth Estimation via Self-Modulation
Scaling Whole-Slide Pathology Foundation Model Pretraining with Billions Off-the-Shelf Tokens
FlexiBrain: Resolution-Agnostic Voxel-Level Encoding for Native fMRI
FreqPhys: Repurposing Implicit Physiological Frequency Prior for Robust Remote Photoplethysmography
Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution
Quantile‑Adaptive Temperature Scaling for Confidence Calibration
STARLINC: Satellite Trail Artifact Removal using Inter-Frame Correlation
Masked BRep Autoencoder via Hierarchical Graph Transformer
RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement
RainODE: Continuous-Time Precipitation Forecasting with Latent Neural ODEs
TopoGAT: Plug-and-Play Topological Graph Attention for Fine-Grained 3D Segmentation
Learn to See the Unseen in Low-light Spike Streams
2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction
NGPS: Structure-Preserving Self-Supervised Denoising via Neighbor-Guided Patch Sampling
Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation
CUST : Clustered Unit-level Similarity Transformer for Lightweight Image Super-Resolution
DeLux: Cross-Modal Local Artifact Restoration in Video Using Neuromorphic Data
LIIFusion: Coarse-to-fine Framework for Generative MEF via Implicit Neural Representation
Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping
XPos3R: Cross-Modal Transformer for Intraoperative 2D/3D Registration
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA
Semantic-Anchored Evidential Fusion for Domain-Robust Whole-Slide Survival Analysis
UniH3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
Pol-CACTI: A System and dataset forHigh-Speed Polarized Video Compressive Imaging
MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation
Foundation-Guided Representation Alignment for Multimodal Medical Image Registration
Φeat: Physically-Grounded Material Feature Representation
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
Physics-Grounded Disentangled Flow Modeling for Brain Disease Progression Trajectory
Ice Cloud Geometry Retrieval with Calibrated Uncertainty from Passive Satellite Imagery
From Minimal Clinical Prompts to 3D: Spacing-Aware Prompt Propagation for Multimodal Prostate Lesion Segmentation in bpMRI
Event-based Sparse-view Background-Oriented Schlieren Tomography
CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories in Latent Space
LineGraph2Road: Structural Graph Reasoning on Line Graphs for Road Network Extraction
WARP: Wide Attention with Rich Projections for Image Super-Resolution
MorphJEPA: Morphology-Aware Latent Prediction for Hyperspectral Images
G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening
Proximity-Constrained Counterfactual Decoding for Hallucination-Robust Medical VQA
MOOZY: A Patient-First Foundation Model for Computational Pathology
Finding Highlight Images In Your Albums:From Benchmark To MLLM
Zero-shot Depth from Defocus
HybridSim: A Physics–Learning Hybrid Digital Twin for mmWave Human Sensing
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
MedRepBench: Benchmarking Structured Understanding of Medical Report Images
Reconstructing Dense Depth of Dark Scenes with Sparse LiDAR, Noisy Events, and Blurry RGB
Seeing What Matters: Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Report Generation
SPHERE: From MRI Sampling Mechanisms to Spatial Priors for Generalizable Brain Tumor Segmentation
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
MCPNet:Masked Coordinate Pooling-based Attention Network for Medical Landmark Detection
SVI360: Spherical Video Interpolation
Task-driven Processing with Coarse-to-Fine Glimpse-based Active Perception
From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation
EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation
SurvMILKD: A Weakly Supervised Survival Analysis Framework for Multi-Teacher Knowledge Distillation using Pathology Foundation Models
DualResPS: Dual-Resolution Photometric Stereo Using a Frame-Event Hybrid Camera
PolarAPP: Beyond Polarization Demosaicking for Polarimetric Applications
Fast and Accurate Image Restoration with Rank Enhanced Linear Attention
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation
Integrated Forward–Inverse Network for Reconstruction for Lensless Image Reconstruction
Gaussian Volumetric Representation for Efficient Shear–Warp Visualization
RobustRDP: Advancing Reaction Diagram Parsing via Synthetic-to-Real Data Scaling and Robustness-Oriented Training
Physics-Guided Deep Learning for Linear Mueller Matrix Acquisition
EventVGGT: Exploring Cross-Modal Distillation for Consistent Event-based Depth Estimation
Towards Reliable Multi-Label Classification via Conditional Dependency Modeling
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
Stokes-Informed Diffusion for Robust Linear Polarization Estimation
RIGS: Radar-Informed Gaussian Splatting for Uncertainty-Aware 3D Occupancy and Motion Prediction
Structured SIR: Efficient and Expressive Importance-Weighted Inference for High-Dimensional Image Registration
GeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence Sampling
ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation
3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems
NeuralDMD: Interpretable Untrained Neural Network for Imaging from Sparse and Noisy Observations
Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers
Off the Planckian Locus: Using 2D Chromaticity to Improve In-Camera Color
BiCE-HG: A Bi-Conditional Egocentric Hand Gesture Dataset for Intelligent Reality Systems
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept Erasure
Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
Escaping the Low-Frequency Bias: Adversarial Frequency Perturbation for Generalisable Gaze Estimation
Fully Rotation-Equivariant Spectral-Spatial Learning for Multispectral Object Detection
On the Reliability of Cue Conflict and Beyond
This Looks Distinctly Like That: Grounding Interpretable Recognition in Stiefel Geometry against Neural Collapse
Learn to Rank: Visual Attribution by Learning Importance Ranking
Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning
Scaling Laws for Black-box Adversarial Attacks
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
Data Circuit Breaker: Identifying Training, Test, and Generated Data in Image Generative Models
Denoised Variance-Based Pruning with Optimal Brain Bias Compensation
Adaptive Neural Dynamics for Robust Geometric LiDAR-Inertial State Estimation on UAVs
Compact and Structurally Transparent Cervical Cytology with Geometry-Driven Features and Closed-Form Attention
TR-MoE: Temporal Reliability-Aware Mixture-of-Experts for Robust Tracking
Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
Frequency Director: Learnable Mixture of Frequency Experts for Unified Concealed Scene Segmentation
Learning Probabilistic Prompt for Continual Learning
Intrinsically Stable Spiking Neural Networks: Overcoming the Performance Barrier in the Absence of Batch Normalization
R-ESC: Robustly Erasing Space Concepts via Stochastic Feature Remapping
Defending from GeoLocalization through Adversarial Road Trips
IConE: Batch Independent Collapse Prevention for Self-Supervised Representation Learning
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in the Semantic-Residual Space
Fast and Flexible Robustness Certificates for Semantic Segmentation
On the Plasticity Collapse in Continual Machine Unlearning
Weight Feedback Computes the Exact Jacobian Transpose in Modern Deep Networks
Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning
Proteus: Model Leakage-Induced Adversarial Attack in Federated Learning
Geometry Aware Reliable Instance Selection for Noisy Partial Label Learning
Rethinking Temporal Modeling in Visual Object Tracking via Decoupled Auxiliary Supervision
ORBIT: Overcoming Hallucination Risks via Bi-manifold Interaction and Traction
Beyond Filter Pruning: Top-K Spatial Selection for Efficient Neural Networks
Enhancing Pretrained Model-based Continual Representation Learning via Guided Random Projection
BackTranslation2.0 - A Linguistically Motivated Metric to Assess Sign Language Production
AracNet: Revealing Debiasing Signals across Layers with Shallow Monitors
Exposing Implicit Vulnerabilities in Text-to-Image Models via Adversarial Agentic Probing
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
Closing the Capacity–Convergence Gap: Globally Optimal Configuration of Implicit Neural Representations
Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations
RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations
Explainability-aware Frustum Attack: Exposing Structural Vulnerabilities in LiDAR-Based 3D Object Detectors
Robustness Emerges Early in Training Dynamics, but Is Not Preserved
REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
When Higher Order Hurts: Pre-Asymptotic Order Collapse in Generative ODE Sampling — A Theory of Discretization–Learning Interaction
UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models
SPICE: Simple Polysemantic feature Interpretation via Clustering-based Explanations
BrainRiem: Riemannian Prototype Learning for Source-Free Cross-Site Brain Network Diagnosis
Bridging Theory and Practice in Source-Free Domain Adaptation via Adversarial Proxy Perturbation
Structured-Noise Masked Modeling for Video, Audio and Beyond
SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning
Causal Intervention in Concept Bottleneck Models
iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders
Improving Adversarial Robustness by Mitigating Instability through Relearning
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection
Dynamic-V2C: Editable and Continual Vision-to-Concept Bottleneck Models via Influence Functions
FaceArmor: A Universal Facial Image Protection Against Diffusion-Based Manipulations
Multi-Anchor Distillation with Text-Guided Analytic Classifier for Continual Learning
Deep Noise Label Learning via Effective Rank Reduction
Adversarial Attack and Disturbance Detection by Hadamard-Coded Output Representations for Object Detection and Semantic Segmentation
Breaking High Confidence: Practical Face Impersonation under High-Security Thresholds
Simple Filtering Improves Masked Autoencoders
Leveraging Dark Knowledge for Intrinsic Multimodal Out-of-Distribution Detection
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
Hyperbolic Hierarchical Clustering for Visual Representation Learning
VQT: Vector Quantization Tuning for Efficient Fine-tuning and Compression of Pre-trained Vision Transformers
Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
Steerable Vision Transformers
Task Alignment: A simple and effective proxy for model merging in computer vision
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Diffusion-Based Immersive Visual Reasoning
Efficient Quantization-Aware Adaptation for Visual Foundation Models
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
DnA: Denoising Attention for Visual Tasks
Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs
Enlightening Photographic Style Transfer with a Self-Supervised Photographic Embedding
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
Verifying Cancer Segmentation in Vision Transformers via Internal Concepts
Quick ViTs: Speeding up Vision Transformers through Equivariance
Auto-Prompting: Layer-Specific Prompt Fusion Discovery via Differentiable Search
DeCoPatch: Revealing Causal Latent Subspaces in Vision-Language Models for GUI Grounding
Rethinking Attention Reallocation for Multimodal Emotion Recognition
Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
Invisible Shortcuts: Why Vision Encoders Know Your Camera
Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
World Knowledge in the Weights: Reading Concept Circuits of Vision Transformers
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
CL4D: Contrastive Language–4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
Why Can Accurate Models Be Learned from Inaccurate Annotations?
MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein
VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision
MuSViT: A Foundation Vision Model for Sheet Music Representation
Moving Beyond More Views: Redundancy-Aware Ego–Exo Fusion for Proficiency Estimation
MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
Background Blurring Matters: Improving Visual Grounding by Merging Text-Irrelevant Tokens
Contrastive-Guided Self-Supervised Latent Visual Reasoning for Hallucination Mitigation
SyncVL: Synchronizing Vision ⟷ Language Using Unsupervised Adaptation
Test-Time Registers as Global Priors for Tokenized Image Generation
Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
SRRA: Stable-Rank-Based Residual Adaptation for Generalizable Deepfake Detection
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
Mechanistic interventions for explainable digital pathology uncovers adversarial vulnerabilities
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
Tesselating The Earth
CLIMP: Contrastive Language-Image Mamba Pretraining
Human-like Object Grouping in Self-supervised Vision Transformers
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
BLOB-Q: Boosting Low Bit ViT Quantization via Global Optimization on Model Distortion
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers
AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware
SafeSAE-VLA: Interpreting OpenVLA Progress Dynamics with Sparse Feature Analysis
When Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High Compression
ECC: Encoder-Centric Corruption for Fine-Grained Vision in VLMs
Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Human Motion
LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
ORFC: Orthogonal Reparameterization for Low-Bitrate ViT Feature Coding
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
Make Geometry Matter for Spatial Reasoning
Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers
TETO: Tracking Events with Teacher Observation for Motion Estimation and Frame Interpolation
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
Syn4D: A Multiview Synthetic 4D Dataset
FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion
MASS: Motion-Aligned Selective Scan for Flow-Based Video Frame Interpolation
4D-VGGT: A SpatioTemporal Foundation Model for Dynamic Scene Geometry Estimation
Auto3R: Automated 3D Reconstruction and Scanning via Data-driven Uncertainty Quantification
PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection
MemPose: Category-level Object Pose Estimation with Memory
RoMa v2: Harder Better Faster Denser Feature Matching
Event-driven Motion Deblurring via Trajectory-based Kernel Reconstruction
AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation
TACO-Net: Topological Signatures Triumph in 3D Object Classification
VectorReLoc: Reliable Vectorized SD Map Visual Re-localization with Contrastive Feature Alignment
LoMa: Local Feature Matching Revisited
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Prosthesis-Aware 3D Human Pose Estimation: A Dataset and Benchmark for RSP Users
360Anything: Geometry-Free Lifting of Images and Videos to 360°
SemGAN: A Semantic and Hierarchical Adversarial Network for 3D Human Pose Estimation
Rolling Shutter Relative Pose Estimation Made Practical
Honey, I Shrunk the Arc de Triomphe!
Gravity-aware partially calibrated absolute pose estimation from affine- or rotation-covariant features
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
General Self-Calibration with Varying Intrinsics
Pointer-CAD v2: Plan-Then-Construct CAD Generation with Dimension-Aware Parametric Precision
UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Driving
Delaunay Canopy: Building Wireframe Reconstruction from Airborne LiDAR Point Clouds via Delaunay Graph
Ray-Path-Aware Virtual Point Removal on 2D Layer-Wise Nearest Point Map
Rolling Shutter Camera Self-Calibration
GARDEN: Gravity-Aligned Reconstruction of Disentangled ENvironments from RGB images
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
Holo360D: A Large-Scale Real-World Dataset with Continuous Trajectories for Advancing Panoramic 3D Reconstruction and Beyond
ObjectForesight: Predicting 3D Object Trajectories from Human Videos
Sequential Visual Place Recognition: Exploiting Trajectory Priors for Robust Localization
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
Towards Reconfigurable Visual Feature Compression
Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation
PriorPose: Reference-Guided Joint Deformation and Alignment for Category-Level Object Pose Estimation
Learning Global Camera Poses from Noisy View-Graphs for Structure from Motion
OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization
VisTa3D: A Dataset and Benchmark for Vision, Tactile, and 3D Point Clouds-based Thin Object Reconstruction
Coarse-to-fine Contrast: A Hybrid Self-supervised Method for Non-rigid 3D Shape Matching
Roam2Room: A Unified Floorplan-to-Furnished Framework for Controllable Indoor Scene Generation
Rethinking Detection Calibration: A Coordinate Perspective
Following Motion for Sequential Modeling in Video Frame Interpolation
Rosetum3D: A Large-Scale 3D Vision Dataset from Preharvest Roses
Mitigating Radar-Inertial Calibration Ambiguities via SO(3) Manifold Steering
Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models
Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
Equivariant Symmetry-Aware Head Pose Estimation for Fetal MRI
LaVPR: Benchmarking Language and Vision for Place Recognition
E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects
Estimating Velocity and Spin of Spherical Objects from Rolling-Shutter Image(s)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras
SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
Lightweight Online Reinforcement Learning for Block Decomposition of CAD Models
Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding
Training-free Controllable Motion Generation under Heterogeneous Constraints
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation
There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion
LivingWorld: Interactive 4D World Generation with Environmental Dynamics
OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
Sector-Level Cross-View Geo-Localization with Implicit Orientation via Azimuthal Scanning
Category-Level Articulated Object Pose Estimation via Pose–Shape Hypothesis Generation and Verification
Ego-Human Motion Prediction with 3D-Aware LLM
TIDES: Time-Derivative Event Simulation via Deformable Reconstruction
SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data
PoseImageNet: Pose Estimation for Extensive Classes Based on Rich Structure Prototypes
Variational Patch Gating for Training-Free Few-Shot Classification
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images
SemLight: Distilled Semantic–Geometric Fusion for Efficient Local Feature Matching
Coordinate Singularities Break Conformal Coverage for Gaze and Head Pose
RADmesh: Remesh-Aware Mesh Deformation
GrowFields: Compositional 4D Neural Fields for Topology-Changing Plant Growth
AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision
IoUCert: Robustness Verification for Anchor-based Object Detectors
Sparsity-Inducing Divergence Losses for Biometric Verification
Leaving the City: A Large-Scale Aerial Dataset for Cross-Season Localization in Unstructured Environments
Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided VLM
StreetForward: Perceiving Dynamic Street with Feedforward Causal Dynamics
MoBa-GS: Learning a Spatially-Varying Motion Basis over a Dynamic Canonical Space for 4D Reconstruction
RegHead: Non-Humanoid Head Blendshapes via Feed-Forward Registration
Tempo-SAM3D: Monocular Video to 4D via Temporal Memory-Guided Generation
UF0-6D: Unified Flow-based Zero-Shot 6D Object Pose Estimation without Refinement
PrimitiveUDF: Primitive-Based Unsigned Distance Fields for Surface Reconstruction from Point Clouds
Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition
Analytic Bayesian Uncertainty for LiDAR Segmentation: A Single-pass Generative Approach
Discovering Geometric Biases in 3D Face Reconstruction: A Curvature-Aware Spectral Framework for Fairness Evaluation
WiFi-JEPA: Self-supervised Learning for WiFi-CSI 3D Human Pose Estimation
PolyLayout: Multi-room Manhattan Layout Estimation
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
Revisiting the Volumetric Data of 4DME: Compression, Extension and Benchmarking for Micro-Expression Analysis
Visible Yet Unrecognizable: Frequency-Selective Facial Privacy via Attention
Learning Manifolds in High-D Point Embedding for Anisotropic Surface Approximation from Unstructured Point Clouds
NEOMAP: Novel-View Synthesis via Noise Initialization by Manifold Alternating Projection
Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
Forecasting Animal Motion
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
GO-Renderer: Generative Object Rendering with 3D-aware Controllable Video Diffusion Models
Walking in the Implicit: Interactive World Exploration via Neural Scene Representation
MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction
VOID: Video Object and Interaction Deletion
StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
One Video, One World: Turning Monocular Video into Physical 4D Scenes
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
DiffProxy: Multi-View Human Mesh Recovery via Diffusion-Generated Dense Proxies
TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
MuCHeR: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking
Human Mesh Modeling for Anny Body
TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
EgoSim: Egocentric World Simulator for Embodiment Interaction Generation
DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
Predictive Structure Improves Video Diffusion Dynamics
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
What Moves? Localized Motion Representations for Compositional Scene Control
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
ECHO: Ego-centric Modeling of Human-Object Interactions
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
MetaPoint: Unlocking Precise Spatial Control in Visual Generation
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation
Narrative-Driven Paper-to-Slide Generation via ArcDeck
Meric: A Unified Framework for Multimodal Music Generation and Retrieval via Representation Space Anchoring
Scalable Cross-embodiment Dexterous Grasping via Morphology-Prior Diffusion
DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models
STANCE: Controllable Video Generation for Structured Dynamics via Sparse-To-dense ANChored Encoding
PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
Learning to Generate Rigid Body Interactions with Video Diffusion Models
Self-supervised Garment Dynamics with Persistent Wrinkles
Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
Egocentric World Model for Photorealistic Hand Object Interaction Synthesis
EMOTE: Expressive Motion and Shape Disentanglement for Human Animation
Dynamic World Generation Made Efficient
Moiré Video Authentication: A Physical Signature Against AI Video Generation
Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Synthesis
PhysPO: Physics-Aware Local Preference Optimization for Physically Consistent Video Diffusion
Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
ID-PreFeR: ID-Preserving Face Restoration with Mixed Data Quality
3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
OctWorld: Long-Range World-Consistent Video Generation with Octree-based 3D Mapping
Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
MemLearner: Learning to Query Context Memory for Video World Models
REON-NVS: Real-Time Online Novel-View Synthesis from Sparse-View Videos
GRE-Diff: Gaussian Room Embeddings for Structured Layout Diffusion
FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World
Video Generation Models Are Inherent Lighting Estimators
IC-World: In-Context Generation for Shared World Modeling
Layer-Aware Video Composition via Split-then-Merge
Perceiving Better Moments: Cover Frame Reselection and Enhancement for Live Photos with the Live2K Dataset
PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark
ReCamDriving: LiDAR-Free Camera-Controlled Video Synthesis for Novel Trajectories
PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Unified Panoramic–Gaussian Representation for Monocular 4D Scene Synthesis
ActionParty: Multi-Subject Action Binding in Generative Video Games
VERTIGO: Visual Preference Optimization for Cinematic Camera Generation
Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
PASTEL: Panoramic Alignment for Monocular 4D Scene Reconstruction
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
CortexVideo: A Semantic-Spatial Dual-Anchor Framework for High-Fidelity fMRI-to-Video Reconstruction
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
LOOM: Weaving Geometry-Consistent Human-Object Interaction Videos via Progressive Curriculum Learning
Head Avatars with Dynamic Explicit Hair
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control
HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
HiChor: Hierarchical Choreography Generation from Pop Music with Choreographic Primitives
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
OmniColor: A Unified Framework for Multi-modal Lineart Colorization
Tuning-free Visual Effect Transfer across Videos
Seeing Fast and Slow: Learning the Flow of Time in Videos
ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Synthetic Data
In-Context Sync-LoRA for Portrait Video Editing
EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning
StyleFusion360: View-Consistent Head Stylization via Adaptive Style Modulation
SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion
AnyView: Synthesizing Any Novel View in Dynamic Scenes
Learn2Fold: Structured Origami Generation with World Model Planning
The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
PhysConvex: Physics-Informed Dynamic Convex Fields for Reconstruction and Simulation
NeuIDO: Neural Intrinsic Dynamics Operator for Physics-Informed 4D World Models
BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection
AutoPhyX: Automatic Text-Condition Physics Property Generation
Who Does What and Where to Go: Orthogonal Alignment and Hierarchical Planning for Multi-Entity Trajectories
MorphGS: Morphology-Adaptive Articulated Motion Transfer from Videos
PartCHOI: Part-Aware Guidance for Clothed Human-Object Interaction Generation
EmoteGPT: 3D Human Facial Expression from Natural Language Descriptions
ROSE: Real-Time Open-World Scene Understanding from Monocular Video via Compact Multimodal 4D Scene Graphs
Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction
ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling
SA-V2V: Training-Free Subject-Aware Video-to-Video Personalization
Multi-Modal Controlled Coherent Motion Generation
KineticGS: Momentum-driven Coherent 4D Gaussian Splatting for Monocular Dynamic Scene Reconstruction
PWM-ArtGen: Part World Model for Articulated Object Generation
Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models
D-Rex : Diffusion Rendering for Relightable Expressive Avatars
LUNA: Learning Universal 3D Human Animation Beyond Skinning
Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset
Towards Real-World Wearable Motion Reconstruction
GraphVid: Interactive Graph-Controllable Video Generation
SignRefine: Adapting Foundational Video Models for Sign Language Generation
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
RealDyadic: Synthesizing Realistic Dyadic 3D Dialogue with Neural Appearance Priors
Taming Dynamic Clutter: Variance-Driven Adaptive Gain Control for Bio-inspired Small Target Detection
Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation
CAST3D: Customizing Arbitrary 2D Assets into 3D World
RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation
JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
SplatCtrlA: Generalizable Single Image to Fully Controllable 3D Avatar
Beyond Pixel Mimicry: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
ExpoMotion: A Large-Scale Benchmark and A Householder Projection Network for Multi-Exposure Fusion
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
A Dual-Transformer Architecture with Cross-Attention for Multi-Camera View Recommendation
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency
ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
LooseControlVideo: Directorial Video Control using Spatial Blocking
DynEval: Holistic Evaluations of T2I Generative Models in the Wild
MotionSplicer: Part-Based Motion Editing for 4D Volumetric Videos
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Versatile Editing of Video Content, Actions, and Dynamics without Training
TAQ: Static-Deployable Temporal-Aware Quantization for Real-World Video Super-Resolution
DIVER: Disentangling Camera–Object and Active–Passive Motion for Video Generation
A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
ARMS: Anchor–Relational Motion Streaming for Seamless Solo-Social Motion Transitions
TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
Music-to-Dance Generation via Atomic Movements
MSEditor: Toward Consistent Multi-Shot Video Editing
OmniLife360: A Benchmark for 3D Reconstruction from In-the-Wild 360° Captures
MotionAnymesh: Physics-Grounded Articulation for Simulation-Ready Digital Twins
SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation
YeTI: You Only Need Two Noisy Images for Real-World sRGB Noise Generation
TextFace: Compositional Text-Guided Identity Preserving Face Synthesis for Face Recognition
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
FlowLess: Controlling Abstract Image Generation
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
WorldCache: Content-Aware Caching for Accelerated Video World Models
Video Generation Models are General-Purpose Vision Learners
Scaling Dense Prediction with Latent Decoding
Kinematics-Agnostic 3D Human Motion Prediction via Equivariant Latent Diffusion
PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection
Dual Masked Generative Adversarial Transformer for Unsupervised Domain Adaptation
Complex-Valued 2D Gaussian Representation for Computer-Generated Holography
Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
On the Diffusibility of High-Dimensional Latents
Diffusion-Based Material Regularization for Physics-Based Inverse Rendering
ReDesign: Recovering Editable Design Structures from Raster Images via Agentic Decomposition
Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat
ShellMaker: Language-Guided Exterior Completion under Structural Constraints
DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference
EVAR: Edge Visual Autoregressive Models via Principled Pruning
JSON: Jigsaw Self-play Optimization for Normalizing Flows
HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching
FlowFace: Rectifying Identity Conditioning with Riemannian Geometry for Face Generation
Bridging Online and Offline Handwriting via Differentiable Physical Rendering
FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy
FontCopilot: Towards Generalist Multimodal Large Language Models for Holistic Chinese Font Engineering
Entropy-Controlled Flow Matching
Exo2EgoPolicy: Pose-Aligned Cross-View Policy Learning
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models
HuCollisionField: Resolving Self-Collisions via Neural Fields for Human Prediction
Multi-View Foundation Models
AV2T-Gen: Aerial Visible to Thermal Generation with Environment and Vehicle State Guidance
ThermoGS: Decoupling Physical Surface Attributes for Spatio-Temporal Thermal Field Emulation via 4D Gaussian Splatting
NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
Adaptive Noise Covariance Scheduling under Riemannian Metrics for Diffusion Models
Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI-to-image Reconstruction via ROI Encoder And visual Mapping
APT: Anchor-aligned Perturbations for Tamper Localization in Fully Regenerated Images
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
DuoFlow: JVP-Free Finite-Difference Mean Flows for One-Step Image Generation
TopoFuse: Topology-Aware Tri-Planar Fusion for 3D Cryo-Electron Tomography Segmentation
Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
Teaching an Agent to Sketch One Part at a Time
DOGE: Differentiable Bézier Graph Optimization for Road Network Extraction
JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising
Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction
Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings
MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes
TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields
STAT: Soft Tail-dropping for Adaptive Visual Tokenization
Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images
PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion
TreeSRNF: Square-Root Normal Fields for Generative Modelling of the Geometric and Structural Variability in Tree-like 3D Objects
Temporally Stable Generative Illumination with a One-Step Diffusion Model
DeCoFlow: Structural Decomposition of Normalizing Flows for Continual Anomaly Detection
URHead: A Unified UV-Space Representation for Joint Mesh–3DGS Optimization in Head Avatars
Hybrid-LUT: Channel-Aware Hybrid Lookup Table and Filtering for Efficient Image Restoration
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
Pretrained Video Models as Differentiable Physics Simulators for Urban Wind Flows
WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
SAM2Matting: Generalized Image and Video Matting
Learning Geometry-Aware Embedding Fields for Intrinsic Riemannian Mappings
Reflecting Process Expertise in Procedural Material Generation
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
DANTE-W: Diffuse Albedo Neural Texturing in the Wild
Language-Guided Transformer Tokenizer for Human Motion Generation
Text-based Tactile Graphics Generation for the Visually Impaired
OmniX: From Unified Panoramic Generation and Perception To Graphics-Ready 3D Scenes
Reconstructing 3D Human-Object Interaction via a Unified Triplane Space
Pixel Ignores, Superpixel Sees: Adverse Weather Image Restoration via Semantic-Center SSM
Repurposing Geometric Foundation Models for Multi-view Diffusion
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
A Scalable Vector Graphics Latent Space
Differentiable Polarized Path Tracing
Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading
Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs
ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection
Deformable and Multi-view Gradient-Aligned Physical Adversarial Camouflage
Through Van Gogh’s Eyes: Global Style Transfer with Diffusion Model
HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization
Generalization and Memorization in Rectified Flow
LumiTokens: 3D Relighting via Token-Space Lighting Transformation
Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation
Penetration-Free Compositional 3D Generation via Gaussian Surface Offset
Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution
PRISM: Latent Composition Consistency for Single-Image Reflection Removal
MagnetGS-Mesh: High-Quality Multi-Object Mesh Reconstruction via Adaptive Surface Optimization
WorldAgents: Can Foundation Image Models be Agents for 3D World Models?
Robust 3DGS-based SLAM via Adaptive Kernel Smoothing
Stitched Embeddings: A Unified Latent Space for 3D Garments and 2D Patterns
TiltDiff: Tilted Weight-Space Diffusion for Neural Network Generation
Why Linear Probing Works: Non-Vacuous Generalization Bounds via Effective Dimension
CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
NeuralGarSim: Geometry-agnostic Garment Simulation with Neural Fields
ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling
InnoText: A Unified Model for Visual Text Generation and Editing
CORE-V: Chain-Of-thought REasoning for Image Editing with Visual Interaction
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
NumColor: Precise Numeric Color Control in Text-to-Image Generation
Vision Bridge Transformer at Scale
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Based Video Editing
Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models
VGEdit: Unlocking Video Generation Priors for Reasoning-Informed Image Editing
MoScale: Autoregressive Next-Scale Prediction for Human Motion Generation and Editing
Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection
Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer
Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation
Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning
MANGO: Unleashing Image Generation Capability of Unified Multimodal Models
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
UEval: A Benchmark for Unified Multimodal Generation
LISA: Locality-Informed Speculative Decoding for Accelerating Autoregressive Image Generation
RePlan: Reasoning-Guided Region Planning for Complex Instruction-Based Image Editing
UNet-Twice: A Simple Structured Reference-based Inpainting Framework
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
Text-Conditioned Background Generation for Editable Multi-Layer Documents
Multi-dimensional Preference Alignment by Conditioning Reward Itself
Follow-Your-Mind: Towards Inversion-Free Brain-Driven Visual Context Synthesis and Editing
Histocomponent-driven Universal Model for Virtual Immunohistochemistry Multiplex Staining via Joint Manifold Evolution
Learning to Stylize by Learning to Destylize: A Scalable Paradigm for Supervised Style Transfer
VSDiffusion: Taming Ill-Posed Shadow Generation via Visibility-Constrained Diffusion
Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
NoisEasier: Test-Time Noise Optimization for Text-to-Video Generation
Training-Free Multi-Concept Image Editing
DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
Q-REAL: Towards Naturalness and Distortion Evaluation for AI-Generated Content
RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models
From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation
NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
SAEdit: Token-Level Control for Continuous Image Editing via Sparse Autoencoder
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models
MPO: Single-Stream Policy Optimization for Efficient Text-to-Image Alignment
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
Revisiting Autoregressive Models for Generative Image Classification
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
FitControler: Toward Fit-Aware Virtual Try-On
A²-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks
LoGAN: Multilingual Font Localization with Generative Agents
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
ELDiff: When Evidential Learning Meets Text-to-Image Diffusion
Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise
Layering Virtual Try-On
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
Policy-Based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
Spanning Tree Autoregressive Visual Generation
POET: Preference Optimization for Enhanced Text-to-Image Generation
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
TOPA: Mitigating Concept Dominance in Diffusion Personalization via Target-Oriented Perturbation Augmentation
Reward Lightning: Fast Video Generation via Homologous Preference Distillation
Reinforcement Learning for Multimodal Diffusion Language Models via Bidimensional Trajectory and Thought Optimization
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
Personalized Reward Modeling for Text-to-Image Generation
From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting
EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
LayerVerse: Finding the Sweet Spot for KV-Injection in Training-Free Image Editing
TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive
Achieving Subcategorical Erasure in Text-to-Image Models
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
FairSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
Unsafe by Reciprocity: How Generation–Understanding Coupling Undermines Safety in Unified Multimodal Models
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
Target-aware Image Editing via Cycle-consistent Constraints
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
SpecV: Specification Verification for Robust Unified Multimodal Evaluation
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
CGCE: Classifier-Guided Concept Erasure in Generative Models
Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation
Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment
FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
Reinforcing Video Reasoning with Focused Thinking
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
UniReflect: Self-Reflection Tuning for Unified Multimodal Understanding and Generation
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
Correlation-Weighted Multi-Reward Optimization for Compositional Generation
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
On-Policy Diffusion Reinforcement Learning Meets Off-Policy Quality Anchoring
Unlocking Complex Image Editing via Natively Interleaved Visual Textual CoT with Deep Confidence Reasoning
OSVE: One Step Video Editing with One Step Diffusion Models
Benchmarking MLLMs on Mistake Recognition and Explanation in Single-Step Components of Cooking
DisRM: Reward Modeling as Discriminative Prediction
InstaEdit: Instant Image Editing via Optimized Noise Prediction
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
PhysEdit: Physically Consistent Image Editing via Causal Enforcement
MVI2V: Human Centric Image to Video Generation with Multiview Consistent Appearance
GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
Zero-Shot Image Personalization from Personas
Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models
The Path to Reconciling Quality and Safety Alignment in Text-to-Image Generation
AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution
Optimization-Guided Diffusion for Interactive Scene Generation
h-Flow: Flexible Flow-based Image Editing via Doob's h-Transform
Reference-Free Quality Assessment for Virtual Try-On via Human Feedback
CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning
NearID: Identity Representation Learning via Near-identity Distractors
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment
Push–Pull Attentional Anchoring for Diffusion Concept Erasure
V-HOLD: Stabilizing Flow Trajectories to Rethink the Edit–Preservation Trade-off
Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs
CHARTSTYLE-100K: A Large-Scale Dataset for Structured Visualization Style Transfer
Consistent Feature Transport for Image Relighting
SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
Multi-History-Step SDE Inversion for Image Editing with Superior Regional Awareness
Self-Improving Diffusion Classifiers with Minority Preference Optimization
MirrorPPR: Exemplar-Based Portrait Photo Retouching
When Rubrics Fail: Error Enumeration as Reward for Reference-Free RL Post-Training
Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
SupGRPO: Enhancing GRPO with Matching-based Online SFT for Text Spotting
i-Design: Step-by-Step Graphic Layout Design with Progressive Aesthetic Policy Optimization
Continuous Speculative Decoding for Autoregressive Image Generation
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
Editing Everything Everywhere All at Once
Learning Consistency in Reward Modeling for Multi-Modal Reasoning
SR-Edit: Region-Aware Image Editing via Self-Refinement
Wan-R1: Verifiable-Reinforcement Learning for Generalizable Video Reasoning
InstantRetouch: Personalized Image Retouching without Test-time Fine-tuning
Semantically Aligned Gradient-Driven Context-Preserving Image Editing
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing
Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples
Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
VLTR: Vision-Language Tool Reasoning for Instruction-Guided Image Editing
To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
The Map Is Not the Territory: Embedding-Coverage Blacklists for Safe Diffusion Steering
M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection
InclusiveHuman-10K: Towards Inclusive Human Parsing Beyond the Intact-Limb Assumption
EDM:Event-guided Diffusion Model for Video Shadow Detection in Complex Dynamic Scenes
Streaming Dense Voxel Representations for 3D Occupancy Prediction
Learning Probabilistic Embeddings for Unsupervised Action Segmentation
Region-Aware Multimodal Interleaving for Animal Re-Identification
Unified Removal of Raindrops and Reflections: A New Benchmark and A Novel Pipeline
Silhouette-based Gait Foundation Model
Free‑CD: Probabilistically Decoupled Training-Free Open-Vocabulary Change Detection with Resolution-Invariant Feature Inversion
STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection
Rethinking Continual Anomaly Detection on the Edge: Benchmarking Under Realistic Industrial Conditions
FoundYou: A Unified Model for Personalized Segmentation and Retrieval
Back-Tracking from Clarity: Self-Learning to See Text from Afar
RePL: Pseudo-label Refinement for Semi-supervised LiDAR Semantic Segmentation
One Slide, Many Views: Unifying Complementary Foundation Model Perspectives for WSI Analysis
High-Throughput Event-Based Feature Detection and Tracking on an Embedded CPU
LogiCo: A Unified Framework for Logical and Structural Anomaly Detection
MiNQVIS: Mitigating Noisy Queries for Robust Online Video Instance Segmentation
CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking
Granular Semantic Cognition for Visible-Infrared Person Re-Identification
Bottom-up modeling of repeated elements via single image analysis-by-synthesis
Uncertainty-aware tree height change regression
Hierarchical Hyperbolic Representation Learning for Aerial-Ground Person Re-Identification
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
Beyond Attention: Convolutional Global Context for Remote Sensing Change Detection
NUN: Nested Unfolding Network for Real-World Concealed Object Segmentation
Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations
ELHINN: Unifying Dense Crowd Simulation Across Scales via Eulerian–Lagrangian Hydrodynamics
QVAM: Query-guided View-aware Adaptive Modulation for Aerial-Ground Person Re-Identification
YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Reasoning
CountEx: Fine-Grained Counting via Exemplars and Exclusion
GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis
InstaPano: Zero-shot Instance Layout Controlled Panorama Generation Via Global Attention Fusion
Sound-based Multi-Person 3D Pose Estimation
CerDETR: Cell-Prior Empowered DETR for Cervical Lesion Detection
Making Partial-Label Datasets Easier: A Simple Yet Highly Effective Data Augmentation for Deep Partial-Label Learning
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking
IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation
EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation
Seek to Segment: Active Perception for Panoramic Referring Segmentation
DeCo: Zero-Shot Anomaly Generation through Decoupling and Recoupling
UniScale: Arbitrary-Scale Anomaly Generation
Environmental Change Detection for Real-World Change Analysis
PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images
SplitHDR: Saturation-Aware HDR Recovery and Denoising for Real-Time Detection
DeltaDeno: Zero-Shot Anomaly Generation via Delta-Denoising Attribution
History-Aware Transformation of ReID Features for Multiple Object Tracking
Beyond Aesthetics: Quantifying Information Loss in Turbid Scenes
Rectified Embedding Flow Learning for Aerial Multi-view Geo-localization
P3-SAM: Native 3D Part Segmentation
Per‑Object IoU Forecasting for Deadline‑Aware Real‑Time Embedded Detection Control
RiO-DETR: DETR for Real-time Oriented Object Detection
ZMIS-SAM: Segment Anything Model Enhanced With Wavelet Transform For Zooplankton Microscopy Image Instance Segmentation
TiCRL: Textual Image Classification with Reinforcement Learning-Based Curriculum Learning
Slim-DETR: Real-Time Tiny Object Detection with Efficient Interaction and Gaussian Query
Decoupling Moment from Event for Video Temporal Grounding
Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels
DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection
CLUE-VAD: Structured Semantic Clues for Understanding Explainable Events in Video Anomaly Detection
ODONet: Online Dynamic Offset Network for Visual Object Tracking
HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking
UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset
Evidence Triangulation for Multimodal Fact-Checking in the Wild
GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks
Constrained Rotation Optimization: Revisiting Crop-Based Gaze Estimation
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
DualCount: Structurally Consistent Density and Point Modeling for Zero-Shot Object Counting
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection
Real-Time Source-Free Object Detection
DETRAM: End-to-end DEtection, Tracking and Recovery of HumAn Meshes
Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection
Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision
SAGE: A Synchronized Action and Gaze Estimation Framework for Comprehensive Human Behavior Analysis
FuDU: A Fuzzy Dual-dimension Uncertainty Framework for Streaming Active Learning in Industrial Defect Detection
Instance Segmentation as Tracking: A New Paradigm for Multi-Small-Object Tracking with Event Cameras
Online 3D Instance Segmentation at task-oriented granularity with Unposed Monocular Video
ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection
Mode-Conditioned Residual Calibration for Multi-Object Tracking
HLRAD: High-dimensional Latent Representation for Unified Anomaly Detection
Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation
WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation
CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning
SARIF: Segment Anything for Robust Image Forensics
On-Orbit Real-Time Wildfire Detection Under On-Board Constraints
RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild
Pseudo-Stereo Inputs: A Solution to the Occlusion Challenge in Self-Supervised Stereo Matching
ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation
RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking
Robust Zero-shot Anomaly Detection under Limited Auxiliary Anomaly Priors
Reliability-Aware 3D Geometric Injection for Universal Person Re-identification
BAAF: Universal Transformation of One-Class Classifiers for Unsupervised Image Anomaly Detection
ModTrack: Sensor-Agnostic Multi-View Tracking via Identity-Informed PHD Filtering with Covariance Propagation
Rapidly Deploying On-Device Eye Tracking by Distilling Visual Foundation Models
CRD-Net: Frequency-Adaptive Feature Injection and Change Decoupling for Building Damage Assessment
Don’t Starve the Boundaries: Boundary-Constrained Label Propagation for Weakly Supervised 3D Segmentation
Motion-aware Sparse Pipeline for Lightweight Object Tracking
CGCC: Towards Generalizable Clothes-Changing Person Re-Identification
Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
Cross-Species Animal Re-Identification with Semantic Consistency Learning
Following the Flow: Advection-Consistent Modeling for Event-based Small Object Detection
EVKit: An Open-source Flexible Toolkit for Efficient Event Camera Data Storage and Loading
OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations
IACD: Iterative Adversarial Collaborative Detection via Dual-Perspective Blind Spot Discovery
VIGA: View-Conditioned and Identity-Guided Adaptation for Aerial-Ground Person Re-Identification
Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers
ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs
Segmentation-Guided Homography Estimation for Long-Term Planar Tracking
DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints
Diffusion-based dual-view reflection removal
ParaFlow: Parallel Sampling for Flow Matching Models
Direct Autoregressive Diffusion Distillation via Error-aware Causal Pretraining
Training-Free Refinement of Flow Matching with Divergence-based Sampling
LACON: Training Text-to-Image Model from Uncurated Data
Discrete Diffusion Bridges for Spatiotemporally Aligned Image Translation and Generation
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?
AViTS:Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
Expert Weaving: Marrying Masked AutoRegressive and Diffusion Models for Unified Image Restoration
Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention
Prompt2Effect: Training-Free LoRA Synthesis for Controllable Video Effects
Allo{SR}2: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows
NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control
UltraGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
ARVAR: Accelerating Visual Autoregressive Model via Attention Retrospect
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators
HNDiff: Haze-Noise Diffusion for Image Dehazing
Difficulty-Conditioned Attribute-Specific Restoration for Low-Light Image Enhancement
Learning to Corrupt for Better Restoration
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
Physics Meets Perception: A Reinforcement Learning Framework for Unpaired Real-World Image Dehazing
RefAlign: Representation Alignment for Reference-to-Video Generation
Glance: Accelerating Diffusion Models with 1 Sample
Towards Scalable Pre-training of Visual Tokenizers for Generation
Spectral Consistent Flow for One-step 3D Medical Image Translation
Adversarial Score Distillation for Stable One-Step Diffusion in Real-World Image Super-Resolution
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
From Noise to Events: Conditional Diffusion for Event Data Augmentation
Region-Aware Test-Time Scaling for Compositional Image Generation
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers
Histogram-constrained Image Generation
Parsimonious Flow Matching for Efficient Image Generation
AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
SkelEM: Explicit Decoupling of Topology and Details for Self-supervised Axial Super-Resolution in Volume Microscopy
UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion
Recolour What Matters: Region-Aware Colour Editing via Token-Level Diffusion
Semantic Browsing: Controllable Diversity for Image Generation
InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
Posterior Augmented Flow Matching
Contrastive Conditional–Unconditional Alignment for Long-tailed Diffusion Model
D3F-IR: Dual-Domain Deterministic Flow Matching for Visible-to-Infrared Translation
GLARE: Towards Generalizable Detection of Latent Diffusion Images with Global-Local Reconstruction Error
ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching
Be Tangential to Manifold: Discovering Riemannian Metric for Diffusion Models
ART-VSR: Adaptive Rectified Trajectories for One-Step Video Super-Resolution
AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation
Zero-Shot Inference-Time Rectification for Real-World Arbitrary-Scale Super-Resolution
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
Representation Alignment for Just Image Transformers is not Easier than You Think
G3AFT: Glance Guided Gradient Aligned Fine-Tuning for Visual Autoregressive Models
Hi-DiT: Hybrid Latent-Pixel Diffusion Transformer for Image Generation
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
Jumping the Landing Phase: Noise Variance Matching Enables Accurate Few-Step Inversion
AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer
FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal
MaterialFlow: Attribute-Disentangled Material Transfer via Trajectory-Aware Velocity Modulation
FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution
Diffusion Image Generation with Explicitly Modeling of Data Manifold Geometry
Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation
Introspective Attention Modulation for Safe Text-to-Image Generation
Continuous Adversarial Flow Models
ReAL: Reference-to-Image (R2I) Aware Latent Diffusion for Image Super-Resolution
Stochastic Optimal Control Sampling for Diffusion Inverse Problems
Fair and Faithful: A Diffusion-Enhanced Dataset and Hybrid State-Space Mamba for Face Super-Resolution
High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control
Image Warping for Image-to-Image Translation
ConceptWeaver: Weaving Disentangled Concepts with Flow
TEASR: Training-Efficient Any-Step Diffusion Transformer for Real-World Image Super-Resolution
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
Spectral Prior for Reducing Exposure Bias in Diffusion Models
Unified Multi-plane Autoregressive Diffusion for 3D Multi-Contrast MRI Synthesis
Y-diff: Structure-Texture Decoupled Diffusion Distillation for H&E-to-pCLE Translation
DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent
Dual-End Consistency Model
GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors
Controllable Generative Reference for Stereo Image Compression via Reliability-Aware Gating
Cross-Resolution Distribution Matching for Diffusion Distillation
Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration
FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring
Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
GMODiff: One-Step Gain Map Refinement with Diffusion Priors for Efficient HDR Reconstruction
BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
Co-evolving Representations in Joint Image-Feature Diffusion
Extreme Face Super-Resolution through Identity Fitting and Decoupling
Momentum Guidance: Plug-and-Play Guidance for Flow Models
Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models
Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
MeanTalker: Efficient and Expressive Speech-Driven 3D Facial Animation via Geometric-Aware Mean Flow
Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution
Short-to-Long Functional Connectivity Transfer via Structure-Aware Latent Diffusion
RefDiT: Local Attribute Guidance in Reference-Based Image Generation
Improving Image-to-Image Translation via a Rectified Flow Reformulation
DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution
Rdm: Re-conceptualizing Distribution Matching as a Reward for Diffusion Distillation
ELT: Elastic Looped Transformers for Visual Generation
AdaBridge-SR: Adaptive Bridge Matching for Real-World Image Super-Resolution
Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness
Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories
TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization
Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis
DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models
LUA: Latent Upscaling Adapter for Diffusion-Based Image Synthesis
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
Steering Diffusion Models via Class-Contrastive Influence for Few-Shot Classification
Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
Generative Manifold Distillation: Aligning Restoration Trajectories with the Natural Image Prior
Wavelet-Driven Cross-Domain Consistency for Mixed-Supervised 3D Tumor Segmentation
Thermo-JEPA: Learning a Geometry-Grounded Thermal World Model via Cross-Modal Privileged Masking
Adaptive Spectrum-Aware Feature Disentangled Network for Small Object Detection
ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
OpenPanoD: Aligning Multimodal Prompts and Spherical Representations for Open-Vocabulary Panoramic Detection
Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction
LGD-Net: Leader-Guided Cross-Modal Dynamics for Hyperspectral and Panchromatic Image Fusion
RT-SDGOD: Real-Time Single-Domain Generalized Object Detection
A Dual-space Patch-driven Complementary Learning Framework for Semi-supervised Multi-organ Segmentation
Selective Synergistic Learning for Video Object-Centric Learning
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
Disentangling and Reusing Interaction Cues for Zero-Shot HOI Detection
BEVOpen3D: Towards Open-World 3D Object Detection in Bird's-Eye-View
SGP2: Coarse-to-Fine Controllable Multimodal Remote Sensing Image Generation
Degradation-Agnostic Clarity Learning for Unpaired Image Dehazing
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
ReMoMask: Retrieval-Augmented Masked Motion Generation
Decoding Multimodal Causality: End-to-End Multimodal Mediation Pathways Inference
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
VarProtoAD: Variational Prototype-Conditioned Prompting for Zero-Shot Anomaly Detection
What Images Cannot Say: Language-Guided Olfactory Representation Learning
Distribution-Aware Feature Selection for Post-hoc Out-of-Distribution Detection
DAP: Doppler-aware Point Network for Heterogeneous mmWave Action Recognition
FSD-Net: Foundation-Guided Spatiotemporal Distillation for Video Polyp Segmentation
From Visual Primitives to Semantic Masks: Fine-Grained Visual-Linguistic Alignment for Open-Vocabulary Remote Sensing Image Segmentation
PhysFlowNet: Learning Canonical Latent Manifolds via Spatio-Spectral Physics Priors for Underwater Object Detection
DA-F2F: Domain-Adaptive Object Detection with Feature-to-Feature Modulation and Alignment
Intra-Class Consistency Guided Class-Agnostic Event Segmentation
Towards Unsupervised Multi-modal Semantic Segmentation
XSemanticFlow: Cross Object Semantic Alignment for Zero-shot Manipulation
Progressively Spiral Mamba Fusion for Multimodal Tracking
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Safe Generalization: Mitigating Catastrophic Forgetting in Single-Source Multi-Organ Segmentation via Collaborative Causal Learning
DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning
Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time
CHIMERA: Adaptive Cache Injection and Semantic Anchor Prompting for Zero-shot Image Morphing with Morphing-oriented Metrics
When Specialists Meet Generalists: Segmenter-Coordinated Asymmetric Learning for Label-Deficient Concealed Object Segmentation
Enhancing prompt-image alignment evaluations via cyclic mutual information maximization
InfraNet: Quality-Aware RGB Guidance for Infrared Object Detection
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
Hierarchical Style Aggregation for Versatile Chinese Handwriting Generation
Fragmented Text Is Insufficient for Image Representation: Fine-Grained Correspondence in Multimodal Dataset Distillation
RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation
NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models
MoMCE: Mixture of Modality and Cue Experts for Multimodal Deception Detection
When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization
VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
Maximum Spanning Tree Guided Confidence and Sparse Graph for Robust Noisy Label Learning
Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges
Preserving Knowledge across Space and Time for Continual Video Deepfake Detection
ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation
Amplify, Aggregate, and Adjust: VideoMAE-based Holistic-Subtle Aggregation for Micro-Action Recognition
Local-to-global Cross-modal Coordination for Self-supervised RGB-T Tracking
Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation
Beyond Alignment: A Generative Matching Paradigm via Flow Matching for Zero-Shot Skeleton-Based Action Recognition
Saber: Anchoring Semantics to Scale-Aware Kinetic Salience for Zero-Shot Skeleton Action Recognition
Path-JEPA: Path Signature Based Predictive Learning for Skeleton Action Recognition
OP3DSG: Open-vocabulary Part-aware 3D Scene Graph Generation for Real-world Environments
Mitigate Modality-Asymmetric Forgetting via Stabilizing Visual Representations in CLIP-Based Class-Incremental Learning
Towards Sparsely Annotated Open World Object Detection
MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker
Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension
Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition
Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts
Mask-guided Semantic Alignment: Robust Learning with Noisy Labels via Temporal Attention Stability
Rethinking Cross-Spectral Image Generation via Shared-Specific Representation
Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small-Object Detection
MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts
ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving
Warp-free Cross-view Geo-localization via Feature-space Consensus Mining
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
Rethinking Pseudo-Labels: Multi-Granularity Supervision for Domain Adaptive Object Detection
Hierarchical Prompt Injector for Domain Generalization Segmentation
Group3D: MLLM-Guided Semantic Grouping for Open-Vocabulary 3D Object Detection
Context-Aware Joint Alignment for Cross-Scene Hyperspectral Image Classification
AquaStereo: Enabling Underwater Stereo Matching via Depth-Conditioned Diffusion and Geometry Self-Distillation
HVGCD:Rethinking Generalized Category Discovery through Hypothesis–Verification
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
Liquid Fusion of Heterogeneous Representations Towards General Salient Object Detection
DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation
Beyond Categorical Matching: Intra-Class Graded Relevance Estimation for Cross-Modal 3D Retrieval
SOVTrack: Open-Vocabulary Multi-Object Tracking with Self-Supervised Pseudo Labeling and Feature Distillation
SARA: Structure-Aware Riemannian-Guided Alignment for Drone Image-Text Retrieval
What is the Right Embedding Space for Contrastive Learning in Referring Expression Counting?
Iterative Refinement of Semantic and Spatial Representations for Open-Vocabulary Camouflaged Object Segmentation
HER-Count: Learning Hyper-Exemplar Representation for Generalized Zero-Shot Object Counting
FST-SAM3: Taming SAM~3 with Frequency-Spatio-Temporal Refinement for Video Polyp Segmentation
Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection
Unleashing the Power of Large-Scale ViT in Zero-Shot SBIR: A Strong Baseline with Multi-Layer Feature Aggregation
InSeg: Interactive Refinement via Intent Propagation for Point Cloud Semantic Segmentation
CaPCL: Caption-Preserved Continual Learning for Text-to-Image Retrieval
Label-Free Text Prototype Adaptation for Open Vocabulary Segmentation
SFKD: Spatial–Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction
Toward Robust In-Context Segmentation via Concept Guidance
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
Open-Weather Robust 3D Detection via Dual-Critic Diffusion Alignment
Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation
Degradation-Robust and Temporally Consistent Infrared–Visible Video Fusion via One-step Diffusion Framework
CascadeProto: Cascaded Cross-Modal Prototype Purification via Entropy-Aware Learning for Few-Shot 3D Point Cloud Segmentation
Proposal Score Realignment Guided by Semantic Completeness for Weakly Supervised Temporal Action Localization
IP-SAM: Rethinking Prompt-Conditioned Segmentation for Prompt-Absent Deployment
DETR is Secretly a Multispectral Detector: Zero-Parameter Adaptation via Semantic Alignment
NegROI: Click-Centric Uncertainty-Guided Refinement with Scene-Conditioned Negative Prompts for Robust Interactive 3D Segmentation
DE2TR: Dual Evidence Detection Transformer for Video Temporal Grounding
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
P²Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
PASR: Pattern-Aware Scene-Conditioned Reasoning for Camouflaged Object Detection
μFlow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors
ORACLE-3D: Open-world Region-aligned Cross-modal Learning for Label-efficient 3D Scene Understanding
Mitigating Pose–Scale Discrepancy Bias and Reforming Multi-Support Reasoning for Few-Shot Semantic Segmentation
Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching
Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation
ReynoldsFlow: Physics-Inspired Spatiotemporal Flow Representation for Video Understanding
CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Image Segmentation
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch‑Level Incoherence and Temporal Roughness
Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
Breaking the Model Forgetting Cycle in Long-Incremental 3D Object Detection
TerrainGraphNet: Terrain-Constrained Graph Reasoning for Landslide Segmentation
Fourier Self-Supervision for Fine-Grained Generalized Category Discovery
Calibrate Before Adapt: Training-Free Pseudo-Label Calibration for Semi-Supervised Cross-Domain Few-Shot Detection
Capturing Spectral and Spatial Patterns for Federated Remote Sensing Segmentation
FAIR: Feature-Augmented Implicit Regularization for AI-generated Fake Image Detection
Learning to Attract and Repel: Dual Quality Margin Learning for Face Recognition (DQM-Face)
From Local Geometry to Global Pseudo-Labeling for Robust Positive–Unlabeled Learning under Covariate Shift
OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification
Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs
Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning
HEM: a margin-based loss for visual categorisation tasks
Beyond Disjoint Tasks: Towards More Natural Continual Learning for Vision-Language Models
The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning
FaceMoE: Mixture of Experts for Low-Resolution Face Recognition
Dual-Generalization-aware Minimization for Continual Fine-Tuning of Vision-Language Models
Graph Coloring for Multi-Task Learning
S2-FracMix: Self-Saliency Fractal Mixup
SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting
Shared LoRA Subspaces for almost Strict Continual Learning
Improving Adversarial Robustness via Activation Amplification and Attenuation
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
Learning to Recover Task Experts from a Multi-Task Merged Model
Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
Residual-Guided Expert Specialization for Incomplete Multimodal Learning
On the Vulnerability of Parameter-Level Defenses to Model Merging
Unified Multi-Layer Subspace Modeling for Cross-Domain OOD Detection
GCMRD: Global Consistency Multi-teacher Robustness Distillation
Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation
Benchmarking Federated Learning & Knowledge Distillation for Point Cloud Classification
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection
VLMSysTrojan: Stealthy System-Aware Backdoor Attacks Against Vision-Language Models
Holistic Optimal Label Selection for Robust Prompt Learning under Partial Labels
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Dive into the implicit biases of low-rank vision-language alignment
Neural Collapse-Inspired Multi-Label Federated Learning under Label-Distribution Skew
MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation
CollectionLoRA: Collecting 50 Effects in 1 LoRA for Deployment
FedMental: Topology-Aware Federated Prototype Learning for Polymorphic Multimodal Psychiatry
Prevention over Correction: Learning Aligned Representations in One-shot Federated Learning
TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction
MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
Training-Free Task Classification for Multi-Task Model Merging
Rank-Aware Hyperbolic Alignment for Vision–Language Dataset Distillation
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation
SUM: Unified Geometric Surgery on Spatio-Temporal Adaptation Vectors for Federated Class Incremental Learning
SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation
Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning
SCoT: Similarity-guided Conflict-aware Task Consolidation for Continual VQA
DA-MergeLoRA: Hypernetwork-Based LoRA Merging for Few-Shot Test-Time Domain Adaptation
Importance-Aware Low-Rank Distillation of Diffusion Transformers
MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation
BATQuant: Outlier-Resilient MXFP4 Quantization via Learnable Block-wise Optimization
Molecular Identifier Visual Prompting and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing
Prototype Normalization: Optimizing Prototype Separation for Heterogeneous Federated Learning
ReTarget: Representation Transformation via Adversarial Regularization for Geometric Misalignment
Flash-DD: An Ultra Parameter-Efficient Approach to Dataset Distillation
DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images
TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
TouchAnything: Diffusion-Guided 3D Reconstruction from Sparse Robot Touches
Seen2Scene: Completing Realistic 3D Scenes with Visibility-Guided Flow
WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation
SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction
FILT3R: Latent State Adaptive Kalman Filter for Streaming 3D Reconstruction
GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training
White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing
Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
UniQueR: Unified Query-based Feedforward 3D Reconstruction
SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation
LESV:Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
SHReg: Strictly Rotation-Equivariant Point Cloud Registration via Spherical Harmonics
Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs
Map2World: Segment Map Conditioned Text to 3D World Generation
Towards Practical Lossless Neural Compression for LiDAR Point Clouds
BitRIC: Efficient Neural Compression of LiDAR Range Images via Hierarchical Bitplanes
Uncertainty-Driven Gaussian Sphere Propagation for 3D Semantic Segmentation
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
Zero-Shot Novel Depth Synthesis Using Foundation Models Scene Representations
SLAM-Former: Putting SLAM into One Transformer
Pixel-wise Geo-registration of Drone and Satellite Images
G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation
MessyKitchens: Contact-rich object-level 3D scene reconstruction
MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
Occlusion-Resilient Category-Agnostic Pose Estimation with Conditional Flow Matching
Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding
SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for RGB+Thermal 3D Reconstruction
Sparse-Aware Vector Quantization for Bandwidth-Efficient Collaborative 3D Semantic Occupancy Prediction
Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity
Bootstrapping Articulated 3D Reconstruction from 2D Image Collections
DRS-VPT: Directly Re-localizing in Scenes using a Vision and Point Transformer
DisentangledTMR: Privacy-Preserving Skeleton Motion Retargeting via Factorized Transformers
Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
Transferability Between Understanding and Generation in Unified Multimodal Models
Unfold The World: Factorize 4D Properties in Reinforcing Spatial Understanding
What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA
Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
Towards Interactive Global Geolocation Assistant
SWIFT: Spatial-Window Integrated Frequency-aware Token Pruning for Efficient MLLMs on Edge Devices
PhysAlign: Learning Physical Priors for Dynamical Event-Driven Video Generation via Representation Alignment
DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception–Memory Integration in Embodied Environments
FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity
AMCI: Unlock the Potential of Large Multimodal Models for Fine-grained Open-world Classification via Adaptive Memory Context Injection
Multiple Images Distract Large Multimodal Models via Attention Fragmentation
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
RAGrasp: A Retrieval-Augmented Framework with Diversity-Aware Modeling for Dexterous Grasp Generation
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning
Spectral Evolution-Guided Token Pruning in Large Multimodal Models
HERO: Enhancing Multimodal Faithfulness via Dynamic Entropy-Aware Reinforcement Learning
Video-Text Alignment Model for Sign Language Translation
Evidence-Backed Video Question Answering
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
Enhancing Interpretability in CLIP with Optimal Transport-based Submodular Optimization for Ophthalmic Imaging
Learning Active Perception for Pixel-Space Reasoning via Visual-Intent Stratified GRPO
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
Counting Trees from Satellite Imagery with Noisy Supervision
Toward Interpretable Analysis of Whole-slide Pathology Images via Large Language Model-based Agentic Reasoning
ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering
GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision
PanoRec: Spatially-Structured Sequence Modeling for Multi-Granularity Panoramic Retrieval
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
HippoCamp: Benchmarking Contextual Agents on Personal Computers
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
Staying VIGILant: Mitigating Visual Laziness in MLLMs via Information-Theoretic Alignment
Co-Steer: Cross-Modal Collaborative Steering for Jailbreaking MLLMs
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
Knowledge-Centric Agents for Workflow Generation in ComfyUI
Incentivizing Vision Language Models to Search for Long Video Question Answering
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Learning Sample-wise Rank-Aware Interpolation Weights for Composed Visual Data Retrieval
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
Mixture of Specialized Vision Experts: Unlocking Complementary Visual Insights for Faithful MLLM Reasoning
Experts-Guided Unbalanced Optimal Transport for ISP Learning from Unpaired and/or Paired Data
AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Large Language Models and Reinforcement Learning
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus
Visual Spatial Tuning
HumanOmni-Speaker: Identifying Who said What and When
BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
Unbalanced Optimal Transport for Efficient Visual Document Retrieval
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
DocLayout-VL: A Foundational Model for Hierarchical, Open-set, and Promptable Document Layout Segmentation
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training
VisCritic: Visual State Comparison as Process Reward for GUI Agents
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
EatVid-Bench: A Multimodal Fine-Grained Eating Behavior Video Dataset
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
Multi-label Instance-level Generalised Visual Grounding in Agriculture
SEERBench: A Spatial Ego-Exo Reasoning Benchmark for MLLMs with a Simple Yet Effective Baseline
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
See & Sniff: Learning Visuo-Olfactory Representations
Personalizing MLLMs via Reinforced Multimodal Reference Game
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
CMDR: Contextual Multimodal Document Retrieval
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
HighlightBench: Benchmarking and Diagnosing Markup-Driven Table Reasoning in Scientific Documents
Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
From Illusion to Intention: Visual Rationale Learning for Reliable Evidence Acquisition
Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray
InstrAct: Towards Action-Centric Understanding in Instructional Videos
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
NAPA: Natively Multimodal Autoregressive Perception Architecture
Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT–Quantization Embedding
Information-Regularized Attention for Visual-Centric Reasoning
Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
Mitigating Sycophancy in Multimodal Chart Understanding via Vision-Grounded Verification
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs
The Telephone Game: Evaluating Semantic Drift in Unified Models
Wavelet-based Intra-video Counterfactual Reasoning for Video Question Grounding
How Far Are Video Models from True Multimodal Reasoning?
SE-DETR: Explicit Semantic Exploration for Generalizability and Distinguishability in Video Temporal Grounding
MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
How to Teach Large Multimodal Models New Skills
Towards High-Resolution Visual Perception via Hierarchical Entity Exploration
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
Iterative Perceptual Alignment for VLMs via Deterministic Reconstruction Feedback
Caption Bottleneck Models
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
CogniCred: A Dataset and Benchmark for Cognitive Credential Forgery Detection
Error-Driven Scene Editing for 3D Grounding in Large Language Models
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Event-Driven Video Generation
Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
Geo-DPO: Aligning Semantic Intent with Geometry for 3D Affordance Segmentation
ViTAL‑X: Video-Text Alignment with Cross‑Modal Temporal Edits
Learning Spectral and Polarimetric Clues for One-to-Multimodal Novel View Synthesis
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
ChronoFlow Policy: Unifying Past-Future Interaction Flow in Visuomotor Policy Learning
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
RoMan-4D: Learning Robot Arm Manipulation from 4D World Models
OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction
Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving
Panoramic Affordance Prediction
CabinSI: Omni-Cabin Spatial Reasoning through Explicit Visual Cognitive Maps
Towards Generalizable Robotic Manipulation in Dynamic Environments
EffiDINO: Task-Specific Model Pruning via Gram Anchoring Subspace Consistency
Spatial Amsan: A Benchmark for Perception-Grounded Spatial Reasoning and Action Evaluation in Egocentric Manipulation
DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
ASSCG: Just-Right Gating over Chattering for Fast–Slow LLM Planning in Autonomous Driving
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
ESTANet: Efficient Online Error Detection in Procedural Videos via Prediction Inconsistency
A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
ORION: Ordinal Neural Collapse as a Representation Prior for Visual Navigation
Grasp-Oriented Non-Prehensile Manipulation via Learning a Graspability Field
Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving
Learning Transferable Dynamics Priors from Action to World Modeling
Unordered Landmark Visual Navigation
AgentVLN: Towards Agentic Vision-and-Language Navigation
CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models
Walk through Paintings : Ego-centric World models from Internet Priors
TEX-Drive: Temporal Perception Meets Experience-Guided Mixture-of-Experts for End-to-End Autonomous Driving
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception
PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving
Doe-2: 3D Representation World Model for Unified Driving Scene Forecasting
PersonaDrive: Controllable Trajectory Prediction with Multi-Dimensional Driving Personas
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation
CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement
RelAfford6D: Relational 6D Affordance Graphs for Constraint-Driven Robotic Manipulation
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving
When the City Teaches the Car: Label-Free 3D Perception from Infrastructure
UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection
ZTRS: Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer
Driving like yourself: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather
SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection
UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving
StructPolicy: Structure-Guided Imitation Learning Robust to Visual Domain Shifts
RadarGen: Automotive Radar Point Cloud Generation from Cameras
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
World-in-Loop: Online Correction via Event-Triggered World Models for Robust VLA Policies
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving
C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
NeSy-Route: A Neural-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
Guide, Think, Act: Interactive Embodied Reasoning for Vision-Language-Action Model
Noise is a Good Teacher: A Noise-Driven Framework for Robust Collaborative Perception
Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Reconstruction
TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning
Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D
RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception
CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation
DiverseAD: A Large-Scale Driving Dataset with Diverse Atmospheric Conditions
Markov-Renewal Single-Photon LiDAR Simulator
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
What if? Emulative Simulation with World Models for Situated Reasoning
TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction
SIMON: SImultaneous Multi-Object Navigation
Reflection-Aware Reasoning for Non-Line-of-Sight Pedestrian Localization
3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints
TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
VVSim: A Large-Scale Aerial-Ground Dataset and Benchmark for Cooperative Perception
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
LaGen: Towards Autoregressive LiDAR Scene Generation
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Two-Way Street: Efficient VSLAM using Collaborative In-Sensor and Off-Sensor processing
ContextFlow: In-Context Flow Matching for Robot Manipulation
VOCA: Visual Odometry with Codec Awareness
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
Agent-OBJ: Prompt-Driven 3D Adversaries for Multi-Modal Perception
RAE-NWM: Navigation World Model in Dense Visual Representation Space
WildCity: A Real-World Dataset for City-Scale Rendering and Beyond
GameWorlds: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection
Self-Evolving Just-In-Time Memory for Proactive Embodied Safety
Drive2Danger: Deceive End-to-End Autonomous Driving with Risky Instance Recognition
Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-shot Real-World Deployment
Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
AMCoNav: Asynchronous Multi-module Collaborative Framework for Embodied Visual Navigation
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management
HitMem: Hierarchical Temporal 3D Memory with Multi-Modal Context-Aware Retrieval for Dynamic Environments
HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning
Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models
The Language of Visual Attention: Modeling Scanpaths via Autoregressive Token Prediction
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
EvoWorld: A World-Model-Centric Framework for Continuous Self-Evolution of Modular Embodied Skills
Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
RGBT-GroundBench: Visual Grounding Beyond RGB in Complex Real-World Scenarios
Fast and Scalable LiDAR Data Generation for Autonomous Driving Simulation without Raycasting
SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos
HERO: Heterogeneous Evidential Robust Object-Level Collaborative Perception
RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
Hi-Nav: Hierarchical Framework for Continuous Vision-Language Navigation via Map Guidance and Waypoint Reasoning
UECP: Uncertainty-Enhanced Collaborative Perception
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation
LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
CausalVAE as a Plug-in for World Models: Towards Reliable Counterfactual Dynamics
Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation
Towards Metric-Agnostic Trajectory Forecasting
DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving
SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
AeroVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
ZAP: Zero-Shot Assembly Planning with Large Language Models
RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection
AdaDexGrasp: Adaptive Dexterous Grasping via 3D Visuo-Tactile Representation Fusion
EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
Tactile Modality Fusion for Vision-Language-Action Models
BeyondSight: Object Permanence for End-to-End Autonomous Driving
VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation
WALL-EVE: World Alignment with Rule Learning in Visual Environments
Scale3D: Autoregressive Modeling for Large Outdoor Scene Generation
Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Re-rendering
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming
Generalizable Neural Reconstruction of High-Fidelity Surfaces via Sparse Volumetric Representations
MotionEditGS: Editing Motion and Appearance of 4D Scenes from Monocular Video via Semantically Anchored Gaussians
Diffusion to Obfuscation: Time-Adaptive Synthesized Generation Against Gradient Leakage Attacks in Federated Learning
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Generation
TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis
Show Me Examples: Inferring Visual Concepts from Image Sets
Large-Scale Light Field Synthesis from Videos Enables Geometrically Consistent Bokeh Editing
Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification
P-CORE: Self-Supervised Surface Consistency for Point-Based Neural Editing
Less is More: A Simple yet Effective Object-Centric Prompting Strategy for Vision-Language Reasoning in Autonomous Driving
Beyond Isolated Scans: Cross-Phase Alignment of Structure and Topology for 3D Medical Pretraining
PrintAnything: Learning Geometric Plan Map for 3D Printing G-code Generation from Unoriented Point Clouds
H-SFP: Hierarchical Federated Learning with Decoupled Split-Model Prototyping
Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving
Flash-Refine: Frustum-Guided Local Incremental Learning for Efficient 3D Gaussian Splatting Completion
MLVC: A Multi-platform Learned Video Codec for Real-World Deployment
We use cookies to store which papers have been visited.
I agree
Successful Page Load
ECCV uses cookies for essential functions only. We do not sell your personal information.
Our Privacy Policy »
Accept
We use cookies to store which papers have been visited.
I agree