| Show Detail |
Timezone: Europe/Stockholm
|
Filter Rooms:
TUE 8 SEP
9 a.m.
Workshop:
(ends 1:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 6:00 PM)
Tutorial:
(ends 1:00 PM)
2 p.m.
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Tutorial:
(ends 6:00 PM)
Tutorial:
(ends 6:00 PM)
WED 9 SEP
9 a.m.
Workshop:
(ends 1:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 6:00 PM)
Tutorial:
(ends 1:00 PM)
Tutorial:
(ends 1:00 PM)
Tutorial:
(ends 6:00 PM)
2 p.m.
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
THU 10 SEP
8 a.m.
(ends 9:00 AM)
9 a.m.
Spotlights 9:00-10:15
[9:00]
GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction
[9:05]
Wat3R: Underwater 3D Geometry Learning without Underwater Annotations
[9:10]
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
[9:15]
DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling
[9:20]
Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
[9:25]
PriSplat: Propagating Reliable Multi-view Information for Distractor-Free 3DGS
[9:30]
ReSplat: Learning Recurrent Gaussian Splatting
[9:35]
SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
[9:40]
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
[9:45]
SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization
[9:50]
Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering
[9:55]
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
[10:00]
CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
[10:05]
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
[10:10]
TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid
(ends 10:30 AM)
Orals 9:00-10:30
[9:00]
Provable and Robust Wavefront Sensing via Self-Reference Interferometry
[9:15]
Broadband Wide Field of View Imaging with Computational Mirrors
[9:30]
Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation
[9:45]
A second-order theory of texture for depth from focus
[10:00]
PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment
[10:15]
Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
(ends 10:30 AM)
Spotlights 9:00-10:10
[9:00]
CFM: Language-aligned Concept Foundation Model for Vision
[9:05]
AdaBoosting Text Prompts for Vision-Language Models
[9:10]
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
[9:15]
Molmo-Point: Better Pointing for VLMs with Grounding Tokens
[9:20]
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
[9:25]
Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
[9:30]
Learning to Deny: Action Denial in Multimodal Large Language Models
[9:35]
PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
[9:40]
GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
[9:45]
On Test-Time Scaling for Vision-Language Models
[9:50]
URoPE: Universal Relative Position Embedding across Geometric Spaces
[9:55]
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
[10:00]
EmbedCopilot: Evaluating Vision-Language Models for Hardware-Aware Embedded System Development
[10:05]
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
(ends 10:30 AM)
10:30 a.m.
Posters 10:30-12:30
RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes
SkyLume: A Large-Scale Multi-Illumination Aerial Benchmark for Urban Scene Reconstruction and Beyond
(ends 12:30 PM)
(ends 12:30 PM)
noon
1:30 p.m.
Orals 1:30-2:45
[1:30]
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
[1:45]
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
[2:00]
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
[2:15]
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
[2:30]
FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30]
On the Reliability of Cue Conflict and Beyond
[1:35]
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
[1:40]
TR-MoE: Temporal Reliability-Aware Mixture-of-Experts for Robust Tracking
[1:45]
Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
[1:50]
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in the Semantic-Residual Space
[1:55]
Proteus: Model Leakage-Induced Adversarial Attack in Federated Learning
[2:00]
Beyond Filter Pruning: Top-K Spatial Selection for Efficient Neural Networks
[2:05]
Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations
[2:10]
When Higher Order Hurts: Pre-Asymptotic Order Collapse in Generative ODE Sampling — A Theory of Discretization–Learning Interaction
[2:15]
Structured-Noise Masked Modeling for Video, Audio and Beyond
[2:20]
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
[2:25]
Multi-Anchor Distillation with Text-Guided Analytic Classifier for Continual Learning
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30]
Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
[1:35]
Mimicking Radiologists: A Coarse-to-Fine Framework with Structural Sparse Tokens for Dual-LLM Computed Tomography Report Generation
[1:40]
VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
[1:45]
RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement
[1:50]
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
[1:55]
Ice Cloud Geometry Retrieval with Calibrated Uncertainty from Passive Satellite Imagery
[2:00]
Event-based Sparse-view Background-Oriented Schlieren Tomography
[2:05]
Zero-shot Depth from Defocus
[2:10]
Bayesian Self-Attention with Local Pixel Correlations for Lightweight Denoising Transformers
[2:15]
PolarAPP: Beyond Polarization Demosaicking for Polarimetric Applications
[2:20]
Stokes-Informed Diffusion for Robust Linear Polarization Estimation
[2:25]
Off the Planckian Locus: Using 2D Chromaticity to Improve In-Camera Color
(ends 3:00 PM)
3 p.m.
4 p.m.
4:30 p.m.
(ends 6:30 PM)
FRI 11 SEP
9 a.m.
Spotlights 9:00-10:15
[9:00]
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
[9:05]
Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding
[9:10]
GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
[9:15]
OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
[9:20]
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
[9:25]
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
[9:30]
Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
[9:35]
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
[9:40]
EgoExoMoCap: Distributed Human Motion Capture via Ego- and Exocentric Body Tracking from Head-Mounted Devices
[9:45]
Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction
[9:50]
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
[9:55]
CameraAnything: Refilming Videos with Arbitrary Camera Control
[10:00]
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
[10:05]
A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
[10:10]
Grounding World Simulation Models in a Real-World Metropolis
(ends 10:30 AM)
Orals 9:00-10:30
[9:00]
Steerable Vision Transformers
[9:15]
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
[9:30]
World Knowledge in the Weights: Reading Concept Circuits of Vision Transformers
[9:45]
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
[10:00]
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
[10:15]
Make Geometry Matter for Spatial Reasoning
(ends 10:30 AM)
Spotlights 9:00-10:15
[9:00]
RoMa v2: Harder Better Faster Denser Feature Matching
[9:05]
LoMa: Local Feature Matching Revisited
[9:10]
Gravity-aware partially calibrated absolute pose estimation from affine- or rotation-covariant features
[9:15]
Rolling Shutter Camera Self-Calibration
[9:20]
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
[9:25]
OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization
[9:30]
E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects
[9:35]
SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
[9:40]
Training-free Controllable Motion Generation under Heterogeneous Constraints
[9:45]
OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
[9:50]
VKSR: Scalable Kernel Surface Reconstruction Using Vecchia's Approximation
[9:55]
RADmesh: Remesh-Aware Mesh Deformation
[10:00]
Face Anything: 4D Face Reconstruction from Any Image Sequence
[10:05]
HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations
[10:10]
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
(ends 10:30 AM)
10:30 a.m.
(ends 12:30 PM)
Posters 10:30-12:30
AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware
CortexVideo: A Semantic-Spatial Dual-Anchor Framework for High-Fidelity fMRI-to-Video Reconstruction
(ends 12:30 PM)
noon
1:30 p.m.
Orals 1:30-2:45
[1:30]
Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat
[1:45]
NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
[2:00]
TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields
[2:15]
Repurposing Geometric Foundation Models for Multi-view Diffusion
[2:30]
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30]
Silhouette-based Gait Foundation Model
[1:35]
CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
[1:40]
QVAM: Query-guided View-aware Adaptive Modulation for Aerial-Ground Person Re-Identification
[1:45]
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
[1:50]
PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images
[1:55]
RiO-DETR: DETR for Real-time Oriented Object Detection
[2:00]
Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels
[2:05]
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
[2:10]
Instance Segmentation as Tracking: A New Paradigm for Multi-Small-Object Tracking with Event Cameras
[2:15]
SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
[2:20]
PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification
[2:25]
Event-based Gaze Control Systems for Real-time Spin Estimation in Professional Ball Games
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30]
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
[1:35]
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
[1:40]
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
[1:45]
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
[1:50]
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
[1:55]
EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
[2:00]
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
[2:05]
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
[2:10]
GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
[2:15]
Push–Pull Attentional Anchoring for Diffusion Concept Erasure
[2:20]
When Rubrics Fail: Error Enumeration as Reward for Reference-Free RL Post-Training
[2:25]
Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
(ends 3:00 PM)
3 p.m.
4 p.m.
(ends 6:00 PM)
(ends 6:00 PM)
6 p.m.
SAT 12 SEP
9 a.m.
10 a.m.
10:30 a.m.
(ends 12:30 PM)
Posters 10:30-12:30
When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators
MeanTalker: Efficient and Expressive Speech-Driven 3D Facial Animation via Geometric-Aware Mean Flow
Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training
Saber: Anchoring Semantics to Scale-Aware Kinetic Salience for Zero-Shot Skeleton Action Recognition
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
ORACLE-3D: Open-world Region-aligned Cross-modal Learning for Label-efficient 3D Scene Understanding
(ends 12:30 PM)
noon
1:30 p.m.
Spotlights 1:30-2:40
[1:30]
SPEAR: A Simulator for Photorealistic Embodied AI Research
[1:35]
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
[1:40]
Unordered Landmark Visual Navigation
[1:45]
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
[1:50]
Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
[1:55]
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
[2:00]
Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
[2:05]
TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
[2:10]
VOCA: Visual Odometry with Codec Awareness
[2:15]
PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving
[2:20]
Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
[2:25]
MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
[2:30]
CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
[2:35]
VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation
(ends 3:00 PM)
Spotlights 1:30-2:45
[1:30]
DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
[1:35]
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity
[1:40]
SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction
[1:45]
PanoRec: Spatially-Structured Sequence Modeling for Multi-Granularity Panoramic Retrieval
[1:50]
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
[1:55]
Incentivizing Vision Language Models to Search for Long Video Question Answering
[2:00]
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
[2:05]
HumanOmni-Speaker: Identifying Who said What and When
[2:10]
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
[2:15]
See & Sniff: Learning Visuo-Olfactory Representations
[2:20]
Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
[2:25]
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
[2:30]
When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs
[2:35]
How to Teach Large Multimodal Models New Skills
[2:40]
Event-Driven Video Generation
(ends 3:00 PM)
Orals 1:30-3:00
[1:30]
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
[1:45]
Geometric Context Transformer for Streaming 3D Reconstruction
[2:00]
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
[2:15]
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
[2:30]
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
[2:45]
Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
(ends 3:00 PM)
3 p.m.
(ends 5:00 PM)
(ends 5:00 PM)
Successful Page Load