Orals
Provable and Robust Wavefront Sensing via Self-Reference Interferometry
Nebiyou Yismaw ⋅ Vishwanath Saragadam ⋅ Aswin C. Sankaranarayanan ⋅ M. Salman Asif
Wavefront sensing involves estimating the phase and intensity of light, enabling a wide range of imaging applications, from adaptive optics and astronomy to biomedical imaging. Since conventional image sensors can only measure the spatial intensity distribution, phase retrieval arises as the central problem in wavefront sensing. Conventional interferometric approaches like phase-shifting interferometry (PSI) can recover phase information, but they rely on a stable reference beam that is difficult to realize in practical settings. To overcome this limitation, we propose a novel self-reference framework that relies on interference between shifted copies of the incoming wave; this results in pairwise phase differences between shifted pixels. We formulate an analytical solution for the complete phase retrieval based on the propagation of these differences across a connected graph. Furthermore, we provide a theoretical analysis of optimal measurement patterns, proving that co-prime shifts guarantee a connected graph and bound worst-case error accumulation, yielding a provably robust method. Extensive simulations demonstrate that complete phase profiles can be recovered from as few as eight shifted measurements, outperforming several existing approaches. Finally, we validate our framework using a hardware prototype, demonstrating real experiments for optical phase profile recovery, auto-refocusing, and imaging through scattering media.
Show more
Broadband Wide Field of View Imaging with Computational Mirrors
Vishwanath Saragadam ⋅ Niki Nezakati ⋅ Amit Roy-Chowdhury ⋅ Vivek Boominathan
Traditional glass-based optics are typically optimized for nar-row spectral bands, such as the visible (400–700nm) or shortwave infrared(1000–1800nm). While the emergence of VIS-SWIR sensors (400–1700nm)offers transformative potential, refractive optics struggle to focus thisentire range simultaneously. Mirrors represent a promising achromaticalternative; however, they are often sidelined by field curvature, and off-axis aberrations. This paper introduces Computational Mirrors, aframework that enables high-resolution, full-field-of-view imaging acrossthe complete VIS-SWIR spectrum using a single sensor. Our method isbuilt on the observation that distinct regions of the field of view reachfocus at varying distances from the mirror. By capturing a minimal fo-cal stack (2–4 images), we utilize a computational backend to recovera sharp, all-in-focus image. A key contribution of this paper is Seidel-Conv, a novel, physics-inspired, spatially-varying point spread function(PSF) model designed to accurately characterize and correct the off-axisaberrations inherent in simple concave mirrors. We demonstrate the ef-ficacy of our approach using a first-of-its-kind 50mm F/1 optical systemequipped with a VIS-SWIR sensor. Our system produces sharp imagesacross RGB, NIR, and SWIR wavelengths without requiring refocusing,revealing material details invisible within individual spectral bands. Wefurther validate the scalability of our approach with a 100mm F/2 systemoptimized for long-range imaging.
Show more
Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation
Irene Kim ⋅ Sai Tanmay Reddy Chakkera ⋅ Alexandros Graikos ⋅ Dimitris Samaras ⋅ Akshat Dave
Monocular surface normal estimators trained on large-scale RGB-normal data often perform poorly in the edge cases of reflective, textureless, and dark surfaces. Polarization encodes surface orientation independently of texture and albedo, offering a physics-based complement for these cases. Existing polarization methods, however, require multi-view capture or specialized training data, limiting generalization. We introduce Poppy, a training-free framework that refines normals from any frozen RGB backbone using single-shot polarization measurements at test time. Keeping backbone weights frozen, Poppy optimizes perpixel offsets to the input RGB and output normal along with a learned reflectance decomposition. A differentiable rendering layer converts the refined normals into polarization predictions and penalizes mismatches with the observed signal. Across seven benchmarks and three backbone architectures (diffusion, flow, and feed-forward), Poppy reduces mean angular error by 23–26% on synthetic data and 6–16% on real data. These results show that guiding learned RGB-based normal estimators with polarization cues at test time refines normals on challenging surfaces without retraining. An interactive demo and code are available at https://irnkim.github.io/poppy/.
Show more
A second-order theory of texture for depth from focus
Sreekar Ranganathan ⋅ Ioannis Gkioulekas
We present a theory of textured appearance of optically roughsurfaces based on wave optics, emphasizing the role of texture for passivedepth from focus. Our theory shows that even surfaces that traditionalcomputer vision would consider textureless can produce textured appear-ance, due to subjective speckle from surface microgeometry. We analyzethe properties of this second-order texture, and show that we can enhanceits contrast under natural ambient lighting by simply using a narrowbandspectral filter. Doing so results in dramatic improvements in passive depthreconstruction of seemingly textureless scenes, as we demonstrate throughextensive theory, simulations, and real-world experiments.
Show more
PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment
Aymeric Fleith ⋅ Julian Zirbel ⋅ Daniel Cremers ⋅ Niclas Zeller
We introduce PRISM-VO, a novel pure optimization-basedsparse photometric visual odometry framework for focused plenopticcameras. The core of PRISM-VO is a novel photometric plenoptic bun-dle adjustment which jointly optimizes camera poses and inverse depthvalues of points in a sliding window. By combining geometric depth froma single plenoptic image with temporal multi-view constraints, PRISM-VO achieves accurate and drift-resilient motion estimation. Through ex-plicit modeling of the plenoptic projection, PRISM-VO provides reliablemetric-scale reconstructions, overcoming the scale ambiguity of monocu-lar SLAM algorithms. Importantly, our approach relies solely on a singleplenoptic sensor and avoids complex initialization, as depth priors arecomputed directly from plenoptic imaging.Experiments show that PRISM-VO outperforms the current state-of-the-art plenoptic visual odometry method on indoor and outdoor scenes.The proposed approach rivals other optimization- and learning-basedmethods while accurately and reliably recovering a metric scale of thescene.Project page: https://prism-vo.github.io/.
Show more
Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
Taeho Kang ⋅ Nairan Zhang ⋅ Yelin Kim ⋅ Yujiao Shi ⋅ Youngki Lee
Accurate yaw estimation is a bottleneck in cross-view lo-calization between ground view and Bird’s Eye View (BEV). Existingmethods couple yaw with translation and rely on height or projectionassumptions that degrade under large yaw ambiguity. We disentangleyaw from location accuracy and introduce LAYS, a radially invariantline-consensus voting method. By exploiting the radial invariance of ourformulation, we achieve sub-degree yaw precision via 3D voting over allcandidate poses, while eliminating the need for accurate location. Ourkey observation is that a ground-image column matched to BEV pixelsinduces the same yaw across all camera positions along the radial direc-tion of the pixels. LAYS matches BEV pixels to ground columns usingfeature similarity and accumulates the induced yaw votes into discrete3D bins, where correct correspondences along the radial line concentrateinto a sharp peak for the correct yaw. Experiments on Mapillary, Ford,KITTI, and VIGOR show significant gains under unknown yaw, particu-larly for normal FoV with unknown yaw (+28∼45%p), and using LAYSas a yaw prior improves downstream 3-DoF localization.
Show more
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
Shizhan Liu ⋅ Xinran Deng ⋅ Zhuoyi Yang ⋅ Jiayan Teng ⋅ Xiaotao Gu ⋅ Jie Tang
Latent diffusion models pair VAEs with diffusion backbones,and the structure of VAE latents strongly influences the difficulty of dif-fusion training. However, existing video VAEs typically focus on recon-struction fidelity, overlooking latent structure. We present a statisticalanalysis of video VAE latent spaces and identify two spectral propertiesessential for diffusion training: a channel-wise eigenspectrum dominatedby a few modes, and a spatio-temporal frequency spectrum biased towardlow frequencies. To induce these properties, we propose two lightweight,backbone-agnostic regularizers: Latent Masked Reconstruction and Lo-cal Correlation Regularization. Experiments show that our Spectral-Structured VAE (SSVAE) achieves a 3× speedup in text-to-video gen-eration convergence and a 10% gain in video reward, outperformingstrong open-source VAEs. Code is available at: https://github.com/zai-org/SSVAE.
Show more
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
JUNFU PU ⋅ Yuxin Chen ⋅ Teng Wang ⋅ Ying Shan
Current multimodal large language models (MLLMs) havedemonstrated remarkable capabilities in short-form video understand-ing, yet translating long-form cinematic videos into detailed, temporallygrounded scripts remains a significant challenge. This paper introducesthe novel video-to-script (V2S) task, aiming to generate hierarchical,scene-by-scene scripts encompassing character actions, dialogues, expres-sions, and audio cues. To facilitate this, we construct a first-of-its-kindhuman-annotated benchmark and propose a temporally-aware hierarchi-cal evaluation framework. Furthermore, we present OmniScript, an 8B-parameter omni-modal (audio-visual) language model tailored for long-form narrative comprehension. OmniScript is trained via a progressivepipeline that leverages chain-of-thought supervised fine-tuning for plotand character reasoning, followed by reinforcement learning using tem-porally segmented rewards. Extensive experiments demonstrate that de-spite its parameter efficiency, OmniScript significantly outperforms largeropen-source models and achieves performance comparable to state-of-the-art proprietary models, including Gemini 3-Pro, in both temporallocalization and multi-field semantic accuracy. The code is available athttps://github.com/TencentARC/OmniScript.
Show more
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
Yaofeng Su ⋅ Yuming Li ⋅ Zeyue Xue ⋅ Jie Huang ⋅ Siming Fu ⋅ Haoran Li ⋅ Haoyang Huang ⋅ Nan Duan
Recent joint audio-visual diffusion models achieve remark-able generation quality but suffer from high latency due to their bidirec-tional attention dependencies, hindering real-time applications. We pro-pose OmniForcing, the first framework to distill an offline, dual-streambidirectional diffusion model into a high-fidelity streaming autoregressivegenerator. However, naively applying causal distillation to such dual-stream architectures triggers severe training instability, due to the ex-treme temporal asymmetry between modalities and the resulting tokensparsity. We address the inherent information density gap by introducingan Asymmetric Block-Causal Alignment with a zero-truncation GlobalPrefix that prevents multi-modal synchronization drift. The gradient ex-plosion caused by extreme audio token sparsity during the causal shift isfurther resolved through an Audio Sink Token mechanism equipped withan Identity RoPE constraint. Finally, a Joint Self-Forcing Distillationparadigm enables the model to dynamically self-correct cumulative cross-modal errors from exposure bias during long rollouts. Empowered by amodality-independent rolling KV-cache inference scheme, OmniForcingachieves state-of-the-art streaming generation at ∼25 FPS on a singleGPU, maintaining multi-modal synchronization and visual quality on parwith the bidirectional teacher. Project Page: https://omniforcing.com.
Show more
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
Yeeun Choi ⋅ Youngbeom Yoo ⋅ Joon-Young Lee ⋅ Hyolim Kang ⋅ Seon Joo Kim
When videos extend from hours to days, directly process-ing them end-to-end becomes impractical for current Multi-modal LargeLanguage Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction topre-model high-level relations in videos, despite not knowing the down-stream query at build time. We instead prioritize high-recall retrievabil-ity during memory building, and defer query-specific, high-level relationcomposition to inference time. To this end, we propose MERIT (Multi-key Episodic Retrieval with Inference-time Temporal expansion), a sim-ple yet effective agentic framework for ultra-long video understanding.First, we formulate an episodic multi-key representation that enablesprecise retrieval of fine-grained memories through a simple key-matchingmechanism. Second, we introduce a neighbor filtering mechanism to cap-ture broader semantic context without the massive computational over-head of global memory construction. This is achieved by expanding thetemporal scope exclusively around the retrieved segments at inferencetime. By leveraging simple key-matching with this on-demand tempo-ral expansion, MERIT achieves state-of-the-art performance across threelong-video benchmarks: EgoLifeQA, LVBench, and Video-MME(Long).
Show more
FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
Eadom Dessalene ⋅ Botao He ⋅ Michael Maynord ⋅ Yonatan Tussa ⋅ Pavan Mantripragada ⋅ Yianni Karabatis ⋅ Nirupam Roy ⋅ Yiannis Aloimonos
We introduce FEEL (Force-Enhanced Egocentric Learning), the first large-scale dataset pairing force measurements gathered from custom piezoresistive gloves with egocentric video. Our gloves enable scalable data collection, and FEEL contains approximately 2 million force-synchronized frames of natural unscripted manipulation in kitchen environments, with ∼45% of frames involving hand-object contact. Because force is the underlying cause that drives physical interaction, it is a critical primitive for physical action understanding. We demonstrate the utility of force for physical action understanding through application of FEEL to two families of tasks: (1) contact understanding, where we jointly perform temporal contact segmentation and pixel-level contacted object segmentation; and, (2) action representation learning, where force prediction serves as a self-supervised pretraining objective for video backbones. We achieve state-of-the-art temporal contact segmentation results and competitive pixel-level segmentation results without any need for manual contacted object segmentation annotations. Furthermore we demonstrate that action representation learning with FEEL improves transfer performance on action understanding tasks without any manual labels over EPIC-Kitchens, SomethingSomething-V2, EgoExo4D and Meccano.
Show more
Steerable Vision Transformers
Jona Ruthardt ⋅ Manu Gaur ⋅ Deva Ramanan ⋅ Makarand Tapaswi ⋅ Yuki Asano
Pretrained Vision Transformers (ViTs) such as DINOv2 andMAE provide generic image features that can be applied to a varietyof downstream tasks such as retrieval, classification, and segmentation.However, such representations tend to focus on the most salient visualcues in the image, with no way to direct them toward less prominentconcepts of interest. In contrast, Multimodal LLMs can be guided withtextual prompts, but the resulting representations tend to be language-centric and lose their effectiveness for generic visual tasks. To addressthis, we introduce Steerable Visual Representations, a new class of vi-sual representations, whose global and local features can be steered withnatural language. While most vision-language models (e.g., CLIP) fusetext with visual features after encoding (late fusion), we inject text di-rectly into the layers of the visual encoder (early fusion) via lightweightcross-attention. We introduce benchmarks for measuring representationalsteerability, and demonstrate that our steerable visual features can focuson any desired object in an image while preserving the underlying rep-resentation quality. Our method also matches or outperforms dedicatedapproaches on anomaly detection and personalized object discrimination,exhibiting zero-shot generalization to out-of-distribution tasks.Project Website: jonaruthardt.github.io/project/SteerViTPrompt CLS AttentionSteerable Visual RepresentationsGoal: control what vision features encode New Pareto FrontierEval: retrieving images w/ prompted objects Previous SoTAUse: task-specific adaptationFig. 2: SteerViT produces high-quality visual representations that can besteered by text. Left: Traditional (non-steerable) representations like DINOv2 tendto focus on the dominant object in an image and retrieve images with the same object.SteerViT can adapt to a text prompt, enabling retrieval of images even with smallobjects of interest. Right: We compare SteerViT to prior work in terms of its ability toadapt to text (measured by text-guided image retrieval (cf. Sec. 4.1)) and the qualityof the visual representation (measured by the accuracy of linear probing for the CLSfeature and semantic segmentation for patch features). While models typically tradeoff steerability for representation quality, SteerViT preserves both. By modulating agating factor (Eq. (2)), SteerViT achieves a new Pareto frontier.
Show more
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
Weichen Zhou ⋅ Yawen Zou ⋅ Chunzhi Gu ⋅ Ran Dong ⋅ Haoran Xie ⋅ Chao Zhang
We introduce a controlled subspace intervention framework to investigate how self-supervised Vision Transformers (ViTs) encode dense geometric information. While linear probing is widely used to assess geometric representations, it treats features as a black box, failing to disentangle the underlying topology. To address this issue, we decompose the weights of converged linear probes to isolate the low-rank subspaces containing explicit geometric signals using Singular Value Decomposition (SVD). Our perspective yields three key insights: (1) Pre-training objectives determine how features are encoded. DINOv2 aligns spatial features for efficient linear extraction, while Masked Autoencoders (MAE) tend to disperse these signals, requiring a broader spatial context. (2) Explicit geometric representations are highly compressible, suggesting dense predictive heads could potentially be constrained to low-rank subspaces with minimal performance loss. (3) The layer-wise task affinity suggests that geometric precision peaks at intermediate layers before yielding to semantic abstraction in the final layers. By connecting internal encoding mechanics with downstream performance, these findings provide a basis for effective feature selection and lightweight decoder design. The source code is available at https://github.com/Zhou-Weichen/Geosubprobe.
Show more
World Knowledge in the Weights: Reading Concept Circuits of Vision Transformers
Yanlin Chen ⋅ Tang Li ⋅ Xi Peng
Vision transformers (ViTs) have achieved remarkable gen-eralization across visual domains, yet little is known about how theyinternally represent the structure of the world. To address this gap, weuse Cross-Layer Transcoders (CLTs) to read concept circuits from ViTs:directed graphs whose nodes correspond to sparse, interpretable con-cepts and edges capture concept interactions across layers. Our methodyields two complementary views of model behavior. The global conceptcircuit is input-invariant and can be recovered directly from learnedcross-layer weights, exposing the reusable “world knowledge” encodedin the model. The instance concept circuit is input-dependent and iden-tifies the concepts and pathways actually used for a specific prediction,enabling faithful example-level explanations. We demonstrate the util-ity of concept circuits in three ways: (1) Automatic spurious correla-tion discovery: leveraging the statistics of our global concept circuitsto identify shortcut dependencies within the model. (2) Spurious cor-relation removal: intervening on the instance concept circuit to steerthe model towards correct predictions. Empirical results show that ourmethod outperforms existing counterparts by 11.0% on the Waterbirddataset. (3) Model comparison: contrasting the global concept circuits ofdifferent foundation models (e.g., CLIP vs. DINO) to reveal how super-vision paradigms shape representational structure. Our code is availableat https://github.com/deep-real/VisionCLT
Show more
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
Chen-yi Lu ⋅ Yueh-Shao Chen ⋅ Somali Chaterji
Contrastive vision-language models such as CLIP map se-mantically opposite phrases (e.g., “a dog” vs. “not a dog”) to nearly iden-tical embeddings, rendering them insensitive to negation. We attributethis failure to a phenomenon we call Representational Collapse: by track-ing compositional divergence and visual alignment across the CLIP textencoder, we show that middle layers build compositional syntax, but thefinal layers collapse this structure as visual alignment rises, producinga syntax-blind final representation. To recover the lost negation signalwithout altering pretrained weights, we propose PeakPatch, a lightweightpost-hoc correction system that intercepts the encoder at its composi-tional peak while keeping CLIP fully frozen. An Embedding CorrectionNetwork (ECN) uses cross-attention to extract a negation-specific sig-nal from the peak layer, anchored to a stable baseline, and predicts adeviation vector that re-injects the lost syntax into the final-layer em-bedding space. A complementary Score Correction Network (SCN) pre-dicts bounded scalar score offsets for discriminative tasks. Both modulesare trained jointly end-to-end while all CLIP parameters remain frozen,adding only 5.2M parameters (3.5% of the backbone) and preserving thestandard cosine similarity interface. On NegBench, PeakPatch achieves74.3% on COCO MCQ (+35.1 over CLIP, +17.8 over the best encoderfine-tuning method) and 65.5% on VOC MCQ, while outperforming allfine-tuning baselines on fully out-of-distribution negation retrieval de-spite training only 3.5% of the parameters. The corrected embeddingsalso transfer to text-to-image generation (+18.4 negation score) and gen-eralize across ViT-B/32, ViT-L/14, and SigLIP backbones.Project page: https://stevencylu.github.io/PeakPatch/
Show more
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
Yiming Qin ⋅ Bomin Wei ⋅ Jiaxin Ge ⋅ Konstantinos Kallidromitis ⋅ Stephanie Fu ⋅ Trevor Darrell ⋅ XuDong Wang
Vision–Language Models (VLMs) excel at reasoning in lin-guistic space but struggle with perceptual understanding that requiresdense visual perception, e.g., spatial reasoning and geometric aware-ness. This limitation stems from the fact that current VLMs have lim-ited mechanisms to capture dense visual information across spatial di-mensions. We introduce Chain-of-Visual-Thought (CoVT), a frameworkthat enables VLMs to reason not only with discrete text tokens but alsothrough continuous visual tokens—compact latent representations thatencode rich perceptual cues. With a small budget of roughly 20 tokens,CoVT distills knowledge from lightweight vision experts that capturecomplementary properties such as 2D appearance, 3D geometry, spatiallayout, and edge structure. During training, the VLM with CoVT au-toregressively predicts these visual tokens to reconstruct dense supervi-sion signals (e.g., depth, segmentation, edges, and DINO features). At in-ference, the model reasons directly in the continuous visual latent space,preserving efficiency while optionally decoding dense predictions for in-terpretability. Evaluated across more than ten diverse benchmarks, in-cluding CV-Bench, MME-RealWorld, MMVP, RealWorldQA, MMStar,WorldMedQA, and HRBench, integrating CoVT into strong VLMs suchas Qwen2.5-VL and LLaVA consistently improves performance by 3% to16% and demonstrates that compact continuous visual thinking enablesmore precise, grounded, and interpretable multimodal intelligence.
Show more
Make Geometry Matter for Spatial Reasoning
Shihua Zhang ⋅ Qiuhong Shen ⋅ Shizun Wang ⋅ Tianbo Pan ⋅ Xinchao Wang
Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to handle this limitation by injecting geometry tokens from pretrained 3D foundation models into VLMs. Nevertheless, we observe that naive token fusion followed by standard finetuning in this line of work often leaves such geometric cues underutilized for spatial reasoning, as VLMs tend to rely heavily on 2D visual cues. In this paper, we propose GeoSR, a framework designed to make geometry matter by encouraging VLMs to actively reason with geometry tokens. GeoSR introduces two key components: (1) Geometry-Unleashing Masking, which strategically masks portions of 2D vision tokens during training to weaken non-geometric shortcuts and force the model to consult geometry tokens for spatial reasoning; and (2) Geometry-Guided Fusion, a gated routing mechanism that adaptively amplifies geometry token contributions in regions where geometric evidence is critical. Together, these designs unleash the potential of geometry tokens for spatial reasoning tasks. Extensive experiments on both static and dynamic spatial reasoning benchmarks demonstrate that GeoSR consistently outperforms prior methods and establishes new state-of-the-art performance by effectively leveraging geometric information. The project page is available at https://suhzhang.github.io/GeoSR/.
Show more
Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat
Simone Foti ⋅ Caner Korkmaz ⋅ Stefanos Zafeiriou ⋅ Tolga Birdal
3D Gaussian Splatting has recently revolutionised novel viewsynthesis as well as many other 3D vision methods and applications.Drawing inspiration from this representation, we now rethink texturesto overcome the main issues of UV mapping while considerably lower-ing their memory footprint. Heat Kernel Textures (HKTex) eliminateUV unwrapping as well as their persistent issues of wasted UV space,seams, distortions, vertex-duplication, and varying resolution. Groundedin discrete Riemannian geometry and intrinsically defined on any man-ifold surface discretised as a triangular mesh, HKTex uses anisotropicheat kernels as geodesic equivalents to Gaussians. Like our kernels, alsothe optimisation of their position and the adaptive densification strate-gies were redefined to operate on the surface of the object to be tex-tureised. Our novel representation is also fully integrated with a physi-cally based renderer and can be optimised either from existing texturesor multi-view images. Our project page and code are available at circle-group.github.io/research/HeatKernelTextures.
Show more
NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
Jingye Qiu ⋅ Shizhe Zhou
Differentiable rendering of planar rational splines remains largely underexplored, despite their widespread use in vector graphics and design. Existing differentiable vector renderers primarily focus on Bézier curves and rely on analytic rasterization, which can suffer from gradient instability and limited flexibility. We propose NURBS Splatting, a unified framework that represents planar rational curves as continuous Gaussian fields. By sampling Gaussians along the curve parameter domain and inside closed regions, rendering is reformulated as a smooth accumulation process with stable gradients. Our method naturally supports long splines, rational weights, non-uniform knots, and closed-region filling. We demonstrate its effectiveness in calligraphy reconstruction, vectorization frameworks, and long-spline image abstraction, showing improved stability and reconstruction quality over existing approaches.
Show more
TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields
Haoxuan Li ⋅ Ziya Erkoç ⋅ Daniele Sirigatti ⋅ Vladislav Rosov ⋅ Lei Li ⋅ Angela Dai ⋅ Matthias Niessner
We present TriFlow, a new generative approach for produc-ing compact 3D meshes with artist-like triangle topology directly frominput geometry conditions such as signed distance fields. Our key insightis to represent mesh topology as a nearest-vertex vector field (NVF) de-fined over the surface, where each point encodes its association to thenearest triangle vertex in the local barycentric frame. We train a latentflow-matching model to synthesize this field, enabling topology genera-tion conditioned on the input geometry. To extract a coherent mesh, wecluster surface regions using the generated NVF and guide a constrainedquadric error metric mesh simplification with topology-aware optimiza-tion. This yields output meshes that closely match the input geometrywhile exhibiting structured, artist-like connectivity. Experiments demon-strate that TriFlow achieves stronger generalization and significantlyimproved topology quality compared to state-of-the-art learning-basedapproaches, alongside 90% lower Chamfer Distance and an 8× speedup.
Show more
Repurposing Geometric Foundation Models for Multi-view Diffusion
Wooseok Jang ⋅ Seonghu Jeon ⋅ Jisang Han ⋅ Jinhyeok Choi ⋅ Minkyung Kwon ⋅ Seungryong Kim ⋅ Saining Xie ⋅ Sainan Liu
The latent space of diffusion models fundamentally deter-mines their learning efficiency and generation quality. While recent ad-vances in the latent space have driven substantial progress in single-imagegeneration, the optimal latent space for novel view synthesis (NVS) re-mains largely unexplored. In particular, NVS requires geometrically con-sistent generation across viewpoints, but existing approaches typicallyoperate in a view-independent latent space. In this paper, we proposeGeometric Latent Diffusion (GLD), a framework that repurposesthe feature space of a geometric foundation model as the latent space formulti-view diffusion. We show that the features of the geometric founda-tion model not only support high-fidelity RGB reconstruction but alsoencode strong cross-view geometric correspondences, providing a well-suited latent space for NVS. Through experiments, GLD outperformsboth VAE and RAE on 2D image quality and 3D consistency metrics,accelerating training by more than 4.4× compared to the VAE latentspace. Notably, GLD remains competitive with state-of-the-art methodsthat leverage large-scale text-to-image pretraining, despite training itsdiffusion model from scratch without such generative pretraining.
Show more
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
Yang Chen ⋅ Xiaowei Xu ⋅ Shuai Wang ⋅ Xinwen Zhang ⋅ Qiushi Guo ⋅ Tiezheng Ge ⋅ Limin Wang
Normalizing Flows (NFs) are powerful generative models ca-pable of exact density estimation and sampling. However, their strictinvertibility often forces the model to exhaust its capacity on low-levelpixel details, hindering the capture of high-level semantic structures.While Masked Image Modeling (MIM) has excelled in representationlearning, its integration into generative pipelines has remained largelymodular and disjointed. In this paper, we propose MIMFlow, a unifiedend-to-end framework that jointly optimizes latent semantics, pixel re-construction, and generative flow. By employing a VAE encoder to infersemantic latent from masked images, MIMFlow achieves a principleddecoupling of the generative task: the Normalizing Flow focuses on mod-eling a simplified, low-frequency semantic manifold, while a specializeddecoder handles high-frequency synthesis. This design effectively resolvesthe inherent capacity bottleneck of NFs, allowing the model to prioritizeglobal structural coherence over redundant noise. Empirical results onImageNet 256×256 show that MIMFlow-L reaches 71.3% linear prob-ing accuracy and an FID of 2.50. Despite using only 128 tokens (50%fewer than standard models), it yields a 32.8% performance gain oversimilar-scale NF baselines.
Show more
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
Yue Pan ⋅ Tao Sun ⋅ Liyuan Zhu ⋅ Lucas Nunes ⋅ Iro Armeni ⋅ Jens Behley ⋅ Cyrill Stachniss
Point cloud registration aligns multiple unposed point clouds into a common reference frame and is a core step for 3D reconstruction and robot localization when no initial pose guess is available. In this work, we cast point cloud registration as conditional generation: a learned, continuous point-wise velocity field transports noisy points to a registered scene, from which the pose of each view is recovered. Unlike prior methods that perform correspondence matching to estimate pairwise transformations and then optimize a pose graph for multi-view registration, our model directly generates the registered point cloud, yielding both efficiency and point-level global consistency. By scaling the training data and conducting test-time rigidity enforcement, our approach achieves state-of-the-art average performance on existing pairwise registration benchmarks and on our proposed cross-domain multi-view registration benchmark. The superior zero-shot performance on this benchmark demonstrates that our method generalizes across view counts, scene scales, and sensor modalities even with low overlap.
Show more
Geometric Context Transformer for Streaming 3D Reconstruction
Lin-Zhuo Chen ⋅ Jian Gao ⋅ Shangzhan Zhang ⋅ Yihang Chen ⋅ Nan Xue ⋅ Jianyuan Wang ⋅ Christian Rupprecht ⋅ Xun Cao ⋅ Xing Zhu ⋅ Yujun Shen ⋅ Yao Yao ⋅ YINGHAO XU
Streaming 3D reconstruction aims to recover 3D information, such as camera poses and point clouds, from a video stream, which necessitates geometric accuracy, temporal consistency, and computational efficiency. Motivated by the principles of Simultaneous Localization and Mapping (SLAM), we introduce GCT, geometric context transformer, a feed-forward 3D foundation model for reconstructing scenes from streaming data. A defining aspect of GCT lies in its carefully designed attention mechanism, which integrates an anchor context, a pose-reference window, and a trajectory memory to address coordinate grounding, dense geometric cues, and long-range drift reduction, respectively. This design keeps the streaming state compact while retaining rich geometric context, enabling stable efficient inference at around 20 FPS on 518×378 resolution inputs over long sequences exceeding 10,000 frames. Extensive evaluations across a variety of benchmarks demonstrate that our approach achieves superior performance compared to both existing streaming and iterative optimization-based approaches.
Show more
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
Zhengqin Li ⋅ Cheng Zhang ⋅ Jakob Engel ⋅ Dong Zhao
We introduce the Large Sparse Reconstruction Model tostudy how scaling transformer context windows affects feed-forward 3Dreconstruction. Although recent object-centric feed-forward methods pro-duce robust, high-quality reconstructions, they still lag behind dense-view optimization in recovering fine-grained texture and appearance. Weshow that expanding the context window—by substantially increasingthe number of active object and image tokens—narrows this gap andenables high-fidelity 3D object reconstruction and inverse rendering. Toscale effectively, we adapt native sparse attention [68] for 3D reconstruc-tion with three key contributions: (1) an efficient coarse-to-fine pipelinethat focuses computation on informative regions by predicting sparsehigh-resolution residuals; (2) a 3D-aware spatial routing mechanism thatestablishes accurate 2D-3D correspondences using explicit geometric dis-tances rather than standard attention scores; and (3) a custom block-aware sequence-parallel strategy with an All-gather-KV protocol to bal-ance dynamic, sparse workloads across GPUs. As a result, LSRM handles20× more object tokens and >2× more image tokens than prior state-of-the-art (SOTA) methods. Extensive evaluations on standard novel-viewsynthesis benchmarks show substantial gains over the current SOTA,yielding >2.4 dB higher PSNR and >40% lower LPIPS. Furthermore,when extending LSRM to inverse rendering, qualitative and quantitativeevaluations on widely used benchmarks demonstrate consistent improve-ments in texture and geometry details, achieving an LPIPS that matchesor exceeds that of SOTA dense-view optimization methods. Code andmodel weights are available on our project page.
Show more
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
Nicolas von Lützow ⋅ Barbara Roessle ⋅ Katharina Schmid ⋅ Matthias Niessner
Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer-based model that directly generates 3D Gaussians via next-token prediction, thus facilitating full 3D scene generation. We first compress Gaussian primitives into a discrete latent grid using a sparse 3D convolutional autoencoder with vector quantization. The resulting tokens are serialized and modeled using a causal transformer with 3D rotary positional embedding, enabling sequential generation of spatial structure and appearance. Unlike diffusion-based methods that refine scenes holistically, our formulation constructs scenes step-by-step, naturally supporting completion, outpainting, controllable sampling via temperature, and flexible generation horizons. This formulation leverages the compositional inductive biases and scalability of autoregressive modeling while operating on explicit representations compatible with modern neural rendering pipelines, positioning autoregressive transformers as a complementary paradigm for controllable and context-aware 3D generation.
Show more
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
Kadir Yilmaz ⋅ Adrian Kruse ⋅ Tristan Höfer ⋅ Daan de Geus ⋅ Bastain Leibe
Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong domain priors. This keeps the field isolated from the broader Transformer ecosystem, limiting the transfer of research advances from other domains and the benefits of increasingly optimized software and hardware stacks. To bridge this gap, we propose the Volume Transformer (Volt), which adapts the vanilla Transformer encoder to 3D scenes with minimal modifications. Specifically, Volt partitions 3D scenes into volumetric patch tokens, processes them with full global self-attention, and injects positional information through 3D rotary positional embeddings (RoPE). Our initial experiments reveal that naively training Volt on standard 3D benchmarks leads to poor generalization, highlighting the limited scale of current 3D supervision. To overcome this, we introduce a data-efficient training recipe based on strong 3D augmentations, regularization, and distillation from a convolutional teacher, making Volt competitive with state-of-the-art methods. We then scale supervision through joint training on multiple datasets and show that Volt benefits more from increased scale than domain-specific 3D backbones, achieving state-of-the-art results across indoor and outdoor semantic segmentation benchmarks. Finally, we use Volt as a drop-in backbone in a standard 3D instance segmentation pipeline, where it also achieves new state-of-theart results, highlighting its potential as a simple, scalable, and generalpurpose backbone for 3D scene understanding.
Show more
Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
Oleksii Nasypanyi ⋅ Jaemin Cho ⋅ Utku Ozbulak ⋅ Byungkon Kang ⋅ Francois Rameau
Scene Coordinate Regression (SCR) methods are increasingly adopted for visual localization. In these approaches, the scene is implicitly encoded within a neural network that regresses a 3D world coordinate for each image pixel. Because the scene is represented only through the network parameters and not stored explicitly as images or maps, such methods are often assumed to be privacy-preserving. In this work, we show that this assumption is incorrect in practice. Specifically, we introduce a query-based attack that reconstructs the 3D geometry of the training environment from an SCR model under different levels of model access. To do so, we repeatedly query the model with batches of proxy images unrelated to the target scene to obtain dense pixel-wise 3D coordinates. Reliable points are identified through their stability under small input perturbations and can be further refined in a white-box setting. These stable points are accumulated across independent query batches to recover the scene geometry. From the recovered 3D representation, we also invert the network features to synthesize images from arbitrary viewpoints, revealing additional appearance information. Experiments on indoor and outdoor datasets demonstrate that substantial portions of training environments can be reconstructed with high geometric fidelity. Beyond geometry, we also recover an approximate color appearance, which exposes recognizable layout and potentially sensitive scene elements. This directly contradicts claims in the literature that SCR representations are privacy-preserving by design, and reveals a real risk when such systems are deployed in private or security-critical spaces. The project page is available here.
Show more
Successful Page Load