Large-Scale Light Field Synthesis from Videos Enables Geometrically Consistent Bokeh Editing
Abstract
Focus and depth of field (DoF) define where attention falls and how much of a scene appears sharp. Adjusting both after capture provides creators two-dimensional control over visual attention. However, existing focus and DoF editing tools fail when applied to scenes with complex geometries, such as reflections and fine features. This failure is a result of fundamental limitations in existing training data generation pipelines: current pipelines either sacrifice geometric fidelity for scalability, or sacrifice scalability for geometric fidelity. To address this, we present Vid2Bokeh, a scalable bokeh data acquisition pipeline that reconstructs dense 25×25 light fields from casual videos via feed-forward 3D reconstruction, bypassing depth estimation entirely and avoiding the geometric errors it introduces. The resulting Vid2Bokeh Dataset comprises over 100K light fields on diverse real-world scenes, simultaneously achieving scene diversity, optical density, and geometric fidelity that no existing bokeh dataset or bokeh data acquisition pipeline provides. By training on this geometrically faithful bokeh data, we introduce a diffusion model for full bidirectional focus–DoF editing that outperforms depth-based baselines on complex geometries. This result demonstrates that geometric fidelity in large-scale training data is the key to geometrically consistent bokeh editing.