Parse
Select informative keyframes, associate instances across views, and construct a verified scene graph with reference views, masks, partial point clouds, 3D boxes, and referring cues.
Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
¹ ByteDance Seed · ² Peking University · ³ Zhejiang University
Lucida reconstructs a real indoor scene as complete, editable object assets arranged as observed. It follows a parse, generate, and place pipeline, reserving precise alignment for the final closed-loop stage.
Each stage consumes evidence available in a cluttered real capture. The final stage uses observation feedback to close the loop instead of assuming perfect upstream geometry.
Select informative keyframes, associate instances across views, and construct a verified scene graph with reference views, masks, partial point clouds, 3D boxes, and referring cues.
Use each evidence bundle to synthesize an occlusion-free object image and lift it into a complete, editable 3D asset.
Initialize the asset coarsely, then let GizmoAct issue executable pose edits until the rendered asset aligns with the captured scene.
Lucida represents the reconstructed scene as a collection of complete object assets. Each object remains a separate, editable mesh ready for simulation.
A VLM operates a 3D editor in a closed loop. At each turn, it receives a set of rendered observations, issues one executable edit in the object's local frame, and decides when to stop.
Easy, medium, and hard cases show how the same policy handles progressively larger pose residuals.
For each case, the same policy starts from pose initializations produced by Boxer, Any6D*, and SAM 3D.
* Any6D-style depth-and-mask initialization only; FoundationPose prediction and refinement are excluded. All variants use the same GizmoAct policy, with up to four views and 12 refinement steps.
Lucida improves scene-level 3D object detection, object pose estimation, and complete scene reconstruction across R2S, CA-1M, and Aria Digital Twin.
Lucida produces a single deduplicated inventory of object instances with oriented 3D boxes before asset generation.
Mean class-agnostic AP over 3D IoU thresholds from 0.05 to 0.50. Higher is better.
all and key indicate which frames are used for the initial object prompts. The filtered protocol evaluates objects recovered by at least one method in the comparison.
GizmoAct improves strict surface alignment and oriented-box overlap on R2S-Object, CA-1M, and Aria Digital Twin.
Scene reconstruction on R2S-Scene. Scene CD uses squared Euclidean distances in metric scene coordinates; object-level metrics are computed after per-object normalization.
The ground-truth asset pose is rendered in blue. Mutual occlusion between the prediction and the blue reference reveals any residual alignment error.
Select a dataset above, then click the figure to inspect it at full resolution.
Compared with single-view scene reconstruction, Lucida better preserves object identity, scale, orientation, and overall arrangement.
We thank Wei Li, Heng Dong, and Qifeng Zhang for providing the pretrained VLM and for their help with model training. We thank Lihao Liu, Baifeng Xie, and Yiming Qiao for their help with simulation data, evaluation, and physical-plausibility post-processing.