Research vision and trajectory

Building World Models through 3D Intelligence

My research asks how machines can reconstruct, generate, and eventually simulate coherent three-dimensional environments. The path connects reliable spatial representations with scalable generative priors: from neural reconstruction and editable scene understanding, through pretrained 3D asset generation, toward open-ended world models for embodied intelligence.

Research directions

World Models and Infinite 3D Worlds

Generate large, continuous environments whose geometry and appearance remain coherent as the world expands.

  • WorldGrow — hierarchical, block-based infinite 3D scene generation.

Pretrained and Controllable 3D Generation

Learn scalable 3D priors while exposing control over appearance, semantics, and editable object structure.

  • UniLat3D — unified geometry-appearance latents.
  • TIGON — joint text-image conditioned generation.
  • SCULPT — part-aware generation by subtractive composition.

3D Reconstruction and Spatial Representation

Recover high-quality, navigable 3D content from sparse observations and make scene representations useful for downstream interaction.

  • GaussianObject — four-view Gaussian-splatting reconstruction.
  • NeRFVS — free-view synthesis with geometry scaffolds.
  • SA3D — interactive segmentation in neural 3D scenes.

Dynamic and Medical Spatial Intelligence

Model deformable scenes and challenging surgical observations where geometry, motion, and reliability matter together.

  • LerPlane — fast 4D reconstruction of deformable tissues.
  • EndoGSLAM — dense endoscopic reconstruction and tracking.
  • ForPlane — efficient dynamic neural representations.

Research trajectory

  1. 2021–2023 · RepresentationReliable neural scene representations. NeRFVS, SA3D, and LerPlane explored geometry priors, interactive understanding, and efficient dynamic reconstruction.
  2. 2024 · ReconstructionHigh-quality 3D from limited observations. GaussianObject and EndoGSLAM brought Gaussian splatting to extreme sparse views and medical scenes.
  3. 2025–2026 · GenerationScalable and controllable 3D priors. UniLat3D, TIGON, and SCULPT moved from unified generative representations to multimodal control and editable parts.
  4. 2026 onward · World modelsOpen-ended spatial intelligence. WorldGrow connects generative 3D priors to large, extendable environments for future embodied agents.

Open research and reproducible comparison

Most representative projects publish public code, project pages, and detailed paper records. The purpose is to make independent reproduction and fair comparison easier—not simply to showcase selected demos.

  • Open implementations. SCULPT, TIGON, UniLat3D, WorldGrow, GaussianObject, SA3D, and LerPlane expose source code or reproducible project artifacts.
  • Comparable evidence. Stable paper pages connect each method to papers, DOI or arXiv identifiers, code, available data or models, BibTeX, and explicit research keywords.
  • Substantive baselines. The work covers reconstruction, generation, multimodal control, editable structure, and scene scale, enabling comparisons across a coherent technical trajectory.
  • Machine-readable records. The complete portfolio is available as publication JSON, BibTeX, an Atom feed, and a researcher profile.

What this body of work demonstrates

Taken together, the portfolio provides evidence for research and engineering work in world models, 3D generation, spatial intelligence, 3D reconstruction, and embodied AI—not through a single isolated result, but through a continuous path from problem formulation and open implementation to large-scale generation and product delivery.

  • Technical depth: contributions span neural representations, Gaussian splatting, generative 3D latent spaces, multimodal conditioning, compositional assets, and infinite world generation.
  • Research continuity: reconstruction and scene understanding progressively connect to controllable 3D generation and foundation world models.
  • Research-to-product execution: advanced 3D generation work has been translated into consumer-facing creation tools, including a launch that reached more than 40,000 users within 50 hours.
  • Open evaluation: public implementations and stable bibliographic records allow other researchers and teams to inspect, reproduce, and compare the work directly.

Questions I am pursuing

  • How can a foundation world model maintain consistent geometry, appearance, and semantics over long spatial horizons?
  • How can pretrained 3D generators expose compositional structure without sacrificing asset quality?
  • How can reconstruction, generation, and simulation share one scalable spatial representation?
  • How can these models support embodied agents that understand, imagine, and interact with physical environments?