Research
I work on 3D vision, embodied intelligence, and generative scene modeling. My recent work explores efficient 3D Gaussian Splatting for dynamic and active reconstruction, single-image 3D scene generation, and long-horizon embodied navigation in open-world scenarios.
|
|
DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction
Hongbo Duan, Pengting Luo, Chengzhi Zhao, Yuanhao Chiang, Fangming Liu, Xueqian Wang
ACM MM, 2026
An active reconstruction framework based on 3DGS for dynamic scenes. It decouples structural and motion uncertainty to guide view selection and path planning, forming a closed-loop active reconstruction pipeline that outperforms existing baselines in accuracy and exploration efficiency.
|
|
NavCrafter: Exploring 3D Scenes from a Single Image
Hongbo Duan, Peiyu Zhuang, Yi Liu, Zhengyang Zhang, Yuxin Zhang, Pengting Luo, Fangming Liu, Xueqian Wang
ICRA, 2026
A single-image 3D scene exploration framework that combines video diffusion priors with geometry-aware expansion strategies. It features a camera trajectory modulation module and a collision-aware planner for controllable large-viewpoint novel view synthesis, and achieves SOTA 3DGS reconstruction on RealEstate10K.
|
|
CausalNav: A Framework for Long-term Embodied Navigation of Autonomous Vehicles in Outdoor Open Scenarios
Hongbo Duan, Shangyi Luo, Zhiyuan Deng, Yanbo Chen, Yuanhao Chiang, Yi Liu, Fangming Liu, Xueqian Wang
IEEE Robotics and Automation Letters (RA-L)
An open-world semantic navigation framework driven by geodata augmentation and RAG. It builds multi-level embodied graphs and robust planning algorithms to solve long-distance navigation in dynamic outdoor environments, supporting dynamic retrieval of both coarse buildings and fine-grained objects.
|
|
GSWAM: A Unified Gaussian World Action Model for Robotic Manipulation
Chengzhi Zhao*, Hongbo Duan*, Yueyang Weng, Xiaopeng Zhang, Mu Yongjin, Bin Qian, Yanjie Li
CORL, 2026
GSWAM is a 3D Gaussian‑based world‑action model. It unifies scene dynamics and robot actions in a shared 3D latent space, enabling joint action generation and scene prediction, and achieves strong results on simulated and real‑world robotic manipulation.
|
|
G2P-WAM: From Static–Dynamic Geometry Distillation to Preference Alignment of World-Action Models
Hongbo Duan*, Bin Qian*, Chengzhi zhao*, Fan Du, Yan Gao, Fangming Liu, Xueqian Wang
ICLR, 2027 Under Review
G2P-WAM is a geometry-to-preference post-training framework that connects static–dynamic representation distillation to action-policy preference alignment. It features three stages: GeoSFT grounds WAM representations with removable static-3D and dynamic-4D supervision, reward diagnosis validates geometric scores on the policy's own rollouts, and GeoDPO aligns the action stream with same-context preference pairs. It improves closed-loop manipulation success across benchmarks with no inference overhead.
|
|
STaR-WAM: Static and Temporal Representation Alignment for World-Action Models
Bin Qian*, Hongbo Duan*, Feng yan, Guangxin Wu, Zhijie Song, Fan Du, Mingxin Wang, Keru Zhou, Junwei Li, Chenxi Wu, Chengzhi zhao, Xueqian Wang, Yan Wang, Hao Wang, HENG YANG
AAAI, 2026 Under Review
We propose STaR‑WAM, a training‑only geometric supervision framework for World‑Action Models. Static and temporal geometry alignment losses distill 3D structural and motion knowledge without manual annotations. All auxiliary teachers are removed after training, bringing no inference overhead. It achieves higher success rates on geometry‑sensitive robotic manipulation tasks and faster convergence.
|
|
GGD-SLAM: Monocular 3DGS SLAM Powered by Generalizable Motion Models for Dynamic Environment
Yi Liu, Haoxuan Xu, Hongbo Duan, Keyu Fan, Zhengyang Zhang, Peiyu Zhuang, Pengting Luo, Houde Liu
ICRA, 2026
A 3D Gaussian‑based visual SLAM framework for dynamic environments. It eliminates the need for predefined semantics or depth inputs, exploits sequential attention and dynamic feature enhancement to separate static‑dynamic components, and adopts occlusion filling and distractor‑adaptive SSIM loss for strong robustness, achieving state‑of‑the‑art localization and dense reconstruction results.
|
|
EventGS: Event-based 3D Gaussian Splatting SLAM with Diffusion Model
Chengzhi Zhao*, Hongbo Duan*, Ruixiang Wang, Yanjie Li, Fangming Liu, Xueqian WANG
IROS, 2026
EventGS is a near real‑time 3DGS‑based SLAM system for monocular event streams. It leverages diffusion‑restored intensity and event‑accumulated observations together with progressive refinement to enable accurate pose tracking and high‑quality renderable reconstruction under extreme HDR and fast‑motion conditions.
|
|
LIO-Track: Tightly-Coupled LiDAR-Inertial Odometry with Multi-Object Tracking
Zhengcheng Yu, Hongbo Duan, Xueqian WANG, Bin Liang
IROS, 2026
LIO‑Track is a tightly‑coupled LiDAR‑Inertial odometry with multi‑object tracking built upon hierarchical factor‑graph optimization and ESKF, which jointly optimizes SLAM and MOT for improved pose accuracy and robust dynamic‑object processing.
|
|
GEN3D: Generating Domain-Free 3d Scenes From a Single Image
Yuxin Zhang, Ziyu Lu, Hongbo Duan, Keyu Fan, Pengting Luo, Peiyu Zhuang, Mengyu Yang, Houde Liu
ICASSP, 2026
Gen3d generates high‑quality generic 3D Gaussian‑based world models from a single image. It lifts RGBD input into point clouds, expands and optimizes the scene representation to synthesize consistent, high‑fidelity novel views.
|
|
FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis
Zhengyang Zhang, Ziyu Lu, PengCheng Li, Hongbo Duan, Yi Liu, Pengting Luo, Peiyu Zhuang, Xinghui Li, Shaohua Ma
ICASSP, 2026
We introduce a Fourier motion modeling module for 4D Gaussian Splatting, decomposing motion into sinusoidal frequency components to capture complex dynamics. Combined with frequency‑aware regularization, it enhances complex‑motion fitting and long‑term temporal coherence while preserving real‑time rendering.
|
|
Tencent Hunyuan
TEG Multi‑Modal Model Department, Algorithm Research Intern
2026.03 – Present
|
|
Huawei 2012 Labs
Central Media Technology Institute, Algorithm Research Intern
2025.06 – 2025.12
|
|
Education
Tsinghua University Shenzhen International Graduate School · Direct‑track PhD, Electronic Information (Intelligent Robotics) · 2024.09 – 2029
Harbin Institute of Technology, Weihai · B.Eng., Automation · 2020.09 – 2024.06
|
|
Awards
National Scholarship (Undergraduate)
Interdisciplinary Contest in Modeling (ICM) Finalist
Outstanding Graduate, Harbin Institute of Technology
|
|
Academic Services
Conference Reviewer: ICLR 2026, NeurIPS 2026, ECCV 2026, ICRA 2026, IROS 2026
Journal Reviewer: IEEE Transactions on Multimedia, IEEE Robotics and Automation Letters
|
|