Skip to content

Faculty Spotlight

Dr. Huaizu Jiang, Assistant Professor

A robot navigating a cluttered room in real time. A humanoid that moves and interacts with objects on command. Both depend on the same capability: a machine that can recover the three-dimensional world from flat, imperfect images. That challenge drives the research of Huaizu Jiang, assistant professor in the Khoury College of Computer Sciences at Northeastern University and a core faculty member at the Institute for Experiential Robotics (IER).

Jiang directs the 3D Visual Intelligence Lab, where his team develops algorithms to reconstruct the geometry and dynamics of the 3D world from images, and builds agents that act on that understanding, from humanoids that learn skills from human demonstrations to perception systems running onboard robots. He joined Northeastern in 2021 after working as a postdoctoral researcher at Caltech and a visiting researcher at NVIDIA. He received his PhD in Computer Science from UMass Amherst, where he was advised by Erik Learned-Miller, and completed research internships at NVIDIA Research and Facebook AI Research (FAIR).

Image 1. Huaizu Jiang

Research Foundations: Coordination Without a Central View

Figure 2: SV4D generates temporally consistent novel-view videos of dynamic 3D objects from a single monocular video, developed in collaboration with Stability AI (ICLR 2025).

Most vision systems treat images as flat grids of pixels. Jiang’s lab goes further, building algorithms that recover the full three-dimensional structure and dynamics of a scene. LASER (CVPR 2026) performs streaming 4D reconstruction at real-time speeds without any task-specific training, and Point4Cast (CVPR 2026) extends reconstruction to forecasting how dynamic scenes will evolve. Beyond camera images, machines also perceive the 3D world through direct sensory data such as point clouds; UniCorrn (CVPR 2026) introduces a single transformer that establishes dense correspondences across both 2D images and 3D point clouds, unifying two problems traditionally solved by separate systems. He has accumulated over 10,000 citations, making a sustained impact across the field.

The lab also develops generative models that synthesize visual content beyond what cameras observe. SV4D (ICLR 2025), developed in collaboration with Stability AI, generates novel-view videos of dynamic 3D objects from a single monocular video, making temporally consistent 4D content generation practical at scale. StreamForce (SIGGRAPH Asia 2026) goes a step further, turning generation into a real-time interactive process: from a single image, the model streams video that responds causally to user-applied physical forces. These efforts build on a line of work in video synthesis stretching back to Super SloMo (CVPR 2018, Spotlight), developed with NVIDIA Research, which has since been integrated into NVIDIA’s NGX platform and cited over 2,000 times.

3D Visually Intelligent Agents

Beyond reconstructing the 3D world, Jiang’s lab builds agents that act within it. A central focus is generating realistic humanoid motion and interaction. OmniControl (ICLR 2024) enables fine-grained control over human motion generation, specifying the position of any joint at any point in time. SK-HOI (SIGGRAPH Asia 2026) introduces a surface keypoint representation that generates multi-object and articulated human-object interactions, moving beyond the single-object, rigid-body settings of prior work.

Through IER, Jiang collaborates with faculty, including Hanumant Singh and Taskin Padir on perception systems that run onboard physical robots. NeuFlow (IROS 2024) delivers real-time optical flow on edge devices, and StereoVoxelNet (ICRA 2023) performs real-time obstacle detection from stereo cameras using learned occupancy voxels. Both are designed for the latency and power constraints of deployed robots rather than server-side benchmarks.

Figure 3: SK-HOI enables: (a) whole-body interaction of a variable number of rigid objects from text prompts; (b) interaction with articulated objects exhibiting diverse joint dynamics (e.g., laptop, drawers, and knobs); and (c) text-driven whole-body motion synthesis guided by sparse object waypoints or full object motion trajectories.

Mentorship and Student Outcomes

The 3D Visual Intelligence Lab is an active group with PhD students working across 3D reconstruction, visual generative modeling, spatial reasoning, and medical AI. Recent graduates and alumni have placed well across industry and academia: Yiming Xie, whose Apple AI/ML Fellowship supported work on SV4D and OmniControl, is joining a stealth startup; Qianru Lao is a Member of Technical Staff at OpenAI; Hongyu Li is a PhD student at Brown University; Lei Zhong, a Google PhD Fellow, is a PhD student at the University of Edinburgh; and Tianye Ding is pursuing a PhD at UT Austin. Jiang teaches CS 7150: Deep Learning and CS 5330: Pattern Recognition and Computer Vision at Khoury.

Looking Ahead

As robots move from scripted routines into open-ended tasks, they will need to understand their 3D surroundings in real time, reason about how to interact with objects and people, and plan physically grounded actions. These are the problems Jiang’s two research threads are converging on: 3D visual intelligence provides the perception, and humanoids trained on human demonstrations interact with the physical world naturally. Through IER’s collaboration, Jiang is working to close that loop, toward building robots that learn from human behavior and extend our ability to act in environments that are difficult, dangerous, or beyond human reach.

We use cookies to improve your experience on our sites. By continuing to use our sites, you agree to our Privacy Statement.