← Search

Shengkai Zhang

8 accepted papers

2026

MOGS: Monocular Object-Guided Gaussian Splatting in Large Scenes

ICRA 2026poster

Recent advances in 3D Gaussian Splatting (3DGS) deliver striking photorealism, and extending it to large scenes opens new opportunities for semantic reasoning and prediction in applications such as autonomous driving. Today's state-of-the-art systems for large scenes primarily originate from LiDAR-b…

2025

VISC: mmWave Radar Scene Flow Estimation using Pervasive Visual-Inertial Supervision

IROS 2025

This work proposes a mmWave radar’s scene flow estimation framework supervised by data from a widespread visual-inertial (VI) sensor suite, allowing crowdsourced training data from smart vehicles. Current scene flow estimation methods for mmWave radar are typically supervised by dense point clouds f

Cited by 0SourceScholar
2024

Enhancing mmWave Radar Point Cloud via Visual-inertial Supervision

ICRA 2024poster

Complementary to prevalent LiDAR and camera systems, millimeter-wave (mmWave) radar is robust to adverse weather conditions like fog, rainstorms, and blizzards but offers sparse point clouds. Current techniques enhance the point cloud by the supervision of LiDAR’s data. However, high-performance LiD…

Cited by 0SourcecodeScholar
2024

Resolving Loop Closure Confusion in Repetitive Environments for Visual SLAM through AI Foundation Models Assistance

ICRA 2024poster

In visual SLAM (VSLAM) systems, loop closure plays a crucial role in reducing accumulated errors. However, VSLAM systems relying on low-level visual features often suffer from the problem of perceptual confusion in repetitive environments, where scenes in different locations are incorrectly identifi…

Cited by 5SourceScholar
2022

BodyGAN: General-Purpose Controllable Neural Human Body Generation

CVPR 2022poster

Recent advances in generative adversarial networks (GANs) have provided potential solutions for photorealistic human image synthesis. However, the explicit and individual control of synthesis over multiple factors, such as poses, body shapes, and skin colors, remains difficult for existing methods.…

Cited by 11PDFScholar
2022

DC-Loc: Accurate Automotive Radar Based Metric Localization with Explicit Doppler Compensation

ICRA 2022poster

Automotive mmWave radar has been widely used in the automotive industry due to its small size, low cost, and complementary advantages to optical sensors (e.g., cameras, LiDAR, etc.) in adverse weathers, e.g., fog, raining, and snowing. On the other side, its large wavelength also poses fundamental c…

Cited by 19SourcecodeScholar
2021

Conquering Textureless with RF-referenced Monocular Vision for MAV State Estimation

ICRA 2021poster

The versatile nature of agile micro aerial vehicles (MAVs) poses fundamental challenges to the design of robust state estimation in various complex environments. Achieving high-quality performance in textureless scenes is one of the missing pieces in the puzzle. Previously proposed solutions either…

Cited by 5SourcecodeScholar
2021

UltraPose: Synthesizing Dense Pose With 1 Billion Points by Human-Body Decoupling 3D Model

ICCV 2021poster

Recovering dense human poses from images plays a critical role in establishing an image-to-surface correspondence between RGB images and the 3D surface of the human body, serving the foundation of rich real-world applications, such as virtual humans, monocular-to-3d reconstruction. However, the popu…

Cited by 20PDFcodeScholar