← Search

Kai Cheng

13 accepted papers

2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

ICML 2026poster

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improvement. However, existing embodied benchmarks fail to provide actionable insights because they focus on task-level evaluation rather than discovering capability bottlenecks. To address t…

Cited by 0SourceScholar
2025

EfficientEQA: An Efficient Approach to Open-Vocabulary Embodied Question Answering

IROS 2025

Embodied Question Answering (EQA) is an essential yet challenging task for robot assistants. Large vision-language models (VLMs) have shown promise for EQA, but existing approaches either treat it as static video question answering without active exploration or restrict answers to a closed set of ch

Cited by 11SourcecodeScholar
2025

GaussianRoom: Improving 3D Gaussian Splatting with SDF Guidance and Monocular Cues for Indoor Scene Reconstruction

ICRA 2025

Embodied intelligence requires precise reconstruction and rendering to simulate large-scale real-world data. Although 3D Gaussian Splatting (3DGS) has recently demonstrated high-quality results with real-time performance, it still faces challenges in indoor scenes with large, textureless regions, re

Cited by 38SourcecodeScholar
2025

ResGS: Residual Densification of 3D Gaussian for Efficient Detail Recovery

ICCV 2025poster

Recently, 3D Gaussian Splatting (3D-GS) has prevailed in novel view synthesis, achieving high fidelity and efficiency. However, it often struggles to capture rich details and complete geometry. Our analysis reveals that the 3D-GS densification operation lacks adaptiveness and faces a dilemma between…

Cited by 0SourcePDFScholar
2025

SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction

CVPR 2025poster

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion triplets, has opened new avenues for training, yielding promising r…

2025

Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios

ICRA 2025

Monocular depth estimation from RGB images plays a pivotal role in 3D vision. However, its accuracy can deteriorate in challenging environments such as nighttime or adverse weather conditions. While long-wave infrared cameras offer stable imaging in such challenging conditions, they are inherently l

Cited by 6SourceScholar
2024

DC-Gaussian: Improving 3D Gaussian Splatting for Reflective Dash Cam Videos

NeurIPS 2024poster

We present DC-Gaussian, a new method for generating novel views from in-vehicle dash cam videos. While neural rendering techniques have made significant strides in driving scenarios, existing methods are primarily designed for videos collected by autonomous vehicles. However, these videos are limite…

2024

Denoising Diffusion-Augmented Hybrid Video Anomaly Detection via Reconstructing Noised Frames

IJCAI 2024poster

Video Anomaly Detection (VAD) is crucial for enhancing security and surveillance systems through automatic identification of irregular events, thereby enabling timely responses and augmenting overall situational awareness. Although existing methods have achieved decent detection performances on benc…

Cited by 3SourcePDFScholar
2024

GaussianPro: 3D Gaussian Splatting with Progressive Propagation

ICML 2024poster

3D Gaussian Splatting (3DGS) has recently revolutionized the field of neural rendering with its high fidelity and efficiency. However, 3DGS heavily depends on the initialized point cloud produced by Structure-from-Motion (SfM) techniques. When tackling large-scale scenes that unavoidably contain tex…

2024

UC-NERF: Neural Radiance Field for Under-Calibrated Multi-View Cameras in Autonomous Driving

ICLR 2024poster

Multi-camera setups find widespread use across various applications, such as autonomous driving, as they greatly expand sensing capabilities. Despite the fast development of Neural radiance field (NeRF) techniques and their wide applications in both indoor and outdoor scenes, applying NeRF to multi…

Cited by 9SourcePDFScholar
2023

FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Models

ICCV 2023poster

3D scene reconstruction is a long-standing vision task. Existing approaches can be categorized into geometry-based and learning-based methods. The former leverages multi-view geometry but may face catastrophic failures due to the reliance on accurate pixel correspondence across views, while the latt…

Cited by 17PDFcodeScholar
2023

Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions

NeurIPS 2023poster

Perceiving and manipulating 3D articulated objects in diverse environments is essential for home-assistant robots. Recent studies have shown that point-level affordance provides actionable priors for downstream manipulation tasks. However, existing works primarily focus on single-object scenarios wi…

Cited by 25SourcePDFScholar
2023

Spatial-Temporal Graph Convolutional Network Boosted Flow-Frame Prediction For Video Anomaly Detection

ICASSP 2023accepted

Video Anomaly Detection (VAD) is a critical technology for intelligent surveillance systems and remains a challenging task in the signal processing community. An intuitive idea for VAD is to use a two-stream network to learn appearance and motion normality, respectively. However, existing approaches…

Cited by 0SourceScholar