← Search

Cheng-hao Kuo

24 accepted papers

2026

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

CVPR 2026

Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) have demonstrated promising capabilities in this domain, they primarily excel at object-level grounding and often struggle with part-level grounding--

Cited by 0SourceScholar
2026

Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation

ICRA 2026poster

Remembering where object segments were predicted in the past is useful for improving the accuracy and consistency of class-agnostic video segmentation algorithms. Existing video segmentation algorithms typically use either no object-level memory (e.g. FastSAM) or they use implicit memories in the fo…

2026

Revisiting Model Stitching In the Foundation Model Era

CVPR 2026

Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility. Prior work finds that models trained on the same dataset remain stitchable (negligible accuracy drop) despite differen

Cited by 0SourceScholar
2025

CSCPR: Cross-Source-Context Indoor RGB-D Place Recognition

RA-L 2025

We extend our previous work, PoCo (Liang et al. 2024), and present a new algorithm, Cross-Source-Context Place Recognition (CSCPR), for RGB-D indoor place recognition that integrates global retrieval and reranking into an end-to-end model and keeps the consistency of using Context-of-Clusters (CoCs)

Cited by 1SourceScholar
2025

Details Matter for Indoor Open-vocabulary 3D Instance Segmentation

ICCV 2025poster

Unlike closed-vocabulary 3D instance segmentation that is often trained end-to-end, open-vocabulary 3D instance segmentation (OV-3DIS) often leverages vision-language models (VLMs) to generate 3D instance proposals and classify them. While various concepts have been proposed from existing research,…

Cited by 0SourcePDFScholar
2025

Enhancing Single Image to 3D Generation using Gaussian Splatting and Hybrid Diffusion Priors

IROS 2025

3D object generation from a single unposed RGB image is essential for robotic perception, as reconstructing complete geometry and texture is essential for precise manipulation, grasping, and scene understanding, which is key for autonomous navigation and dexterous interaction. Recent advancements in

Cited by 2SourceScholar
2025

Explicit Memory Through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation

RA-L 2025

Remembering where object segments were predicted in the past is useful for improving the accuracy and consistency of class-agnostic video segmentation algorithms. Existing video segmentation algorithms typically use either no object-level memory (e.g. FastSAM) or they use implicit memories in the fo

Cited by 0SourceScholar
2025

Modeling Uncertainty in 3D Gaussian Splatting Through Continuous Semantic Splatting

ICRA 2025

In this paper, we present a novel algorithm for probabilistically updating and rasterizing semantic maps within 3D Gaussian Splatting (3D-GS). Although previous methods have introduced algorithms which learn to rasterize features in 3D-GS for enhanced scene understanding, 3D-GS can fail without warn

Cited by 14SourceScholar
2025

OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

ICCV 2025poster

Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduce OpenM3D, a novel open-vocabulary multi-view indoor 3D object detector trained without human annotations. In particular…

Cited by 0SourcePDFScholar
2025

POp-GS: Next Best View in 3D-Gaussian Splatting with P-Optimality

CVPR 2025poster

In this paper, we present a novel algorithm for quantifying uncertainty and information gained within 3D Gaussian Splatting (3D-GS) through P-Optimality. While 3D-GS has proven to be a useful world model with high-quality rasterizations, it does not natively quantify uncertainty or information, posi…

Cited by 0SourcePDFScholar
2025

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References

CVPR 2025poster

6D object pose estimation has shown strong generalizability to novel objects. However, existing methods often require either a complete, well-reconstructed 3D model or numerous reference images that fully cover the object. Estimating 6D poses from partial references, which capture only fragments of…

Cited by 0SourcePDFScholar
2025

Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression

CVPR 2025poster

Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compared to 3D-LMMs on 3D question answering tasks, largely due to the greater scale and…

Cited by 0SourcePDFScholar
2024

Configurable Embodied Data Generation for Class-Agnostic RGB-D Video Segmentation

RA-L 2024

This letter presents a method for generating large-scale datasets to improve class-agnostic video segmentation across robots with different form factors. Specifically, we consider the question of whether video segmentation models trained on generic segmentation data could be more effective for parti

Cited by 1SourceScholar
2024

Correspondence-Free SE(3) Point Cloud Registration in RKHS via Unsupervised Equivariant Learning

ECCV 2024poster

"This paper introduces a robust unsupervised SE(3) point cloud registration method that operates without requiring point correspondences. The method frames point clouds as functions in a reproducing kernel Hilbert space (RKHS), leveraging SE(3)-equivariant features for direct feature space registrat…

2024

Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding

ECCV 2024poster

"Self-supervised learning methods have demonstrated impressive performance across visual understanding tasks, including human behavior understanding. However, there has been limited work for self-supervised learning for egocentric social videos. Visual processing in such contexts faces several chall…

Cited by 1SourcePDFScholar
2024

GDA: Generalized Diffusion for Robust Test-time Adaptation

CVPR 2024poster

Machine learning models face generalization challenges when exposed to out-of-distribution (OOD) samples with unforeseen distribution shifts. Recent research reveals that for vision tasks test-time adaptation employing diffusion models can achieve state-of-the-art accuracy improvements on OOD sample…

Cited by 8SourcePDFScholar
2024

GenRC: Generative 3D Room Completion from Sparse Image Collections

ECCV 2024poster

"Sparse RGBD scene completion is a challenging task especially when considering consistent textures and geometries throughout the entire scene. Different from existing solutions that rely on human-designed text prompts or predefined camera trajectories, we propose , an automated training-free pipeli…

2024

No More Ambiguity in 360deg Room Layout via Bi-Layout Estimation

CVPR 2024poster

Inherent ambiguity in layout annotations poses significant challenges to developing accurate 360deg room layout estimation models. To address this issue we propose a novel Bi-Layout model capable of predicting two distinct layout types. One stops at ambiguous regions while the other extends to encom…

Cited by 5SourcePDFScholar
2024

PoCo: Point Context Cluster for RGBD Indoor Place Recognition

IROS 2024

We present a novel end-to-end algorithm (PoCo) for the indoor RGB-D place recognition task, aimed at identifying the most likely match for a given query frame within a reference database. The task presents inherent challenges attributed to the constrained field of view and limited range of perceptio

Cited by 2SourcecodeScholar
2023

Bidirectional Alignment for Domain Adaptive Detection with Transformers

ICCV 2023poster

We propose a Bidirectional Alignment for domain adaptive Detection with Transformers (BiADT) to improve cross domain object detection performance. Existing adversarial learning based methods use gradient reverse layer (GRL) to reduce the domain gap between the source and target domains in feature re…

Cited by 17PDFcodeScholar
2023

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

ICCV 2023poster

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-…

Cited by 11PDFcodeScholar
2023

SupeRGB-D: Zero-Shot Instance Segmentation in Cluttered Indoor Environments

RA-L 2023

Object instance segmentation is a key challenge for indoor robots navigating cluttered environments with many small objects. Limitations in 3D sensing capabilities often make it difficult to detect every possible object. While deep learning approaches may be effective for this problem, manually anno

Cited by 14SourcecodeScholar
2022

Learning Feature Decomposition for Domain Adaptive Monocular Depth Estimation

IROS 2022poster

Monocular depth estimation (MDE) has attracted intense study due to its low cost and critical functions for robotic tasks such as localization, mapping and obstacle detection. Supervised approaches have led to great success with the advance of deep learning, but they rely on large quantities of grou…

Cited by 16SourceScholar
2020

MEBOW: Monocular Estimation of Body Orientation in the Wild

CVPR 2020poster

Body orientation estimation provides crucial visual cues in many applications, including robotics and autonomous driving. It is particularly desirable when 3-D pose estimation is difficult to infer due to poor image resolution, occlusion or indistinguishable body parts. We present COCO-MEBOW (Monocu…

Cited by 41PDFcodeScholar