← Search

Min Sun

66 accepted papers

2026

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

CVPR 2026

Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) have demonstrated promising capabilities in this domain, they primarily excel at object-level grounding and often struggle with part-level grounding--

Cited by 0SourceScholar
2026

Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation

ICRA 2026poster

Remembering where object segments were predicted in the past is useful for improving the accuracy and consistency of class-agnostic video segmentation algorithms. Existing video segmentation algorithms typically use either no object-level memory (e.g. FastSAM) or they use implicit memories in the fo…

2026

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

CVPR 2026

There is growing interest in biomedical vision--language models trained on scientific literature. However, most pipelines compress rich multi-panel figures and long captions into coarse figure-level pairs, discarding the fine-grained correspondences clinicians rely on when zooming into local structu

Cited by 0SourceScholar
2026

Pointing at Parts: Training-Free Few-Shot Grounding in Multimodal LLMs

CVPR 2026

Part-level pointing is important for fine-grained interaction and reasoning, yet existing Multimodal Large Language Models (MLLMs) remain limited to instance-level pointing. Part-level pointing presents unique challenges: annotation is costly, parts are long-tail distributed, and many are difficult

Cited by 0SourceScholar
2026

Revisiting Model Stitching In the Foundation Model Era

CVPR 2026

Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility. Prior work finds that models trained on the same dataset remain stitchable (negligible accuracy drop) despite differen

Cited by 0SourceScholar
2026

The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space

ICML 2026poster

Diffusion-based Vision-Language-Action policies achieve remarkable success in robotic manipulation, yet commit a fundamental geometric error we term the \textbf{Euclidean Fallacy}: representing SE(3) poses as flat $\mathbb{R}^{12}$ vectors. This approximation induces (1) manifold drift violating SO(…

Cited by 0SourceScholar
2025

A Novel Wavy Soft Pneumatic Actuator Combining Variable Thickness and an Unconstrained Base

RA-L 2025

This letter proposes novel wavy soft pneumatic actuators (WSPAs) that integrates an unconstrained base with an elastic chamber of varying wall thicknesses. The base plate design eliminates bottom constraints, thereby enhancing the bending performance of WSPAs. Wavy elastomer cavities with varying wa

Cited by 4SourceScholar
2025

CSCPR: Cross-Source-Context Indoor RGB-D Place Recognition

RA-L 2025

We extend our previous work, PoCo (Liang et al. 2024), and present a new algorithm, Cross-Source-Context Place Recognition (CSCPR), for RGB-D indoor place recognition that integrates global retrieval and reranking into an end-to-end model and keeps the consistency of using Context-of-Clusters (CoCs)

Cited by 1SourceScholar
2025

Details Matter for Indoor Open-vocabulary 3D Instance Segmentation

ICCV 2025poster

Unlike closed-vocabulary 3D instance segmentation that is often trained end-to-end, open-vocabulary 3D instance segmentation (OV-3DIS) often leverages vision-language models (VLMs) to generate 3D instance proposals and classify them. While various concepts have been proposed from existing research,…

Cited by 0SourcePDFScholar
2025

ET-Former: Efficient Triplane Deformable Attention for 3D Semantic Scene Completion From Monocular Camera

IROS 2025

We introduce ET-Former, a novel end-to-end algorithm for semantic scene completion using a single monocular camera. Our approach generates a semantic occupancy map from single RGB observation while simultaneously providing uncertainty estimates for semantic predictions. By designing a triplane-based

Cited by 3SourcecodeScholar
2025

Enhancing Single Image to 3D Generation using Gaussian Splatting and Hybrid Diffusion Priors

IROS 2025

3D object generation from a single unposed RGB image is essential for robotic perception, as reconstructing complete geometry and texture is essential for precise manipulation, grasping, and scene understanding, which is key for autonomous navigation and dexterous interaction. Recent advancements in

Cited by 2SourceScholar
2025

Explicit Memory Through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation

RA-L 2025

Remembering where object segments were predicted in the past is useful for improving the accuracy and consistency of class-agnostic video segmentation algorithms. Existing video segmentation algorithms typically use either no object-level memory (e.g. FastSAM) or they use implicit memories in the fo

Cited by 0SourceScholar
2025

Modeling Uncertainty in 3D Gaussian Splatting Through Continuous Semantic Splatting

ICRA 2025

In this paper, we present a novel algorithm for probabilistically updating and rasterizing semantic maps within 3D Gaussian Splatting (3D-GS). Although previous methods have introduced algorithms which learn to rasterize features in 3D-GS for enhanced scene understanding, 3D-GS can fail without warn

Cited by 14SourceScholar
2025

OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

ICCV 2025poster

Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduce OpenM3D, a novel open-vocabulary multi-view indoor 3D object detector trained without human annotations. In particular…

Cited by 0SourcePDFScholar
2025

POp-GS: Next Best View in 3D-Gaussian Splatting with P-Optimality

CVPR 2025poster

In this paper, we present a novel algorithm for quantifying uncertainty and information gained within 3D Gaussian Splatting (3D-GS) through P-Optimality. While 3D-GS has proven to be a useful world model with high-quality rasterizations, it does not natively quantify uncertainty or information, posi…

Cited by 0SourcePDFScholar
2025

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References

CVPR 2025poster

6D object pose estimation has shown strong generalizability to novel objects. However, existing methods often require either a complete, well-reconstructed 3D model or numerous reference images that fully cover the object. Estimating 6D poses from partial references, which capture only fragments of…

Cited by 0SourcePDFScholar
2025

Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression

CVPR 2025poster

Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compared to 3D-LMMs on 3D question answering tasks, largely due to the greater scale and…

Cited by 0SourcePDFScholar
2024

Configurable Embodied Data Generation for Class-Agnostic RGB-D Video Segmentation

RA-L 2024

This letter presents a method for generating large-scale datasets to improve class-agnostic video segmentation across robots with different form factors. Specifically, we consider the question of whether video segmentation models trained on generic segmentation data could be more effective for parti

Cited by 1SourceScholar
2024

Context-Aware Replanning with Pre-Explored Semantic Map for Object Navigation

CoRL 2024poster

Pre-explored Semantic Map, constructed through prior exploration using visual language models (VLMs), has proven effective as a foundational element for training-free robotic applications. However, existing approaches assume the map's accuracy and do not provide effective mechanisms for revising dec…

Cited by 0SourceScholar
2024

Correspondence-Free SE(3) Point Cloud Registration in RKHS via Unsupervised Equivariant Learning

ECCV 2024poster

"This paper introduces a robust unsupervised SE(3) point cloud registration method that operates without requiring point correspondences. The method frames point clouds as functions in a reproducing kernel Hilbert space (RKHS), leveraging SE(3)-equivariant features for direct feature space registrat…

2024

Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding

ECCV 2024poster

"Self-supervised learning methods have demonstrated impressive performance across visual understanding tasks, including human behavior understanding. However, there has been limited work for self-supervised learning for egocentric social videos. Visual processing in such contexts faces several chall…

Cited by 1SourcePDFScholar
2024

GDA: Generalized Diffusion for Robust Test-time Adaptation

CVPR 2024poster

Machine learning models face generalization challenges when exposed to out-of-distribution (OOD) samples with unforeseen distribution shifts. Recent research reveals that for vision tasks test-time adaptation employing diffusion models can achieve state-of-the-art accuracy improvements on OOD sample…

Cited by 8SourcePDFScholar
2024

GenRC: Generative 3D Room Completion from Sparse Image Collections

ECCV 2024poster

"Sparse RGBD scene completion is a challenging task especially when considering consistent textures and geometries throughout the entire scene. Different from existing solutions that rely on human-designed text prompts or predefined camera trajectories, we propose , an automated training-free pipeli…

2024

No More Ambiguity in 360deg Room Layout via Bi-Layout Estimation

CVPR 2024poster

Inherent ambiguity in layout annotations poses significant challenges to developing accurate 360deg room layout estimation models. To address this issue we propose a novel Bi-Layout model capable of predicting two distinct layout types. One stops at ambiguous regions while the other extends to encom…

Cited by 5SourcePDFScholar
2024

PoCo: Point Context Cluster for RGBD Indoor Place Recognition

IROS 2024

We present a novel end-to-end algorithm (PoCo) for the indoor RGB-D place recognition task, aimed at identifying the most likely match for a given query frame within a reference database. The task presents inherent challenges attributed to the constrained field of view and limited range of perceptio

Cited by 2SourcecodeScholar
2023

Bidirectional Alignment for Domain Adaptive Detection with Transformers

ICCV 2023poster

We propose a Bidirectional Alignment for domain adaptive Detection with Transformers (BiADT) to improve cross domain object detection performance. Existing adversarial learning based methods use gradient reverse layer (GRL) to reduce the domain gap between the source and target domains in feature re…

Cited by 17PDFcodeScholar
2023

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

ICCV 2023poster

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-…

Cited by 11PDFcodeScholar
2023

MixFairFace: Towards Ultimate Fairness via MixFair Adapter in Face Recognition

AAAI 2023technical

Although significant progress has been made in face recognition, demographic bias still exists in face recognition systems. For instance, it usually happens that the face recognition performance for a certain demographic group is lower than the others. In this paper, we propose MixFairFace framework…

2022

360-DFPE: Leveraging Monocular 360-Layouts for Direct Floor Plan Estimation

RA-L 2022

We present 360-DFPE, a sequential floor plan estimation method that directly takes 360-images as input without relying on active sensors or 3D information. Our approach leverages a loosely coupled integration between a monocular visual SLAM solution and a monocular 360-room layout approach, which es

Cited by 15SourcecodeScholar
2022

360-MLC: Multi-view Layout Consistency for Self-training and Hyper-parameter Tuning

NeurIPS 2022accept

We present 360-MLC, a self-training method based on multi-view layout consistency for finetuning monocular room-layout models using unlabeled 360-images only. This can be valuable in practical scenarios where a pre-trained model needs to be adapted to a new data domain without using any ground truth…

2022

Autoregressive 3D Shape Generation via Canonical Mapping

ECCV 2022poster

"With the capacity of modeling long-range dependencies in sequential data, transformers have shown remarkable performances in a variety of generative tasks such as image, audio, and text generation. Yet, taming them in generating less structured and voluminous data formats such as high-resolution po…

2022

CC-3DT: Panoramic 3D Object Tracking via Cross-Camera Fusion

CoRL 2022poster

To track the 3D locations and trajectories of the other traffic participants at any given time, modern autonomous vehicles are equipped with multiple cameras that cover the vehicle's full surroundings. Yet, camera-based 3D object tracking methods prioritize optimizing the single-camera setup and res…

Cited by 32SourceScholar
2022

Direct Voxel Grid Optimization: Super-Fast Convergence for Radiance Fields Reconstruction

CVPR 2022oral

We present a super-fast convergence approach to reconstructing the per-scene radiance field from a set of images that capture the scene with known poses. This task, which is often applied to novel view synthesis, is recently revolutionized by Neural Radiance Field (NeRF) for its state-of-the-art qua…

Cited by 1232PDFcodeScholar
2021

Indoor Panorama Planar 3D Reconstruction via Divide and Conquer

CVPR 2021poster

Indoor panorama typically consists of human-made structures parallel or perpendicular to gravity. We leverage this phenomenon to approximate the scene in a 360-degree image with (H)orizontal-planes and (V)ertical-planes. To this end, we propose an effective divide-and-conquer strategy that divides p…

Cited by 16PDFcodeScholar
2021

LED2-Net: Monocular 360deg Layout Estimation via Differentiable Depth Rendering

CVPR 2021poster

Although significant progress has been made in room layout estimation, most methods aim to reduce the loss in the 2D pixel coordinate rather than exploiting the room structure in the 3D space. Towards reconstructing the room layout in 3D, we formulate the task of 360 layout estimation as a problem o…

Cited by 50PDFScholar
2021

Learning 3D Dense Correspondence via Canonical Point Autoencoder

NeurIPS 2021poster

We propose a canonical point autoencoder (CPAE) that predicts dense correspondences between 3D shapes of the same category. The autoencoder performs two key functions: (a) encoding an arbitrarily ordered point cloud to a canonical primitive, e.g., a sphere, and (b) decoding the primitive back to the…

Cited by 28SourcePDFScholar
2021

Robust 360-8PA: Redesigning The Normalized 8-point Algorithm for 360-FoV Images

ICRA 2021poster

In this paper, we present a novel preconditioning strategy for the classic 8-point algorithm (8-PA) for estimating an essential matrix from 360-FoV images (i.e., equirectangular images) in spherical projection. To alleviate the effect of uneven key-feature distributions and outlier correspondences,…

Cited by 7SourcecodeScholar
2021

Specialize and Fuse: Pyramidal Output Representation for Semantic Segmentation

ICCV 2021poster

We present a novel pyramidal output representation to ensure parsimony with our "specialize and fuse" process for semantic segmentation. A pyramidal "output" representation consists of coarse-to-fine levels, where each level is "specialize" in a different class distribution (e.g., more stuff than th…

Cited by 9PDFScholar
2020

360SD-Net: 360° Stereo Depth Estimation with Learnable Cost Volume

ICRA 2020poster

Recently, end-to-end trainable deep neural networks have significantly improved stereo depth estimation for perspective images. However, 360° images captured under equirectangular projection cannot benefit from directly adopting existing methods due to distortion introduced (i.e., lines in 3D are no…

Cited by 82SourcecodeScholar
2020

BiFuse: Monocular 360 Depth Estimation via Bi-Projection Fusion

CVPR 2020poster

Depth estimation from a monocular 360 image is an emerging problem that gains popularity due to the availability of consumer-level 360 cameras and the complete surrounding sensing capability. While the standard of 360 imaging is under rapid development, we propose to predict the depth map of a monoc…

Cited by 229PDFcodeScholar
2020

Mitigating Forgetting in Online Continual Learning via Instance-Aware Parameterization

NeurIPS 2020poster

Online continual learning is a challenging scenario where a model needs to learn from a continuous stream of data without revisiting any previously encountered data instances. The phenomenon of catastrophic forgetting is worsened since the model should not only address the forgetting at the task-lev…

Cited by 50SourcePDFScholar
2019

3D LiDAR and Stereo Fusion using Stereo Matching Network with Conditional Cost Volume Normalization

IROS 2019poster

The complementary characteristics of active and passive depth sensing techniques motivate the fusion of the LiDAR sensor and stereo camera for improved depth perception. Instead of directly fusing estimated depths across LiDAR and stereo modalities, we take advantages of the stereo matching network…

Cited by 54SourceScholar
2019

DuLa-Net: A Dual-Projection Network for Estimating Room Layouts From a Single RGB Panorama

CVPR 2019poster

We present a deep learning framework, called DuLa-Net, to predict Manhattan-world 3D room layouts from a single RGB panorama. To achieve better prediction accuracy, our method leverages two projections of the panorama at once, namely the equirectangular panorama-view and the perspective ceiling-vie…

Cited by 179PDFScholar
2019

HorizonNet: Learning Room Layout With 1D Representation and Pano Stretch Data Augmentation

CVPR 2019poster

We present a new approach to the problem of estimating the 3D room layout from a single panoramic image. We represent room layout as three 1D vectors that encode, at each image column, the boundary positions of floor-wall and ceiling-wall, and the existence of wall-wall boundary. The proposed networ…

Cited by 231PDFcodeScholar
2019

Joint Monocular 3D Vehicle Detection and Tracking

ICCV 2019poster

Vehicle 3D extents and trajectories are critical cues for predicting the future location of vehicles and planning future agent ego-motion based on those predictions. In this paper, we propose a novel online framework for 3D vehicle detection and tracking from monocular videos. The framework can not…

Cited by 284PDFScholar
2019

Plug-and-Play: Improve Depth Prediction via Sparse Data Propagation

ICRA 2019poster

We propose a novel plug-and-play (PnP) module for improving depth prediction with taking arbitrary patterns of sparse depths as input. Given any pre-trained depth prediction model, our PnP module updates the intermediate feature map such that the model outputs new depths consistent with the given sp…

Cited by 25SourceScholar
2018

Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos

CVPR 2018poster

Automatic saliency prediction in 360° videos is critical for viewpoint guidance applications (e.g., Facebook 360 Guide). We propose a spatial-temporal network which is (1) unsupervisedly trained and (2) tailor-made for 360° viewing sphere. Note that most existing methods are less scalable since they…

Cited by 240SourcePDFScholar
2018

DLWV2: A Deep Learning-Based Wearable Vision-System with Vibrotactile-Feedback for Visually Impaired People to Reach Objects

IROS 2018poster

We develop a Deep Learning-based Wearable Vision-system with Vibrotactile-feedback (DLWV2)to guide Blind and Visually Impaired (BVI)people to reach objects. The system achieves high accuracy in object detection and tracking in 3-D using an extended deep learning-based 2.5-D detector and a 3-D object…

Cited by 19SourceScholar
2018

DPP-Net: Device-aware Progressive Search for Pareto-optimal Neural Architectures

ECCV 2018poster

Recent breakthroughs in Neural Architectural Search (NAS) have achieved state-of-the-art performances in applications such as image classification and language modeling. However, these techniques typically ignore device-related objectives such as inference time, memory usage, and power consumption.…

Cited by 272SourcePDFScholar
2018

Efficient Uncertainty Estimation for Semantic Segmentation in Videos

ECCV 2018poster

Uncertainty estimation in deep learning becomes more important recently. A deep learning model can't be applied in real applications if we don't know whether the model is certain about the decision or not. Some literature proposes the Bayesian neural network which can estimate the uncertainty by Mon…

Cited by 134SourcePDFScholar
2018

Leveraging Motion Priors in Videos for Improving Human Segmentation

ECCV 2018poster

Despite many advances in deep-learning based semantic segmentation, performance drop due to distribution mismatch is often encountered in the real world. Recently, a few domain adaptation and active learning approaches have been proposed to mitigate the performance drop. However, very little attenti…

Cited by 1SourcePDFScholar
2018

Liquid Pouring Monitoring via Rich Sensory Inputs

ECCV 2018poster

Humans have the amazing ability to perform very subtle manipulation task using a closed-loop control system with imprecise mechanics (i.e., our body parts) but rich sensory information (e.g., vision, tactile, etc.). In the closed-loop system, the ability to monitor the state of the task via rich sen…

Cited by 9SourcePDFScholar
2018

Omnidirectional CNN for Visual Place Recognition and Navigation

ICRA 2018poster

Visual place recognition is challenging, especially when only a few place exemplars are given. To mitigate the challenge, we consider place recognition method using omnidirectional cameras and propose a novel Omnidirectional Convolutional Neural Network (O-CNN) to handle severe camera pose variation…

Cited by 88SourceScholar
2017

Agent-Centric Risk Assessment: Accident Anticipation and Risky Region Localization

CVPR 2017spotlight

For survival, a living agent (e.g., human in Fig. 1(a)) must have the ability to assess risk (1) by temporally anticipating accidents before they occur (Fig. 1(b)), and (2) by spatially localizing risky regions (Fig. 1(c)) in the environment to move away from threats. In this paper, we take an agent…

Cited by 89PDFScholar
2017

Anticipating Daily Intention Using On-Wrist Motion Triggered Sensing

ICCV 2017spotlight

Anticipating human intention by observing one's actions has many applications. For instance, picking up a cellphone, then a charger (actions) implies that one wants to charge the cellphone (intention). By anticipating the intention, an intelligent system can guide the user to the closest power outle…

Cited by 32PDFcodeScholar
2017

Deep 360 Pilot: Learning a Deep Agent for Piloting Through 360deg Sports Videos

CVPR 2017oral

Watching a 360* sports video requires a viewer to continuously select a viewing angle, either through a sequence of mouse clicks or head movements. To relieve the viewer from this "360 piloting" task, we propose "deep 360 pilot" - a deep learning-based agent for piloting through 360* sports videos a…

Cited by 29PDFScholar
2017

No More Discrimination: Cross City Adaptation of Road Scene Segmenters

ICCV 2017poster

Despite the recent success of deep-learning based semantic segmentation, deploying a pre-trained road scene segmenter to a city whose images are not presented in the training set would not achieve satisfactory performance due to dataset biases. Instead of collecting a large number of annotated image…

Cited by 410PDFScholar
2017

Show, Adapt and Tell: Adversarial Training of Cross-Domain Image Captioner

ICCV 2017poster

Impressive image captioning results are achieved in domains with plenty of training image and sentence pairs (e.g., MSCOCO). However, transferring to a target domain with significant domain shifts but no paired training data (referred to as cross-domain image captioning) remains largely unexplored.…

Cited by 183PDFcodeScholar
2017

Visual Forecasting by Imitating Dynamics in Natural Sequences

ICCV 2017spotlight

We introduce a general framework for visual forecasting, which directly imitates visual sequences without additional supervision. As a result, our model can be applied at several semantic levels and does not require any domain knowledge or handcrafted features. We achieve this by formulating visual…

Cited by 76PDFScholar