← Search

Chi-Wing Fu

70 accepted papers

2026

EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

ICLR 2026poster

Robust 3D hand reconstruction is challenging in egocentric vision due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior works attempt to mitigate the challenges by scaling up training data or incorporating auxiliary cues, often falling short of effectively handling unse…

Cited by 0SourcecodeScholar
2026

IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

AAAI 2026technical

Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistency and coordinating multiple characters across different images. This paper prese

Cited by 0SourcePDFScholar
2026

Rethinking Intermediate Representation for VLM-based Robot Manipulation

CVPR 2026

Vision-Language Model (VLM) is now an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free gramm

Cited by 0SourceScholar
2025

COS3D: Collaborative Open-Vocabulary 3D Segmentation

NeurIPS 2025poster

Open-vocabulary 3D segmentation is a fundamental yet challenging task, requiring a mutual understanding of both segmentation and language. However, existing Gaussian-splatting-based methods rely either on a single 3D language field, leading to inferior segmentation, or on pre-computed class-agnostic…

Cited by 0SourceScholar
2025

EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights

CVPR 2025poster

Traffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints…

2025

Embodiment-agnostic Action Planning via Object-Part Scene Flow

ICRA 2025

Observing that the key for robotic action planning is to understand the target-object motion when its associated part is manipulated by the end effector, we propose to generate the 3D object-part scene flow and extract its transformations to solve the action trajectories for diverse embodiments. The

Cited by 7SourceScholar
2025

Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning

NeurIPS 2025poster

Generalized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency…

Cited by 0SourcecodeScholar
2025

HybridReg: Robust 3D Point Cloud Registration with Hybrid Motions

AAAI 2025technical

Scene-level point cloud registration is very challenging when considering dynamic foregrounds. Existing indoor datasets mostly assume rigid motions, so the trained models cannot robustly handle scenes with non-rigid motions. On the other hand, non-rigid datasets are mainly object-level, so the train…

2025

ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation

AAAI 2025technical

Controversial contents largely inundate the Internet, infringing various cultural norms and child protection standards. Traditional Image Content Moderation (ICM) models fall short in producing precise moderation decisions for diverse standards, while recent multimodal large language models (MLLMs),…

2025

MetaShadow: Object-Centered Shadow Detection, Removal, and Synthesis

CVPR 2025poster

Shadows are often underconsidered or even ignored in image editing applications, limiting the realism of the edited results. In this paper, we introduce MetaShadow, a three-in-one versatile framework that enables detection, removal, and controllable synthesis of shadows in natural images in an objec…

Cited by 2SourcePDFScholar
2025

Not-So-Optimal Transport Flows for 3D Point Cloud Generation

ICLR 2025poster

Learning generative models of 3D point clouds is one of the fundamental problems in 3D generative learning. One of the key properties of point clouds is their permutation invariance, i.e., changing the order of points in a point cloud does not change the shape they represent. In this paper, we analy…

Cited by 0SourcePDFScholar
2025

Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting

CVPR 2025poster

Lifting multi-view 2D instance segmentation to a radiance field has proven effective to enhance 3D understanding. Existing works rely on direct matching for end-to-end lifting, yielding inferior results, or employ a two-stage solution constrained by complex pre- or post-processing. In this work, we…

2025

UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation

CVPR 2025poster

Estimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly handle both scenarios and their performance degrades when appli…

2024

CRAYM: Neural Field Optimization via Camera RAY Matching

NeurIPS 2024poster

We introduce camera ray matching (CRAYM) into the joint optimization of camera poses and neural fields from multi-view images. The optimized field, referred to as a feature volume, can be “probed” by the camera rays for novel view synthesis (NVS) and 3D geometry reconstruction. One key reason for ma…

Cited by 1SourcePDFScholar
2024

HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object Interactions

CVPR 2024poster

Reconstructing 3D hand mesh robustly from a single image is very challenging due to the lack of diversity in existing real-world datasets. While data synthesis helps relieve the issue the syn-to-real gap still hinders its usage. In this work we present HandBooster a new approach to uplift the data d…

2024

Make-A-Shape: a Ten-Million-scale 3D Shape Model

ICML 2024poster

The progression in large-scale 3D generative models has been impeded by significant resource requirements for training and challenges like inefficient representations. This paper introduces Make-A-Shape, a novel 3D generative model trained on a vast scale, using 10 million publicly-available shapes.…

2024

PCF-Lift: Panoptic Lifting by Probabilistic Contrastive Fusion

ECCV 2024poster

"Panoptic lifting is an effective technique to address the 3D panoptic segmentation task by unprojecting 2D panoptic segmentations from multi-views to 3D scene. However, the quality of its results largely depends on the 2D segmentations, which could be noisy and error-prone, so its performance often…

2024

PPN-Pack: Placement Proposal Network for Efficient Robotic Bin Packing

RA-L 2024

Robotic bin packing is a challenging task, requiring compactly packing objects in a container and also efficiently performing the computation, such that the robot arm need not wait too long before taking action. In this work, we introduce PPN-Pack, a novel learning-based approach to improve the effi

Cited by 6SourceScholar
2024

PointRegGPT: Boosting 3D Point Cloud Registration using Generative Point-Cloud Pairs for Training

ECCV 2024poster

"Data plays a crucial role in training learning-based methods for 3D point cloud registration. However, the real-world dataset is expensive to build, while rendering-based synthetic data suffers from domain gaps. In this work, we present , boosting 3D Point cloud Registration using Generative Point-…

2024

SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation

AAAI 2024technical

Estimating 3D hand mesh from RGB images is a longstanding track, in which occlusion is one of the most challenging problems. Existing attempts towards this task often fail when the occlusion dominates the image space. In this paper, we propose SiMA-Hand, aiming to boost the mesh reconstruction perfo…

2024

Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models

ECCV 2024poster

"This paper addresses the limitations of adverse weather image restoration approaches trained on synthetic data when applied to real-world scenarios. We formulate a semi-supervised learning framework employing vision-language models to enhance restoration performance across diverse adverse weather c…

2024

Uncertainty-Aware Suction Grasping for Cluttered Scenes

RA-L 2024

In this work, we present a multi-stage pipeline that aims to accurately predict suction grasps for objects with varying properties in cluttered and complex scenes. Existing methods face difficulties in generalizing to unseen objects and effectively handling noisy depth/point cloud data, which often

Cited by 11SourcecodeScholar
2023

Command-Driven Articulated Object Understanding and Manipulation

CVPR 2023poster

We present Cart, a new approach towards articulated-object manipulations by human commands. Beyond the existing work that focuses on inferring articulation structures, we further support manipulating articulated shapes to align them subject to simple command templates. The key of Cart is to utilize…

2023

Deep Parametric 3D Filters for Joint Video Denoising and Illumination Enhancement in Video Super Resolution

AAAI 2023technical

Despite the quality improvement brought by the recent methods, video super-resolution (SR) is still very challenging, especially for videos that are low-light and noisy. The current best solution is to subsequently employ best models of video SR, denoising, and illumination enhancement, but doing so…

2023

DiffComplete: Diffusion-based Generative 3D Shape Completion

NeurIPS 2023poster

We introduce a new diffusion-based approach for shape completion on 3D range scans. Compared with prior deterministic and probabilistic methods, we strike a balance between realism, multi-modality, and high fidelity. We propose DiffComplete by casting shape completion as a generative task conditione…

Cited by 31SourcePDFScholar
2023

H2ONet: Hand-Occlusion-and-Orientation-Aware Network for Real-Time 3D Hand Mesh Reconstruction

CVPR 2023poster

Real-time 3D hand mesh reconstruction is challenging, especially when the hand is holding some object. Beyond the previous methods, we design H2ONet to fully exploit non-occluded information from multiple frames to boost the reconstruction quality. First, we decouple hand mesh reconstruction into tw…

2023

ISS: Image as Stepping Stone for Text-Guided 3D Shape Generation

ICLR 2023top-25%

Text-guided 3D shape generation remains challenging due to the absence of large paired text-shape dataset, the substantial semantic gap between these two modalities, and the structural complexity of 3D shapes. This paper presents a new framework called Image as Stepping Stone (ISS) for the task by i…

2023

On Improving Boundary Quality of Instance Segmentation in Cluttered and Chaotic Scenarios

ICRA 2023poster

Instance segmentation is a long-standing task for supporting robotic bin picking. However, objects of diverse classes can be closely packed with occlusions in cluttered and chaotic scenes, hence, even recent methods could have difficulty in locating clear and precise boundaries to distinguish nearby…

Cited by 1SourceScholar
2023

Prototypical Variational Autoencoder for 3D Few-shot Object Detection

NeurIPS 2023poster

Few-Shot 3D Point Cloud Object Detection (FS3D) is a challenging task, aiming to detect 3D objects of novel classes using only limited annotated samples for training. Considering that the detection performance highly relies on the quality of the latent features, we design a VAE-based prototype learn…

Cited by 8SourcePDFScholar
2023

SDF-Pack: Towards Compact Bin Packing with Signed-Distance-Field Minimization

IROS 2023poster

Robotic bin packing is very challenging, especially when considering practical needs such as object variety and packing compactness. This paper presents SDF-Pack, a new approach based on signed distance field (SDF) to model the geometric condition of objects in a container and compute the object pla…

Cited by 11SourcecodeScholar
2023

SILT: Shadow-Aware Iterative Label Tuning for Learning to Detect Shadows from Noisy Labels

ICCV 2023poster

Existing shadow detection datasets often contain missing or mislabeled shadows, which can hinder the performance of deep learning models trained directly on such data. To address this issue, we propose SILT, the Shadow-aware Iterative Label Tuning framework, which explicitly considers noise in shado…

Cited by 19PDFcodeScholar
2023

SIRA-PCR: Sim-to-Real Adaptation for 3D Point Cloud Registration

ICCV 2023poster

Point cloud registration is essential for many applications. However, existing real datasets require extremely tedious and costly annotations, yet may not provide accurate camera poses. For the synthetic datasets, they are mainly object-level, so the trained models may not generalize well to real sc…

Cited by 22PDFcodeScholar
2022

A Sim-to-Real Object Recognition and Localization Framework for Industrial Robotic Bin Picking

RA-L 2022

We present a generic and robust sim-to-real deep-learning-based framework, namely S2R-Pick, for fast and accurate object recognition and localization in industrial robotic bin picking. Unlike existing works designed for general everyday environments, objects for industrial bin picking are often text

Cited by 59SourceScholar
2022

Neural Template: Topology-Aware Reconstruction and Disentangled Generation of 3D Meshes

CVPR 2022poster

This paper introduces a novel framework called DT-Net for 3D mesh reconstruction and generation via Disentangled Topology. Beyond previous works, we learn a topology-aware neural template specific to each input then deform the template to reconstruct a detailed mesh while preserving the learned topo…

Cited by 39PDFcodeScholar
2022

SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin Picking

IROS 2022poster

Instance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any…

Cited by 2SourceScholar
2022

Sparse2Dense: Learning to Densify 3D Features for 3D Object Detection

NeurIPS 2022accept

LiDAR-produced point clouds are the major source for most state-of-the-art 3D object detectors. Yet, small, distant, and incomplete objects with sparse or few points are often hard to detect. We present Sparse2Dense, a new framework to efficiently boost 3D detection performance by learning to densif…

2022

TWIST: Two-Way Inter-Label Self-Training for Semi-Supervised 3D Instance Segmentation

CVPR 2022poster

We explore the way to alleviate the label-hungry problem in a semi-supervised setting for 3D instance segmentation. To leverage the unlabeled data to boost model performance, we present a novel Two-Way Inter-label Self-Training framework named TWIST. It exploits inherent correlations between semanti…

Cited by 29PDFcodeScholar
2022

Towards Robust Part-aware Instance Segmentation for Industrial Bin Picking

ICRA 2022poster

Industrial bin picking is a challenging task that requires accurate and robust segmentation of individual object instances. Particularly, industrial objects can have irregular shapes, that is, thin and concave, whereas in bin-picking scenarios, objects are often closely packed with strong occlusion.…

Cited by 15SourceScholar
2021

Accurate Grid Keypoint Learning for Efficient Video Prediction

IROS 2021poster

Video prediction methods generally consume substantial computing resources in training and deployment, among which keypoint-based approaches show promising improvement in efficiency by simplifying dense image prediction to light keypoint prediction. However, keypoint locations are often modeled only…

Cited by 18SourcecodeScholar
2021

CIA-SSD: Confident IoU-Aware Single-Stage Object Detector From Point Cloud

AAAI 2021technical

Existing single-stage detectors for locating objects in point clouds often treat object localization and category classification as separate tasks, so the localization accuracy and classification confidence may not well align. To address this issue, we present a new single-stage detector named the C…

2021

Guided Point Contrastive Learning for Semi-Supervised Point Cloud Semantic Segmentation

ICCV 2021poster

Rapid progress in 3D semantic segmentation is inseparable from the advances of deep network models, which highly rely on large-scale annotated data for training. To address the high cost and challenges of 3D point-level labeling, we present a method for semi-supervised point cloud semantic segmentat…

Cited by 160PDFScholar
2021

One Thing One Click: A Self-Training Approach for Weakly Supervised 3D Semantic Segmentation

CVPR 2021poster

Point cloud semantic segmentation often requires largescale annotated training data, but clearly, point-wise labels are too tedious to prepare. While some recent methods propose to train a 3D network with small percentages of point labels, we take the approach to an extreme and propose "One Thing On…

Cited by 171PDFcodeScholar
2021

SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud

CVPR 2021poster

We present Self-Ensembling Single-Stage object Detector (SE-SSD) for accurate and efficient 3D object detection in outdoor point clouds. Our key focus is on exploiting both soft and hard targets with our formulated constraints to jointly optimize the model, without introducing extra computation in t…

Cited by 460PDFcodeScholar
2021

Seeing Dynamic Scene in the Dark: A High-Quality Video Dataset With Mechatronic Alignment

ICCV 2021poster

Low-light video enhancement is an important task. Previous work is mostly trained on paired static images or videos. We compile a new dataset formed by our new strategy that contains high-quality spatially-aligned video pairs from dynamic scenes in low- and normal-light conditions. We built it using…

Cited by 117PDFcodeScholar
2021

Single-Stage Instance Shadow Detection With Bidirectional Relation Learning

CVPR 2021poster

Instance shadow detection aims to find shadow instances paired with the objects that cast the shadows. The previous work adopts a two-stage framework to first predict shadow instances, object instances, and shadow-object associations from the region proposals, then leverage a post-processing to matc…

Cited by 36PDFScholar
2020

Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization

ECCV 2020poster

The generalization capability of neural networks across domains is crucial for real-world applications. We argue that a generalized object recognition system should well understand the relationships among different images and also the images themselves at the same time. To this end, we present a new…

Cited by 238SourcePDFScholar
2020

PointAugment: An Auto-Augmentation Framework for Point Cloud Classification

CVPR 2020oral

We present PointAugment, a new auto-augmentation framework that automatically optimizes and augments point cloud samples to enrich the data diversity when we train a classification network. Different from existing auto-augmentation methods for 2D images, PointAugment is sample-aware and takes an adv…

Cited by 230PDFcodeScholar
2020

PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation

CVPR 2020oral

Instance segmentation is an important task for scene understanding. Compared to the fully-developed 2D, 3D instance segmentation for point clouds have much room to improve. In this paper, we present PointGroup, a new end-to-end bottom-up architecture, specifically focused on better grouping the poin…

Cited by 519PDFScholar
2019

Deep Floor Plan Recognition Using a Multi-Task Network With Room-Boundary-Guided Attention

ICCV 2019poster

This paper presents a new approach to recognize elements in floor plan layouts. Besides walls and rooms, we aim to recognize diverse floor plan elements, such as doors, windows and different types of rooms, in the floor layouts. To this end, we model a hierarchy of floor plan elements and design a d…

Cited by 159PDFcodeScholar
2019

Deep Multi-Model Fusion for Single-Image Dehazing

ICCV 2019poster

This paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural networ…

Cited by 146PDFScholar
2019

Hierarchical Point-Edge Interaction Network for Point Cloud Semantic Segmentation

ICCV 2019poster

We achieve 3D semantic scene labeling by exploring semantic relation between each point and its contextual neighbors through edges. Besides an encoder-decoder branch for predicting point labels, we construct an edge branch to hierarchically integrate point features and generate edge features. To inc…

Cited by 245PDFScholar
2019

PointWeb: Enhancing Local Neighborhood Features for Point Cloud Processing

CVPR 2019poster

This paper presents PointWeb, a new approach to extract contextual features from local neighborhood in a point cloud. Unlike previous work, we densely connect each point with every other in a local neighborhood, aiming to specify feature of each point based on the local region characteristics for be…

Cited by 953PDFcodeScholar
2019

Underexposed Photo Enhancement Using Deep Illumination Estimation

CVPR 2019oral

This paper presents a new neural network for enhancing underexposed photos. Instead of directly learning an image-to-image mapping as previous work, we introduce intermediate illumination in our network to associate the input with expected enhancement result, which augments the network's capability…

Cited by 1084PDFcodeScholar
2018

Bidirectional Feature Pyramid Network with Recurrent Attention Residual Modules for Shadow Detection

ECCV 2018poster

This paper presents a network to detect shadows by exploring and combining global context in deep layers and local context in shallow layers of a deep convolutional neural network (CNN). There are two technical contributions in our network design. First, we formulate the recurrent attention residual…

2018

Direction-Aware Spatial Context Features for Shadow Detection

CVPR 2018poster

Shadow detection is a fundamental and challenging task, since it requires an understanding of global image semantics and there are various backgrounds around shadows. This paper presents a novel network for shadow detection by analyzing image context in a direction-aware manner. To achieve this, we…

Cited by 484SourcePDFScholar
2018

EC-Net: an Edge-aware Point set Consolidation Network

ECCV 2018poster

Point clouds obtained from 3D scans are typically sparse, irregular, and noisy, and required to be consolidated. In this paper, we present the first deep learning based {em edge-aware} technique to facilitate the consolidation of point clouds. We design our network to process points grouped in local…

Cited by 340SourcePDFScholar
2018

PU-Net: Point Cloud Upsampling Network

CVPR 2018poster

Learning and analyzing 3D point clouds with deep networks is challenging due to the sparseness and irregularity of the data. In this paper, we present a data-driven point cloud upsampling technique. The key idea is to learn multi-level features per point and expand the point set via a multi-branch c…