← Search

Yang Fu

22 accepted papers

2026

3D Aware Region Prompted Vision Language Model

ICLR 2026poster

We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flexible region prompting, allowing users to annotate regions with bounding boxes, segmentation masks on any frame, or direc…

Cited by 0SourcecodeScholar
2026

EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing

CVPR 2026

Video object removal aims to eliminate dynamic target objects and their visual effects, such as deformation, shadows, and reflections, while restoring seamless backgrounds. Recent diffusion-based video inpainting and object removal methods can remove the objects but often struggle to erase these eff

Cited by 0SourcecodeScholar
2026

TIPS: Turn-level Information-Potential Reward Shaping for Search-Augmented LLMs

ICLR 2026poster

Search-augmented large language models (LLMs) trained with reinforcement learning (RL) have achieved strong results on open-domain question answering (QA), but training still remains a significant challenge. The optimization is often unstable due to sparse rewards and difficult credit assignments ac…

Cited by 0SourcecodeScholar
2025

Learning Generalizable Feature Fields for Mobile Manipulation

IROS 2025

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while understanding fine-grained semantics, whereas the former involves capturing the comple

Cited by 49SourceScholar
2025

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

NeurIPS 2025poster

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality 1080P human speech videos w…

Cited by 0SourceScholar
2024

3D Reconstruction with Generalizable Neural Fields using Scene Priors

ICLR 2024poster

High-fidelity 3D scene reconstruction has been substantially advanced by recent progress in neural fields. However, most existing methods train a separate network from scratch for each individual scene. This is not scalable, inefficient, and unable to yield good results given limited views. While le…

2024

COLMAP-Free 3D Gaussian Splatting

CVPR 2024highlight

While neural rendering has led to impressive advances in scene reconstruction and novel view synthesis it relies heavily on accurately pre-computed camera poses. To relax this constraint multiple efforts have been made to train Neural Radiance Fields (NeRFs) without pre-processed camera poses. Howev…

2024

Cut out the Middleman: Revisiting Pose-based Gait Recognition

ECCV 2024poster

"Recent pose-based gait recognition methods, which utilize human skeletons as the model input, have demonstrated significant potential in handling variations in clothing and occlusions. However, methods relying on such skeleton to encode pose are constrained mainly by two problems: (1) poor performa…

2024

HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data

CVPR 2024poster

3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection process. In this paper we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a conditional diffusion model that takes both the 3D hand-obje…

2024

RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos

CVPR 2024poster

We introduce a new RGB-D object dataset captured in the wild called WildRGB-D. Unlike most existing real-world object-centric datasets which only come with RGB capturing the direct capture of the depth channel allows better 3D annotations and broader downstream applications. WildRGB-D comprises larg…

2024

SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models

NeurIPS 2024poster

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT) to enhance VLMs’ spatial perception and reasoning capabilit…

Cited by 61SourcePDFScholar
2023

Framewise Multiple Sound Source Localization and Counting Using Binaural Spatial Audio Signals

ICASSP 2023accepted

Sound source localization is the problem of estimating the positions of one or several sound sources. In terms of binaural audio, localization is a paramount perceptual characteristic which can be assessed subjectively or objectively. For objective evaluation of binaural sound localization, typical…

Cited by 0SourceScholar
2023

MonoNeRF: Learning Generalizable NeRFs from Monocular Videos without Camera Poses

ICML 2023poster

We propose a generalizable neural radiance fields - MonoNeRF, that can be trained on large-scale monocular videos of moving in static scenes without any ground-truth annotations of depth and camera poses. MonoNeRF follows an Autoencoder-based architecture, where the encoder estimates the monocular d…

2023

Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the Wild

ICLR 2023poster

While 6D object pose estimation has wide applications across computer vision and robotics, it remains far from being solved due to the lack of annotations. The problem becomes even more challenging when moving to category-level 6D pose, which requires generalization to unseen instances. Current appr…

2022

Category-Level 6D Object Pose Estimation in the Wild: A Semi-Supervised Learning Approach and A New Dataset

NeurIPS 2022accept

6D object pose estimation is one of the fundamental problems in computer vision and robotics research. While a lot of recent efforts have been made on generalizing pose estimation to novel object instances within the same category, namely category-level 6D pose estimation, it is still restricted in…

2022

DexMV: Imitation Learning for Dexterous Manipulation from Human Videos

ECCV 2022poster

"While in computer vision we have made significant progress on understanding hand-object interactions, it is still very challenging for robots to perform complex dexterous manipulation. In this paper, we propose a new platform and pipeline, DexMV (Dexterous Manipulation from Videos), for imitation l…

2021

CompFeat: Comprehensive Feature Aggregation for Video Instance Segmentation

AAAI 2021technical

Video instance segmentation is a complex task in which we need to detect, segment, and track each object for any given video. Previous approaches only utilize single-frame features for the detection, segmentation, and tracking of objects and they suffer in the video scenario due to several distinct…

2021

Learning to Track Instances without Video Annotations

CVPR 2021poster

Tracking segmentation masks of multiple instances has been intensively studied, but still faces two fundamental challenges: 1) the requirement of large-scale, frame-wise annotation, and 2) the complexity of two-stage approaches. To resolve these challenges, we introduce a novel semi-supervised frame…

Cited by 32PDFScholar
2019

Self-Similarity Grouping: A Simple Unsupervised Cross Domain Adaptation Approach for Person Re-Identification

ICCV 2019oral

Domain adaptation in person re-identification (re-ID) has always been a challenging task. In this work, we explore how to harness the similar natural characteristics existing in the samples from the target domain for learning to conduct person re-ID in an unsupervised manner. Concretely, we propose…

Cited by 606PDFcodeScholar