← Search

Yanwei Fu

127 accepted papers

2026

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

CVPR 2026

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising vision-language-action (VLA) paradigm. However, most existing approaches overlook

Cited by 0SourcecodeScholar
2026

CineBrain: A Large-Scale Multi-Modal Audiovisual Brain Dataset for Brain-Conditioned Video Generation

CVPR 2026

Most research decoding brain signals into images, often using them as priors for generative models, has focused only on visual content. This overlooks the brain's natural ability to integrate auditory and visual information, for instance, sound strongly influences how we perceive visual scenes. To i

Cited by 0SourcecodeScholar
2026

Conformal Reliability: A New Evaluation Metric for Conditional Generation

ICML 2026poster

Conditional generative models have recently achieved remarkable success in various applications. However, a suitable metric for evaluating the reliability of these models, which takes into account their inherent uncertainty, is still lacking. Existing metrics, which typically assess a single output,…

Cited by 0SourceScholar
2026

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among these modalities, sound provides indispensable cues about spat

Cited by 0SourcecodeScholar
2026

FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing

CVPR 2026

First-Frame Propagation (FFP) offers a promising paradigm for controllable video editing, but existing methods are hampered by a reliance on cumbersome run-time guidance. We identify the root cause of this limitation as the inadequacy of current training datasets, which are often too short, low-reso

Cited by 0SourcecodeScholar
2026

Harnessing Spectrum Video for Subject-Level Few-Shot and Cross-Montage EEG Generalization

ICML 2026poster

Existing EEG models are limited by electrode heterogeneity and rigid "channel-first" architectures that treat sensors as independent features. We propose Brain Signal Rendering (BSR), which reinterprets EEG as a physical projection of neural activity and transforms raw signals into geometry-aware Sp…

Cited by 0SourceScholar
2026

MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation

CVPR 2026

Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with only a few sparse RGB views and strict latency constraints. We introduce Multi-view 3D Referring Expression Segmentation

Cited by 0SourcecodeScholar
2026

NI-Tex: Non-isometric Image-based Garment Texture Generation

CVPR 2026

Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited. To acquire more realistic textures, generative methods are often used to extract Physically-based Rendering (PBR) textures and materials from large collections of wil

Cited by 1SourcecodeScholar
2026

OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-To-Robot Action Transfer

ICRA 2026poster

We present OCRA, an Object-Centric framework for video-based human-to-Robot Action transfer that learns directly from human demonstration videos to enable robust manipulation. Object-centric learning emphasizes task-relevant objects and their interactions while filtering out irrelevant background, p…

2026

One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow

AAAI 2026technical

We introduce a one-step generative policy for offline reinforcement learning that maps *noise* directly to *actions* via a *residual reformulation* of MeanFlow, making it compatible with Q-learning. While one-step Gaussian policies enable fast inference, they struggle to capture complex, multimodal

Cited by 0SourcePDFScholar
2026

ReWeaver: Towards Simulation-Ready and Topology-Accurate Garment Reconstruction

CVPR 2026

High-quality 3D garment reconstruction plays a crucial role in mitigating the sim-to-real gap in applications such as digital avatars, virtual try-on and robotic manipulation. However, existing garment reconstruction methods typically rely on unstructured representations, such as 3D Gaussian Splats,

Cited by 0SourcecodeScholar
2026

Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination

AAAI 2026technical

Reasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that rand

Cited by 0SourcePDFScholar
2026

Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery

ICRA 2026poster

We present a framework that integrates EEG-based visual and motor imagery (VI/MI) with robotic control to enable real-time, intention-driven grasping and placement. Motivated by the promise of BCI-driven robotics to enhance human-robot interaction, this system bridges neural signals with physical co…

2026

SCOOP'D: Learning Mixed-Liquid-Solid Scooping Via Sim2Real Generative Policy

ICRA 2026poster

Scooping items with tools such as spoons and ladles is common in daily life, ranging from assistive feeding to retrieving items from environmental disaster sites. However, developing a general and autonomous robotic scooping policy is challenging since it requires reasoning about complex tool-object…

2026

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual System

ICML 2026poster

Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. However, existing task-to-scene generation methods rely exclusively on large language models (LLMs) to predict scene layouts, inevitably yielding object c…

Cited by 0SourceScholar
2026

Spatially Guided Training for Vision-Language-Action Model

ICLR 2026poster

Large vision–language models (VLMs) excel at multimodal understanding but fall short when extended to embodied tasks, where instructions must be transformed into low-level motor actions. We introduce SP-VLA, a dual-system **V**ision–**L**anguage–**A**ction framework that leverages **S**patial **P**r…

Cited by 0SourceScholar
2026

Split Group Knockoffs: Controlling False Discovery Rate in Transformational Group Sparsity

ICML 2026poster

Controlling the false discovery rate (FDR) under complex sparsity structures remains a fundamental challenge in large language model (LLM) analysis. Motivated by multiple comparison problems in LLMs, we consider a setting in which sparsity arises at the group level after a linear transformation of m…

Cited by 0SourceScholar
2026

SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution Alignment

AAAI 2026technical

Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational overhead. While many distillation methods that are solely based on trajectory-preserving or distribution-matching have been

Cited by 0SourcePDFScholar
2026

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

RSS 2026poster

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA…

Cited by 0SourceScholar
2026

V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

CVPR 2026

Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across different viewpoints (e.g., ego-centric and exo-centric). This task poses significant challenges due to drastic viewpoint and

Cited by 0SourcecodeScholar
2026

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

CVPR 2026

Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a geome

Cited by 0SourcecodeScholar
2026

Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective

ICLR 2026poster

Post-training algorithms such as Supervised Fine-Tuning (SFT) and Reinforcement Fine-Tuning (RFT) are widely used to adapt multimodal large language models to downstream tasks. While effective at task adaptation, their impact on prior knowledge remains unclear. In this paper, we introduce jigsaw puz…

Cited by 0SourceScholar
2025

3D StreetUnveiler with Semantic-aware 2DGS - a simple baseline

ICLR 2025poster

Unveiling an empty street from crowded observations captured by in-car cameras is crucial for autonomous driving. However, removing all temporarily static objects, such as stopped vehicles and standing pedestrians, presents a significant challenge. Unlike object-centric 3D inpainting, which relies o…

Cited by 0SourcePDFScholar
2025

Adaptive Pruning of Pretrained Transformer via Differential Inclusions

ICLR 2025poster

Large transformers have demonstrated remarkable success, making it necessary to compress these models to reduce inference costs while preserving their performance. Current compression algorithms prune transformers at fixed compression ratios, requiring a unique pruning process for each ratio, which…

Cited by 0SourcePDFScholar
2025

Bivariate Causal Discovery with Proxy Variables: Integral Solving and Beyond

ICML 2025poster

Bivariate causal discovery is challenging when unmeasured confounders exist. To adjust for the bias, previous methods employed the proxy variable (*i.e.*, negative control outcome (NCO)) to test the treatment-outcome relationship through integral equations -- and assumed that violation of this equat…

Cited by 0SourcePDFScholar
2025

CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image

CVPR 2025highlight

This paper tackles category-level pose estimation of ar- ticulated objects in robotic manipulation tasks and intro- duces a new benchmark dataset. While recent methods es- timate part poses and sizes at the category level, they often rely on geometric cues and complex multi-stage pipelines that firs…

Cited by 0SourcePDFScholar
2025

Concept Replacer: Replacing Sensitive Concepts in Diffusion Models via Precision Localization

CVPR 2025poster

As large-scale diffusion models continue to advance, they excel at producing high-quality images but often generate unwanted content, such as sexually explicit or violent content. Existing methods for concept removal generally guide the image generation process but can unintentionally modify unrelat…

2025

Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification

ICASSP 2025accepted

Cloth-changing person re-identification aims at recognizing the same person with clothing changes across non-overlapping cameras. Advanced methods either resort to identity-related auxiliary modalities (e.g., sketches, silhouettes, and keypoints) or clothing labels to mitigate the impact of clothes.…

Cited by 0SourceScholar
2025

ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context

AAAI 2025technical

Visual storytelling involves generating a sequence of coherent frames from a textual storyline while maintaining consistency in characters and scenes. Existing autoregressive methods, which rely on previous frame-sentence pairs, struggle with high memory usage, slow generation speeds, and limited co…

2025

CrossVTON: Mimicking the Logic Reasoning on Cross-Category Virtual Try-On Guided by Tri-Zone Priors

IJCAI 2025

Despite remarkable progress in image-based virtual try-on systems, generating realistic and robust fitting images for cross-category virtual try-on remains a challenging task. The primary difficulty arises from the absence of human-like reasoning, which involves addressing size mismatches between ga

Cited by 0SourcePDFScholar
2025

CustAny: Customizing Anything from A Single Example

CVPR 2025poster

Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, is still challenging.Object customization, using reference images and textual descriptions, is key to addressing this iss…

2025

Diffusion-Based Imaginative Coordination for Bimanual Manipulation

ICCV 2025poster

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements. While video prediction has been recently studied for repres…

2025

Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

NeurIPS 2025poster

Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need…

Cited by 0SourcecodeScholar
2025

EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation

CVPR 2025poster

Handling complex or nonlinear motion patterns has long posed challenges for video frame interpolation. Although recent advances in diffusion-based methods offer improvements over traditional optical flow-based approaches, they still struggle to generate sharp, temporally consistent frames in scenari…

Cited by 3SourcePDFScholar
2025

HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation

CVPR 2025poster

Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made strides in gesture accuracy, challenges remain in generating diverse and cohere…

2025

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

CVPR 2025poster

We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly enhancing both generalization and 3D consistency in NVS. Our model…

2025

ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives

ICCV 2025poster

Bridging the gap between ego-centric and exo-centric views has been a long-standing question in computer vision. In this paper, we focus on the emerging Ego-Exo object correspondence task, which aims to understand object relations across ego-exo perspectives through segmentation. While numerous segm…

Cited by 0SourcePDFScholar
2025

PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching

NeurIPS 2025poster

Temporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersion of users. Despite its importance, this task remains challenging due to the difficulty in modeling long-term temporal…

Cited by 0SourcecodeScholar
2025

RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base

IROS 2025

Accurate 6D pose estimation is key for robotic manipulation, enabling precise object localization for tasks like grasping. We present RAG-6DPose, a retrieval-augmented approach that leverages 3D CAD models as a knowledge base by integrating both visual and geometric cues. Our RAG-6DPose roughly cont

Cited by 1SourcecodeScholar
2025

ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning

CVPR 2025poster

Open-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous robotics. However, current methods struggle because they rely…

Cited by 0SourcePDFScholar
2025

Sequential Multi-Object Grasping with One Dexterous Hand

IROS 2025

Sequentially grasping multiple objects with multi-fingered hands is common in daily life, where humans can fully leverage the dexterity of their hands to enclose multiple objects. However, the diversity of object geometries and the complex contact interactions required for high-DOF hands to grasp on

Cited by 6SourcecodeScholar
2025

Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

ICCV 2025poster

Visual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limiting their 3D spatial and 4D spatiotemporal awareness. Consequently, these methods struggle to capture the 3D structures an…

Cited by 0SourcePDFScholar
2025

TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making

NeurIPS 2025poster

In daily life, people often move through spaces to find objects that meet their needs, posing a key challenge in embodied AI. Traditional Demand-Driven Navigation (DDN) handles one need at a time but does not reflect the complexity of real-world tasks involving multiple needs and personal choices. T…

Cited by 0SourceScholar
2025

Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

CVPR 2025highlight

Recent advances in image inpainting increasingly use generative models to handle large irregular masks. However, these models can create unrealistic inpainted images due to two main issues: (1) Unwanted object insertion: Even with unmasked areas as context, generative models may still generate arbit…

2025

Towards Reliable and Holistic Visual In-Context Learning Prompt Selection

NeurIPS 2025poster

Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedded in in-context examples, which can be formulated as a global ranking problem of potential candidates. Current VICL meth…

Cited by 0SourceScholar
2025

TransSoft: The Low-Cost, Adaptable, and Radial Reconfigurable Soft Hand for Diverse Object Grasping

IROS 2025

This paper presents TransSoft, a novel soft robotic hand with a reconfigurable design for grasping objects of varying properties. While recent soft robotic hands have improved grasping capabilities, they often struggle with a limited range of manipulable object categories and tasks due to hardware c

Cited by 0SourceScholar
2025

UniDB: A Unified Diffusion Bridge Framework via Stochastic Optimal Control

ICML 2025spotlight

Recent advances in diffusion bridge models leverage Doob’s $h$-transform to establish fixed endpoints between distributions, demonstrating promising results in image translation and restoration tasks. However, these approaches frequently produce blurred or excessively smoothed image details and lack…

2025

You Only Estimate Once: Unified, One-stage, Real-Time Category-Level Articulated Object 6D Pose Estimation for Robotic Grasping

ICRA 2025

This paper addresses the problem of category-level pose estimation for articulated objects in robotic manipulation tasks. Recent works have shown promising results in estimating part pose and size at the category level. However, these approaches primarily follow a complex multi-stage pipeline that f

Cited by 4SourceScholar
2024

Adaptive Slot Attention: Object Discovery with Dynamic Slot Number

CVPR 2024poster

Object-centric learning (OCL) extracts the representation of objects with slots offering an exceptional blend of flexibility and interpretability for abstracting low-level perceptual features. A widely adopted method within OCL is slot attention which utilizes attention mechanisms to iteratively ref…

2024

Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector

ECCV 2024poster

"This paper studies the challenging cross-domain few-shot object detection (CD-FSOD), aiming to develop an accurate object detector for novel domains with minimal labeled examples. While transformer-based open-set detectors, such as DE-ViT, show promise in traditional few-shot object detection, thei…

2024

Doubly Robust Proximal Causal Learning for Continuous Treatments

ICLR 2024poster

Proximal causal learning is a powerful framework for identifying the causal effect under the existence of unmeasured confounders. Within this framework, the doubly robust (DR) estimator was derived and has shown its effectiveness in estimation, especially when the model assumption is violated. Howev…

2024

HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid Explorations

AAAI 2024technical

Existing gait recognition benchmarks mostly include minor clothing variations in the laboratory environments, but lack persistent changes in appearance over time and space. In this paper, we propose the first in-the-wild benchmark CCGait for cloth-changing gait recognition, which incorporates divers…

2024

LAC-Net: Linear-Fusion Attention-Guided Convolutional Network for Accurate Robotic Grasping Under the Occlusion

IROS 2024poster

This paper addresses the challenge of perceiving complete object shapes through visual perception. While prior studies have demonstrated encouraging outcomes in segmenting the visible parts of objects within a scene, amodal segmentation, in particular, has the potential to allow robots to infer the…

Cited by 1SourcecodeScholar
2024

LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model

CVPR 2024poster

This paper introduces LeftRefill an innovative approach to efficiently harness large Text-to-Image (T2I) diffusion models for reference-guided image synthesis. As the name implies LeftRefill horizontally stitches reference and target views together as a whole input. The reference image occupies the…

2024

MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing

NeurIPS 2024poster

Novel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which are discouraged from generalizing to challenging in-the-wild scenes and fail to be employed with 2D synthesis directly. M…

2024

MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo

ICLR 2024poster

Recent advancements in learning-based Multi-View Stereo (MVS) methods have prominently featured transformer-based models with attention mechanisms. However, existing approaches have not thoroughly investigated the profound influence of transformers on different MVS modules, resulting in limited dept…

2024

Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models

IROS 2024poster

This paper investigates the task of the open-ended interactive robotic manipulation on table-top scenarios. While recent Large Language Models (LLMs) enhance robots' comprehension of user instructions, their lack of visual grounding constrains their ability to physically interact with the environmen…

Cited by 4SourcecodeScholar
2024

T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image

ICASSP 2024accepted

Pixel2Mesh (P2M) is a classical approach for reconstructing 3D shapes from a single color image through coarse-to-fine mesh deformation. Although P2M is capable of generating plausible global shapes, its Graph Convolution Network (GCN) often produces overly smooth results, causing the loss of fine-g…

Cited by 0SourceScholar
2024

TaMMa: Target-driven Multi-subscene Mobile Manipulation

CoRL 2024poster

For everyday service robotics, the ability to navigate back and forth based on tasks in multi-subscene environments and perform delicate manipulations is crucial and highly practical. While existing robotics primarily focus on complex tasks within a single scene or simple tasks across scalable scene…

Cited by 2SourceScholar
2024

Test-Time Linear Out-of-Distribution Detection

CVPR 2024poster

Out-of-Distribution (OOD) detection aims to address the excessive confidence prediction by neural networks by triggering an alert when the input sample deviates significantly from the training distribution (in-distribution) indicating that the output may not be reliable. Current OOD detection approa…

2024

Towards Global Optimal Visual In-Context Learning Prompt Selection

NeurIPS 2024poster

Visual In-Context Learning (VICL) is a prevailing way to transfer visual foundation models to new tasks by leveraging contextual information contained in in-context examples to enhance learning and prediction of query sample. The fundamental problem in VICL is how to select the best prompt to activa…

Cited by 4SourcePDFScholar
2024

Unified Lexical Representation for Interpretable Visual-Language Alignment

NeurIPS 2024poster

Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clarity in its representation and similarity scores. On the other hand, lexical representation, a vector whose element repr…

2023

Causally-Aware Intraoperative Imputation for Overall Survival Time Prediction

CVPR 2023poster

Previous efforts in vision community are mostly made on learning good representations from visual patterns. Beyond this, this paper emphasizes the high-level ability of causal reasoning. We thus present a case study of solving the challenging task of Overall Survival (OS) time in primary liver cance…

Cited by 2SourcePDFScholar
2023

Coarse-to-Fine Amodal Segmentation with Shape Prior

ICCV 2023poster

Amodal object segmentation is a challenging task that involves segmenting both visible and occluded parts of an object. In this paper, we propose a novel approach, called Coarse-to-Fine Segmentation (C2F-Seg), that addresses this problem by progressively modeling the amodal segmentation. C2F-Seg…

Cited by 24PDFcodeScholar
2023

Grad-PU: Arbitrary-Scale Point Cloud Upsampling via Gradient Descent With Learned Distance Functions

CVPR 2023poster

Most existing point cloud upsampling methods have roughly three steps: feature extraction, feature expansion and 3D coordinate prediction. However, they usually suffer from two critical issues: (1) fixed upsampling rate after one-time training, since the feature expansion unit is customized for each…

2023

Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative Keypoints

ICCV 2023poster

Learning robust local image feature matching is a fundamental low-level vision task, which has been widely explored in the past few years. Recently, detector-free local feature matchers based on transformers have shown promising results, which largely outperform pure Convolutional Neural Network (CN…

Cited by 11PDFcodeScholar
2023

Language Guided Robotic Grasping with Fine-Grained Instructions

IROS 2023poster

Given a single RGB image and the attribute-rich language instructions, this paper investigates the novel problem of using Fine-grained instructions for the Language guided robotic Grasping (FLarG). This problem is made challenging by learning fine-grained language descriptions to ground target objec…

Cited by 12SourcecodeScholar
2023

Learning Versatile 3D Shape Generation with Improved Auto-regressive Models

ICCV 2023poster

Auto-Regressive (AR) models have achieved impressive results in 2D image generation by modeling joint distributions in the grid space. While this approach has been extended to the 3D domain for powerful shape generation, it still has two limitations: expensive computations on volumetric grids and am…

Cited by 1PDFScholar
2023

PourIt!: Weakly-Supervised Liquid Perception from a Single Image for Visual Closed-Loop Robotic Pouring

ICCV 2023poster

Liquid perception is critical for robotic pouring tasks. It usually requires the robust visual detection of flowing liquid. However, while recent works have shown promising results in liquid perception, they typically require labeled data for model training, a process that is both time-consuming and…

Cited by 6PDFScholar
2023

RankDNN: Learning to Rank for Few-Shot Learning

AAAI 2023technical

This paper introduces a new few-shot learning pipeline that casts relevance ranking for image retrieval as binary ranking relation classification. In comparison to image classification, ranking relation classification is sample efficient and domain agnostic. Besides, it provides a new perspective on…

2023

Rethinking Amodal Video Segmentation from Learning Supervised Signals with Object-centric Representation

ICCV 2023poster

Video amodal segmentation is a particularly challenging task in computer vision, which requires to deduce the full shape of an object from the visible parts of it. Recently, some studies have achieved promising performance by using motion flow to integrate information across frames under a self-supe…

Cited by 10PDFcodeScholar
2023

Rethinking Optical Flow From Geometric Matching Consistent Perspective

CVPR 2023poster

Optical flow estimation is a challenging problem remaining unsolved. Recent deep learning based optical flow models have achieved considerable success. However, these models often train networks from the scratch on standard optical flow data, which restricts their ability to robustly and geometrical…

2023

StyleAdv: Meta Style Adversarial Training for Cross-Domain Few-Shot Learning

CVPR 2023poster

Cross-Domain Few-Shot Learning (CD-FSL) is a recently emerging task that tackles few-shot learning across different domains. It aims at transferring prior knowledge learned on the source dataset to novel target datasets. The CD-FSL task is especially challenged by the huge domain gap between differe…

2023

Unsupervised Open-Vocabulary Object Localization in Videos

ICCV 2023poster

In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via a slot attention approach and then assig…

Cited by 7PDFcodeScholar
2022

DST: Dynamic Substitute Training for Data-Free Black-Box Attack

CVPR 2022poster

With the wide applications of deep neural network models in various computer vision tasks, more and more works study the model vulnerability to adversarial examples. For data-free black box attack scenario, existing methods are inspired by the knowledge distillation, and thus usually train a substit…

Cited by 22PDFcodeScholar
2022

Density-Preserving Deep Point Cloud Compression

CVPR 2022poster

Local density of point clouds is crucial for representing local details, but has been overlooked by existing point cloud compression methods. To address this, we propose a novel deep point cloud compression method that preserves local density information. Our method works in an auto-encoder fashion:…

Cited by 70PDFcodeScholar
2022

H4D: Human 4D Modeling by Learning Neural Compositional Representation

CVPR 2022poster

Despite the impressive results achieved by deep learning based 3D reconstruction, the techniques of directly learning to model 4D human captures with detailed geometry have been less studied. This work presents a novel framework that can effectively learn a compact and compositional representation f…

Cited by 26PDFScholar
2022

High-Fidelity Portrait Editing Via Exploring Differentiable Guided Sketches from the Latent Space

ICASSP 2022accepted

This paper studies the task of sketch-guided high-fidelity portrait editing. Advanced unconditional generators, such as StyleGAN, can generate a high-quality portrait image with great diversity. In previous researches, StyleGAN has successfully been utilized for color-guided image editing through la…

Cited by 0SourceScholar
2022

I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand Sketches

ICRA 2022poster

In this paper, we are interested in the problem of generating target grasps by understanding freehand sketches. The sketch is useful for the persons who cannot formulate language and the cases where a textual description is not available on the fly. However, very few works are aware of the usability…

Cited by 7SourceScholar
2022

Incremental Transformer Structure Enhanced Image Inpainting With Masking Positional Encoding

CVPR 2022poster

Image inpainting has made significant advances in recent years. However, it is still challenging to recover corrupted images with both vivid textures and reasonable structures. Some specific methods can only tackle regular textures while losing holistic structures due to the limited receptive fields…

Cited by 201PDFcodeScholar
2022

Learning 6-DoF Object Poses to Grasp Category-Level Objects by Language Instructions

ICRA 2022poster

This paper studies the task of any objects grasping from the known categories by free-form language instructions. This task demands the technique in computer vision, natural language processing, and robotics. We bring these disciplines together on this open challenge, which is essential to human-rob…

Cited by 23SourceScholar
2022

Learning To Memorize Feature Hallucination for One-Shot Image Generation

CVPR 2022poster

This paper studies the task of One-Shot image Generation (OSG), where generation network learned on base dataset should be generalizable to synthesize images of novel categories with only one available sample per novel category. Most existing methods for feature transfer in one-shot image generation…

Cited by 10PDFScholar
2022

LoRD: Local 4D Implicit Representation for High-Fidelity Dynamic Human Modeling

ECCV 2022poster

"Recent progress in 4D implicit representation focuses on globally controlling the shape and motion with low dimensional latent vectors, which is prone to missing surface details and accumulating tracking error. While many deep local representations have shown promising results for 3D shape modeling…

2022

ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-Wise Semantic Alignment and Generation

CVPR 2022oral

Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical application. In this work, we study a novel task on text-guided image manipulation on the entity level in the real world. Th…

Cited by 19PDFcodeScholar
2022

ONCE-3DLanes: Building Monocular 3D Lane Detection

CVPR 2022poster

We present ONCE-3DLanes, a real-world autonomous driving dataset with lane layout annotation in 3D space. Conventional 2D lane detection from a monocular image yields poor performance of following planning and control tasks in autonomous driving due to the case of uneven road. Predicting the 3D lane…

Cited by 76PDFcodeScholar
2022

RCLane: Relay Chain Prediction for Lane Detection

ECCV 2022poster

"Lane detection is an important component of many real-world autonomous systems. Despite a wide variety of lane detection approaches have been proposed, reporting steady benchmark improvements over time, lane detection remains a largely unsolved problem. This is because most of the existing lane det…

Cited by 32SourcePDFScholar
2022

SAR-Net: Shape Alignment and Recovery Network for Category-Level 6D Object Pose and Size Estimation

CVPR 2022poster

Given a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external real pose-annotated training data. Specifically, beyond the visual cues in RGB images, we rely on the shape information pr…

Cited by 83PDFScholar
2022

Self-supervised Amodal Video Object Segmentation

NeurIPS 2022accept

Amodal perception requires inferring the full shape of an object that is partially occluded. This task is particularly challenging on two levels: (1) it requires more information than what is contained in the instant retina or imaging sensor, (2) it is difficult to obtain enough well-annotated amoda…

Cited by 11SourcePDFScholar
2021

A Simple Feature Augmentation for Domain Generalization

ICCV 2021poster

The topical domain generalization (DG) problem asks trained models to perform well on an unseen target domain with different data statistics from the source training domains. In computer vision, data augmentation has proven one of the most effective ways of better exploiting the source data to impro…

Cited by 246PDFScholar
2021

Delving into Data: Effectively Substitute Training for Black-box Attack

CVPR 2021poster

Deep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, training a substitute model for adversarial attacks has attracted wide attention. Previous substitute training approaches…

Cited by 90PDFScholar
2021

Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object Detection

CVPR 2021poster

The objective of this paper is to learn context- and depth-aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message pro…

Cited by 156PDFcodeScholar
2021

Learning Compositional Representation for 4D Captures With Neural ODE

CVPR 2021poster

Learning based representation has become the key to the success of many computer vision systems. While many 3D representations have been proposed, it is still an unaddressed problem how to represent a dynamically changing 3D object. In this paper, we introduce a compositional representation for 4D c…

Cited by 33PDFScholar
2021

Learning Dynamic Alignment via Meta-Filter for Few-Shot Learning

CVPR 2021poster

Few-shot learning (FSL), which aims to recognise new classes by adapting the learned knowledge with extremely limited few-shot (support) examples, remains an important open problem in computer vision. Most of the existing methods for feature alignment in few-shot learning only consider image-level o…

Cited by 150PDFScholar
2021

Learning Salient Boundary Feature for Anchor-free Temporal Action Localization

CVPR 2021poster

Temporal action localization is an important yet challenging task in video understanding. Typically, such a task aims at inferring both the action category and localization of the start and end frame for each action instance in a long, untrimmed video. While most current models achieve good results…

Cited by 341PDFcodeScholar
2021

Learning a Few-shot Embedding Model with Contrastive Learning

AAAI 2021technical

Few-shot learning (FSL) aims to recognize target classes by adapting the prior knowledge learned from source classes. Such knowledge usually resides in a deep embedding model for a general matching purpose of the support and query image pairs. The objective of this paper is to repurpose the contrast…

Cited by 217SourcePDFScholar
2021

Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With Transformers

CVPR 2021poster

Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for se…

Cited by 4009PDFcodeScholar
2021

The Image Local Autoregressive Transformer

NeurIPS 2021poster

Recently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance compared to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to edit/change local image regions, may suffer from th…

Cited by 13SourcePDFScholar
2020

Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking

ECCV 2020poster

Existing Multiple-Object Tracking (MOT) methods either follow the tracking-by-detection paradigm to conduct object detection, feature extraction and data association separately, or have two of the three subtasks integrated to form a partially end-to-end solution. Going beyond these sub-optimal frame…

2020

DeepSFM: Structure From Motion Via Deep Bundle Adjustment

ECCV 2020poster

Structure from motion (SfM) is an essential computer vision problem which has not been well handled by deep learning. One of the promising trends is to apply explicit structural constraint, e.g. 3D cost volume, into the network. However, existing methods usually assume accurate camera poses either f…

Cited by 127SourcePDFScholar
2020

DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths

ICML 2020poster

Over-parameterization is ubiquitous nowadays in training neural networks to benefit both optimization in seeking global optima and generalization in reducing prediction error. However, compressive networks are desired in many real world applications and direct training of small networks may be trapp…

2020

FM2u-Net: Face Morphological Multi-Branch Network for Makeup-Invariant Face Verification

CVPR 2020poster

It is challenging in learning a makeup-invariant face verification model, due to (1) insufficient makeup/non-makeup face training pairs, (2) the lack of diverse makeup faces, and (3) the significant appearance changes caused by cosmetics. To address these challenges, we propose a unified Face Morpho…

Cited by 23PDFcodeScholar
2020

Neural Pose Transfer by Spatially Adaptive Instance Normalization

CVPR 2020poster

Pose transfer has been studied for decades, in which the pose of a source mesh is applied to a target mesh. Particularly in this paper, we are interested in transferring the pose of source human mesh to deform the target human mesh, while the source and target meshes may have different identity info…

Cited by 71PDFcodeScholar
2020

Sketch-BERT: Learning Sketch Bidirectional Encoder Representation From Transformers by Self-Supervised Learning of Sketch Gestalt

CVPR 2020poster

Previous researches of sketches often considered sketches in pixel format and leveraged CNN based models in the sketch understanding. Fundamentally, a sketch is stored as a sequence of data points, a vector format representation, rather than the photo-realistic image of pixels. SketchRNN studied a g…

Cited by 80PDFScholar
2019

Image Deformation Meta-Networks for One-Shot Learning

CVPR 2019oral

Humans can robustly learn novel visual concepts even when images undergo various deformations and loose certain information. Mimicking the same behavior and synthesizing deformed instances of new concepts may help visual recognition systems perform better one-shot learning, i.e., learning concepts f…

Cited by 303PDFcodeScholar
2019

Meta-Reinforced Synthetic Data for One-Shot Fine-Grained Visual Recognition

NeurIPS 2019poster

This paper studies the task of one-shot fine-grained recognition, which suffers from the problem of data scarcity of novel fine-grained classes. To alleviate this problem, a off-the-shelf image generator can be applied to synthesize additional images to help one-shot learning. However, such synthesi…

2018

MSplit LBI: Realizing Feature Selection and Dense Estimation Simultaneously in Few-shot and Zero-shot Learning

ICML 2018oral

It is one typical and general topic of learning a good embedding model to efficiently learn the representation coefficients between two spaces/subspaces. To solve this task, $L_{1}$ regularization is widely used for the pursuit of feature selection and avoiding overfitting, and yet the sparse estima…

Cited by 23SourcePDFScholar
2018

Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images

ECCV 2018poster

We propose an end-to-end deep learning architecture that produces a 3D shape in triangular mesh from a single color image. Limited by the nature of deep neural network, previous methods usually represent a 3D shape in volume or point cloud, and it is non-trivial to convert them to the more ready-to-…

Cited by 1707SourcePDFScholar
2018

Pose-Normalized Image Generation for Person Re-identification

ECCV 2018poster

Person Re-identification (re-id) faces two major challenges: the lack of cross-view paired training data and learning discriminative identity-sensitive and view-invariant features in the presence of large pose variations. In this work, we address both problems by proposing a novel deep person image…

2018

Stacked Semantics-Guided Attention Model for Fine-Grained Zero-Shot Learning

NeurIPS 2018poster

Zero-Shot Learning (ZSL) is generally achieved via aligning the semantic relationships between the visual features and the corresponding class semantic descriptions. However, using the global features to represent fine-grained images may lead to sub-optimal results since they neglect the discriminat…

2017

Multi-Scale Deep Learning Architectures for Person Re-Identification

ICCV 2017poster

Person Re-identification (re-id) aims to match people across non-overlapping camera views in a public space. It is a challenging problem because many people captured in surveillance videos wear similar clothes. Consequently, the differences in their appearance are often subtle and only detectable at…

Cited by 376PDFScholar
2016

Harnessing Object and Scene Semantics for Large-Scale Video Understanding

CVPR 2016spotlight

Large-scale action recognition and video categorization are important problems in computer vision. To address these problems, we propose a novel object- and scene-based semantic fusion network and representation. Our semantic fusion network combines three streams of information using a three-layer n…

Cited by 113PDFScholar