← Search

Aming WU

25 accepted papers

2026

Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation

ICML 2026poster

Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge from seen tasks. Recent in-context learning approaches leverage seen task demonstrations to generate actions for unseen tasks without parameter updates…

Cited by 0SourceScholar
2026

Geometric-Aware Hypergraph Reasoning for Novel Class Discovery in Point Cloud Segmentation

CVPR 2026

Novel Class Discovery in Point Cloud Segmentation is recently proposed, aiming to leverage knowledge from known classes to automatically segment unlabeled classes within point clouds. The core of this task lies in leveraging the geometric and semantic knowledge of multiple known classes to achieve s

Cited by 0SourcecodeScholar
2026

Simulating Distribution Dynamics: Liquid Temporal Feature Evolution for Single-Domain Generalized Object Detection

AAAI 2026technical

In this paper, we focus on Single-Domain Generalized Object Detection (Single-DGOD), aiming to transfer a detector trained on one source domain to multiple unknown domains. Existing methods for Single-DGOD typically rely on discrete data augmentation or static perturbation methods to expand data div

Cited by 0SourcePDFScholar
2026

Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning

CVPR 2026

Vision-Language Navigation (VLN) aims to enable agents to navigate to a target location based on language instructions. Traditional VLN often follows a close-set assumption, i.e., training and test data share the same style of the input images and instructions. However, the real world is open and fi

Cited by 0SourceScholar
2025

CFD: Learning Generalized Molecular Representation via Concept-Enhanced Feedback Disentanglement

ICLR 2025poster

To accelerate biochemical research, e.g., drug and protein discovery, molecular representation learning (MRL) has attracted much attention. However, most existing methods follow the closed-set assumption that training and testing data share identical distribution, which limits their generalization a…

2025

Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios

ICCV 2025poster

In practice, environments constantly change over time and space, posing significant challenges for object detectors trained based on a closed-set assumption, i.e., training and test data share the same distribution. To this end, continual test-time adaptation has attracted much attention, aiming to…

Cited by 0SourcePDFScholar
2025

Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and Reasoning

NeurIPS 2025poster

In this paper, we focus on Novel Class Discovery for Point Cloud Segmentation (3D-NCD), aiming to learn a model that can segment unlabeled (novel) 3D classes using only the supervision from labeled (base) 3D classes. The key to this task is to setup the exact correlations between the point represent…

Cited by 0SourceScholar
2025

Percept, Memory, and Imagine: World Feature Simulating for Open-Domain Unknown Object Detection

CVPR 2025poster

To accelerate the safe deployment of object detectors, we focus on reducing the impact of both covariate and semantic shifts. And we consider a realistic yet challenging scenario, namely Open-Domain Unknown Object Detection (ODU-OD), which aims to detect unknown objects in unseen target domains with…

2025

Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Grounding

CVPR 2025poster

This paper pays attention to Weakly Supervised Affordance Grounding (WSAG) task that aims to train model to identify affordance regions using human-object interaction images and egocentric images without the need for costly pixel-level annotations. Most existing methods usually consider the affordan…

Cited by 0SourcePDFScholar
2025

Style Evolving along Chain-of-Thought for Unknown-Domain Object Detection

CVPR 2025highlight

Recently, a task of Single-Domain Generalized Object Detection (Single-DGOD) is proposed, aiming to generalize a detector to multiple unknown domains never seen before during training. Due to the unavailability of target-domain data, some methods leverage the multimodal capabilities of vision-langu…

2025

VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual Grounding

ICCV 2025poster

As an important direction of embodied intelligence, 3D Visual Grounding has attracted much attention, aiming to identify 3D objects matching the given language description. Most existing methods often follow a two-stage process, i.e., first detecting proposal objects and identifying the right object…

Cited by 0SourcePDFScholar
2025

Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph Generation

ICCV 2025poster

To promote the deployment of scenario understanding in the real world, Open-Vocabulary Scene Graph Generation (OV-SGG) has attracted much attention recently, aiming to generalize beyond the limited number of relation categories labeled during training and detect those unseen relations during inferen…

2024

Modulated Phase Diffusor: Content-Oriented Feature Synthesis for Detecting Unknown Objects

ICLR 2024poster

To promote the safe deployment of object detectors, a task of unsupervised out-of-distribution object detection (OOD-OD) is recently proposed, aiming to detect unknown objects during training without reliance on any auxiliary OOD data. To alleviate the impact of lacking OOD data, for this task, one…

2024

Prompt-Driven Dynamic Object-Centric Learning for Single Domain Generalization

CVPR 2024poster

Single-domain generalization aims to learn a model from single source domain data attaining generalized performance on other unseen target domains. Existing works primarily focus on improving the generalization ability of static networks. However static networks are unable to dynamically adapt to th…

Cited by 12SourcePDFScholar
2023

Discriminating Known From Unknown Objects via Structure-Enhanced Recurrent Variational AutoEncoder

CVPR 2023poster

Discriminating known from unknown objects is an important essential ability for human beings. To simulate this ability, a task of unsupervised out-of-distribution object detection (OOD-OD) is proposed to detect the objects that are never-seen-before during model training, which is beneficial for pro…

2023

Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph Generation

ICCV 2023poster

The scene graph generation (SGG) task is designed to identify the predicates based on the subject-object pairs. However, existing datasets generally include two imbalance cases: one is the class imbalance from the predicted predicates and another is the context imbalance from the given subject-objec…

Cited by 13PDFcodeScholar
2022

Divide and Conquer: Compositional Experts for Generalized Novel Class Discovery

CVPR 2022poster

In response to the explosively-increasing requirement of annotated data, Novel Class Discovery (NCD) has emerged as a promising alternative to automatically recognize unknown classes without any annotation. To this end, a model makes use of a base set to learn basic semantic discriminability that ca…

Cited by 51PDFcodeScholar
2022

Single-Domain Generalized Object Detection in Urban Scene via Cyclic-Disentangled Self-Distillation

CVPR 2022poster

In this paper, we are concerned with enhancing the generalization capability of object detectors. And we consider a realistic yet challenging scenario, namely Single-Domain Generalized Object Detection (Single-DGOD), which aims to learn an object detector that performs well on many unseen target dom…

Cited by 101PDFcodeScholar
2021

Domain-Smoothing Network for Zero-Shot Sketch-Based Image Retrieval

IJCAI 2021poster

Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a novel cross-modal retrieval task, where abstract sketches are used as queries to retrieve natural images under zero-shot scenario. Most existing methods regard ZS-SBIR as a traditional classification problem and employ a cross-entropy or triplet-…

2021

Generalized and Discriminative Few-Shot Object Detection via SVD-Dictionary Enhancement

NeurIPS 2021poster

Few-shot object detection (FSOD) aims to detect new objects based on few annotated samples. To alleviate the impact of few samples, enhancing the generalization and discrimination abilities of detectors on new objects plays an important role. In this paper, we explore employing Singular Value Decomp…

2021

Vector-Decomposed Disentanglement for Domain-Invariant Object Detection

ICCV 2021poster

To improve the generalization of detectors, for domain adaptive object detection (DAOD), recent advances mainly explore aligning feature-level distributions between the source and single-target domain, which may neglect the impact of domain-specific information existing in the aligned features. Towa…

Cited by 136PDFcodeScholar
2020

Bidirectional Adversarial Training for Semi-Supervised Domain Adaptation

IJCAI 2020poster

Semi-supervised domain adaptation (SSDA) is a novel branch of machine learning that scarce labeled target examples are available, compared with unsupervised domain adaptation. To make effective use of these additional data so as to bridge the domain gap, one possible way is to generate adversarial e…

Cited by 0SourcePDFScholar
2019

Connective Cognition Network for Directional Visual Commonsense Reasoning

NeurIPS 2019poster

Visual commonsense reasoning (VCR) has been introduced to boost research of cognition-level visual understanding, i.e., a thorough understanding of correlated details of the scene plus an inference with related commonsense knowledge. Recent studies on neuroscience have suggested that brain function…