← Search

Yahong Han

24 accepted papers

2026

Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation

ICML 2026poster

Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge from seen tasks. Recent in-context learning approaches leverage seen task demonstrations to generate actions for unseen tasks without parameter updates…

Cited by 0SourceScholar
2026

Geometric-Aware Hypergraph Reasoning for Novel Class Discovery in Point Cloud Segmentation

CVPR 2026

Novel Class Discovery in Point Cloud Segmentation is recently proposed, aiming to leverage knowledge from known classes to automatically segment unlabeled classes within point clouds. The core of this task lies in leveraging the geometric and semantic knowledge of multiple known classes to achieve s

Cited by 0SourcecodeScholar
2026

Simulating Distribution Dynamics: Liquid Temporal Feature Evolution for Single-Domain Generalized Object Detection

AAAI 2026technical

In this paper, we focus on Single-Domain Generalized Object Detection (Single-DGOD), aiming to transfer a detector trained on one source domain to multiple unknown domains. Existing methods for Single-DGOD typically rely on discrete data augmentation or static perturbation methods to expand data div

Cited by 0SourcePDFScholar
2026

Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning

CVPR 2026

Vision-Language Navigation (VLN) aims to enable agents to navigate to a target location based on language instructions. Traditional VLN often follows a close-set assumption, i.e., training and test data share the same style of the input images and instructions. However, the real world is open and fi

Cited by 0SourceScholar
2026

VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving

CVPR 2026

The significance of cross-view 3D geometric modeling capabilities for autonomous driving is self-evident, yet existing Vision-Language Models (VLMs) inherently lack this capability, resulting in their mediocre performance. While some promising approaches attempt to mitigate this by constructing Q&A

Cited by 0SourcecodeScholar
2025

Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios

ICCV 2025poster

In practice, environments constantly change over time and space, posing significant challenges for object detectors trained based on a closed-set assumption, i.e., training and test data share the same distribution. To this end, continual test-time adaptation has attracted much attention, aiming to…

Cited by 0SourcePDFScholar
2025

Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and Reasoning

NeurIPS 2025poster

In this paper, we focus on Novel Class Discovery for Point Cloud Segmentation (3D-NCD), aiming to learn a model that can segment unlabeled (novel) 3D classes using only the supervision from labeled (base) 3D classes. The key to this task is to setup the exact correlations between the point represent…

Cited by 0SourceScholar
2025

Style Evolving along Chain-of-Thought for Unknown-Domain Object Detection

CVPR 2025highlight

Recently, a task of Single-Domain Generalized Object Detection (Single-DGOD) is proposed, aiming to generalize a detector to multiple unknown domains never seen before during training. Due to the unavailability of target-domain data, some methods leverage the multimodal capabilities of vision-langu…

2025

Unknown Text Learning for CLIP-based Few-Shot Open-set Recognition

ICCV 2025poster

Recently, vision-language models (e.g., CLIP) with prompt learning have shown great potential in few-shot learning. However, an open issue remains for the effective extension of CLIP-based models to few-shot open-set recognition (FSOR), which requires classifying known classes and detecting unknown…

2024

Multi-Source Collaborative Gradient Discrepancy Minimization for Federated Domain Generalization

AAAI 2024technical

Federated Domain Generalization aims to learn a domain-invariant model from multiple decentralized source domains for deployment on unseen target domain. Due to privacy concerns, the data from different source domains are kept isolated, which poses challenges in bridging the domain gap. To address t…

Cited by 10SourcePDFScholar
2024

Prompt-Driven Dynamic Object-Centric Learning for Single Domain Generalization

CVPR 2024poster

Single-domain generalization aims to learn a model from single source domain data attaining generalized performance on other unseen target domains. Existing works primarily focus on improving the generalization ability of static networks. However static networks are unable to dynamically adapt to th…

Cited by 12SourcePDFScholar
2023

Reliable and Interpretable Personalized Federated Learning

CVPR 2023poster

Federated learning can coordinate multiple users to participate in data training while ensuring data privacy. The collaboration of multiple agents allows for a natural connection between federated learning and collective intelligence. When there are large differences in data distribution among clien…

Cited by 27SourcePDFScholar
2022

Decision-based Black-box Attack Against Vision Transformers via Patch-wise Adversarial Removal

NeurIPS 2022accept

Vision transformers (ViTs) have demonstrated impressive performance and stronger adversarial robustness compared to Convolutional Neural Networks (CNNs). On the one hand, ViTs' focus on global interaction between individual patches reduces the local noise sensitivity of images. On the other hand, th…

2021

Vector-Decomposed Disentanglement for Domain-Invariant Object Detection

ICCV 2021poster

To improve the generalization of detectors, for domain adaptive object detection (DAOD), recent advances mainly explore aligning feature-level distributions between the source and single-target domain, which may neglect the impact of domain-specific information existing in the aligned features. Towa…

Cited by 136PDFcodeScholar
2020

Bidirectional Adversarial Training for Semi-Supervised Domain Adaptation

IJCAI 2020poster

Semi-supervised domain adaptation (SSDA) is a novel branch of machine learning that scarce labeled target examples are available, compared with unsupervised domain adaptation. To make effective use of these additional data so as to bridge the domain gap, one possible way is to generate adversarial e…

Cited by 0SourcePDFScholar
2020

Extract and Merge: Superpixel Segmentation with Regional Attributes

ECCV 2020poster

For a certain object in an image, the relationship between its central region and the peripheral region is not well utilized in existing superpixel segmentation methods. In this work, we propose the concept of regional attribute, which indicates the location of a certain region in the object. Based…

Cited by 3SourcePDFScholar
2019

Connective Cognition Network for Directional Visual Commonsense Reasoning

NeurIPS 2019poster

Visual commonsense reasoning (VCR) has been introduced to boost research of cognition-level visual understanding, i.e., a thorough understanding of correlated details of the scene plus an inference with related commonsense knowledge. Recent studies on neuroscience have suggested that brain function…

2018

Image-Based PM2.5 Estimation and its Application on Depth Estimation

ICASSP 2018accepted

Air pollution is still a big threat to human health particularly for developing countries. It is highly demanding to measure air quality with daily-used devices such as smartphones. On the other hand, it is difficult to estimate the scene depth under the foul weather using traditional vision-based m…

Cited by 0SourceScholar