← Search

Yunhui Guo

23 accepted papers

2026

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

AAAI 2026technical

Unlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sources. Recent AVS approaches have achieved impressive performance on standard benchmarks. Yet, an important question rema

Cited by 0SourcePDFScholar
2026

Learnability-Driven Submodular Optimization for Active Roadside 3D Detection

CVPR 2026

Roadside perception datasets are typically constructed via cooperative labeling between synchronized vehicle and roadside frame pairs, but real deployment is often limited roadside-only data due to hardware and privacy constraints. The observation that even human experts struggle to produce accurate

Cited by 0SourcecodeScholar
2026

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

CVPR 2026

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical instruments, and human speech). Prior video-conditioned audio genera

Cited by 0SourcecodeScholar
2025

$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time

NeurIPS 2025poster

While recent audio-visual models have demonstrated impressive performance, their robustness to distributional shifts at test-time remains not fully understood. Existing robustness benchmarks mainly focus on single modalities, making them insufficient for thoroughly assessing the robustness of audio-…

Cited by 0SourcecodeScholar
2025

Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation

IROS 2025

Novel Instance Detection and Segmentation (NIDS) aims at detecting and segmenting novel object instances given a few examples of each instance. We propose a unified, simple, yet effective framework (NIDS-Net) comprising object proposal generation, embedding creation for both instance templates and p

Cited by 13SourcecodeScholar
2025

BATCLIP: Bimodal Online Test-Time Adaptation for CLIP

ICCV 2025poster

Although open-vocabulary classification models like Contrastive Language Image Pretraining (CLIP) have demonstrated strong zero-shot learning capabilities, their robustness to common image corruptions remains poorly understood. Through extensive experiments, we show that zero-shot CLIP lacks robustn…

2025

H2ST: Hierarchical Two-Sample Tests for Continual Out-of-Distribution Detection

CVPR 2025poster

Task Incremental Learning (TIL) is a specialized form of Continual Learning (CL) in which a model incrementally learns from non-stationary data streams. Existing TIL methodologies operate under the closed-world assumption, presuming that incoming data remains in-distribution (ID). However, in an ope…

2025

PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time Adaptation

AAAI 2025technical

Real-world vision models in dynamic environments face rapid shifts in domain distributions, leading to decreased recognition performance. Using unlabeled test data, continuous test-time adaptation (CTTA) directly adjusts a pre-trained source discriminative model to these changing domains. A highly e…

2024

Continual Audio-Visual Sound Separation

NeurIPS 2024poster

In this paper, we introduce a novel continual audio-visual sound separation task, aiming to continuously separate sound sources for new classes while preserving performance on previously learned classes, with the aid of visual guidance. This problem is crucial for practical visually guided auditory…

2024

STONE: A Submodular Optimization Framework for Active 3D Object Detection

NeurIPS 2024poster

3D object detection is fundamentally important for various emerging applications, including autonomous driving and robotics. A key requirement for training an accurate 3D object detector is the availability of a large amount of LiDAR-based point cloud data. Unfortunately, labeling point cloud data i…

2024

Unsupervised Feature Learning with Emergent Data-Driven Prototypicality

CVPR 2024poster

Given a set of images our goal is to map each image to a point in a feature space such that not only point proximity indicates visual similarity but where it is located directly encodes how prototypical the image is according to the dataset. Our key insight is to perform unsupervised feature learnin…

Cited by 4SourcePDFScholar
2023

Self-Supervised Unseen Object Instance Segmentation via Long-Term Robot Interaction

RSS 2023poster

We introduce a novel robotic system for improving unseen object instance segmentation in the real world by leveraging long-term robot interaction with objects. Previous approaches either grasp or push an object and then obtain the segmentation mask of the grasped or pushed object after one action. I…

Cited by 9SourcePDFScholar
2022

Unsupervised Hierarchical Semantic Segmentation With Multiview Cosegmentation and Clustering Transformers

CVPR 2022oral

Unsupervised semantic segmentation aims to discover groupings within and across images that capture object- and view-invariance of a category without external supervision. Grouping naturally has levels of granularity, creating ambiguity in unsupervised segmentation. Existing methods avoid this ambig…

Cited by 59PDFcodeScholar
2020

A Broader Study of Cross-Domain Few-Shot Learning

ECCV 2020poster

Recent progress on few-shot learning largely relies on annotated data for meta-learning: base classes sampled from the same domain as the novel classes. However, in many applications, collecting data for meta-learning is infeasible or impossible. This leads to the cross-domain few-shot learning prob…

2020

Improved Schemes for Episodic Memory-based Lifelong Learning

NeurIPS 2020spotlight

Current deep neural networks can achieve remarkable performance on a single task. However, when the deep neural network is continually trained on a sequence of tasks, it seems to gradually forget the previous learned knowledge. This phenomenon is referred to as catastrophic forgetting and motivates…

2019

SpotTune: Transfer Learning Through Adaptive Fine-Tuning

CVPR 2019poster

Transfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning with deep neural networks is to fine-tune a model pretrained on the source task using data from the target task. In this p…

Cited by 640PDFScholar