← Search

Junbao Zhuo

11 accepted papers

2026

MEDUSA: Motion Elimination in Diffusion Using Spectral Attack

ICML 2026poster

With the widespread application of Video Diffusion Models (VDMs), video synthesis has achieved remarkable temporal dynamics. Image-to-Video (I2V) generation allows users to provide reference images, which enables attackers to inject adversarial noise into these conditions. Due to the robust spatio-t…

Cited by 0SourceScholar
2026

Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models

CVPR 2026

As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human-like Theory of Mind (ToM). Most existing ToM evaluations, however, are centered on text-based inputs, while scenarios relying solely on visual infor

Cited by 0SourceScholar
2025

AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios

NeurIPS 2025poster

By sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous work focus on vehicle-to-vehicle and vehicle-to-infrastructure collaboration, with limited attention to aerial perspectiv…

Cited by 0SourcecodeScholar
2025

Image-to-video Adaptation with Outlier Modeling and Robust Self-learning

AAAI 2025technical

The image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Exis…

2025

SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance Segmentation

CVPR 2025poster

In the field of zero-shot 3D instance segmentation, existing 2D-to-3D lifting methods typically obtain 2D segmentation across multiple RGB frames using vision foundation models, which are then projected and merged into 3D space. However, since the inference of vision foundation models on a single fr…

2024

Confusing Pair Correction Based on Category Prototype for Domain Adaptation under Noisy Environments

AAAI 2024technical

In this paper, we address unsupervised domain adaptation under noisy environments, which is more challenging and practical than traditional domain adaptation. In this scenario, the model is prone to overfitting noisy labels, resulting in a more pronounced domain shift and a notable decline in the ov…

2024

Learning Invariant Representation with Consistency and Diversity for Semi-Supervised Source Hypothesis Transfer

ICASSP 2024accepted

Semi-supervised Domain adaptation (SSDA) has shown promising results by leveraging unlabeled data and limited labeled samples in the target domain. However, accessibility to source data is hindered by data privacy concerns, giving rise to Semi-supervised Source Hypothesis Transfer (SSHT). Integratin…

Cited by 0SourceScholar
2022

Learning Linguistic Association towards Efficient Text-Video Retrieval

ECCV 2022poster

"Text-video retrieval attracts growing attention recently. A dominant approach is to learn a common space for aligning two modalities. However, video deliver richer content than text in general situations and captions usually miss certain events or details in the video. The information imbalance bet…

2020

Gradually Vanishing Bridge for Adversarial Domain Adaptation

CVPR 2020poster

In unsupervised domain adaptation, rich domain-specific characteristics bring great challenge to learn domain-invariant representations. However, domain discrepancy is considered to be directly minimized in existing solutions, which is difficult to achieve in practice. Some methods alleviate the dif…

Cited by 352PDFcodeScholar
2020

Towards Discriminability and Diversity: Batch Nuclear-Norm Maximization Under Label Insufficient Situations

CVPR 2020oral

The learning of the deep networks largely relies on the data with human-annotated labels. In some label insufficient situations, the performance degrades on the decision boundary with high data density. A common solution is to directly minimize the Shannon Entropy, but the side effect caused by entr…

Cited by 489PDFcodeScholar
2019

Unsupervised Open Domain Recognition by Semantic Discrepancy Minimization

CVPR 2019poster

We address the unsupervised open domain recognition (UODR) problem, where categories in labeled source domain S is only a subset of those in unlabeled target domain T. The task is to correctly classify all samples in T including known and unknown categories. UODR is challenging due to the domain dis…

Cited by 38PDFcodeScholar