← Search

Junbo Zhang

20 accepted papers

2026

Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentation

CVPR 2026

Audio-Visual Segmentation (AVS) aims to accurately segment sounding objects in video frames by leveraging audio-visual correspondence cues. However, it remains challenging due to the intrinsic semantic incompleteness within a single modality and the semantic gap between audio and visual representati

Cited by 0SourceScholar
2026

Towards Vision-Spatiotemporal Fusion in Traffic Forecasting: A Survey on Cross-Modal Alignment

IJCAI 2026

Traffic forecasting is evolving, with world models emerging as a powerful framework applicable to tasks such as core state, trajectory, event, and demand forecasting. These tasks involve both visual and spatiotemporal data, yet most existing methods treat them separately, hindering a unified underst

Cited by 0Scholar
2026

Unified Vision-Language-Action Model

ICLR 2026poster

Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on the general comprehension capabilities of vision-language models (VLMs) to generate action signals, often overlooking the…

Cited by 0SourcecodeScholar
2025

AirRadar: Inferring Nationwide Air Quality in China with Deep Neural Networks

AAAI 2025technical

Monitoring real-time air quality is essential for safeguarding public health and fostering social progress. However, the widespread deployment of air quality monitoring stations is constrained by their significant costs. To address this limitation, we introduce AirRadar, a deep neural network design…

Cited by 1SourcePDFScholar
2025

Non-collective Calibrating Strategy for Time Series Forecasting

IJCAI 2025

Deep learning-based approaches have demonstrated significant advancements in time series forecasting. Despite these ongoing developments, the complex dynamics of time series make it challenging to establish the rule of thumb for designing the golden model architecture. In this study, we argue that r

2025

Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping

CVPR 2025poster

Class Incremental Learning (CIL) aims to enable models to learn new classes sequentially while retaining knowledge of previous ones. Although current methods have alleviated catastrophic forgetting (CF), recent studies highlight that the performance of CIL models is highly sensitive to the order of…

2024

CED: Consistent Ensemble Distillation for Audio Tagging

ICASSP 2024accepted

Augmentation and knowledge distillation (KD) are well-established techniques employed in audio classification tasks, aimed at enhancing performance and reducing model sizes on the widely recognized Audioset (AS) benchmark. Although both techniques are effective individually, their combined use, call…

Cited by 0SourceScholar
2024

MG-VLN: Benchmarking Multi-Goal and Long-Horizon Vision-Language Navigation with Language Enhanced Memory Map

IROS 2024poster

Vision-Language Navigation (VLN) with high-level language instructions is a crucial task in robotics. Existing VLN benchmarks, such as the REVERIE challenge which has single-goal instructions and limited navigation steps, do not fully encapsulate the complexity of real-world navigation that often re…

Cited by 0SourceScholar
2023

AirFormer: Predicting Nationwide Air Quality in China with Transformers

AAAI 2023technical

Air pollution is a crucial issue affecting human health and livelihoods, as well as one of the barriers to economic growth. Forecasting air quality has become an increasingly important endeavor with significant social impacts, especially in emerging countries. In this paper, we present a novel Trans…

2023

AutoSTL: Automated Spatio-Temporal Multi-Task Learning

AAAI 2023technical

Spatio-temporal prediction plays a critical role in smart city construction. Jointly modeling multiple spatio-temporal tasks can further promote an intelligent city life by integrating their inseparable relationship. However, existing studies fail to address this joint learning problem well, which g…

Cited by 28SourcePDFScholar
2023

Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?

ICLR 2023poster

The success of deep learning heavily relies on large-scale data with comprehensive labels, which is more expensive and time-consuming to fetch in 3D compared to 2D images or natural languages. This promotes the potential of utilizing models pretrained with data more than 3D as teachers for cross-mod…

2023

Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker Extraction

ICASSP 2023accepted

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio an…

Cited by 0SourceScholar
2023

Language-Assisted 3D Feature Learning for Semantic Scene Understanding

AAAI 2023technical

Learning descriptive 3D features is crucial for understanding 3D scenes with diverse objects and complex structures. However, it is usually unknown whether important geometric attributes and scene context obtain enough emphasis in an end-to-end trained 3D scene understanding network. To guide 3D fea…

2023

Spatio-Temporal Self-Supervised Learning for Traffic Flow Prediction

AAAI 2023technical

Robust prediction of citywide traffic flows at different time periods plays a crucial role in intelligent transportation systems. While previous work has made great efforts to model spatio-temporal correlations, existing methods still suffer from two key limitations: i) Most models collectively pred…

2023

Unified Keyword Spotting and Audio Tagging on Mobile Devices with Transformers

ICASSP 2023accepted

Keyword spotting (KWS) is a core human-machine-interaction front-end task for most modern intelligent assistants. Recently, a unified (UniKW-AT) framework has been proposed that adds additional capabilities in the form of audio tagging (AT) to a KWS model. However, previous work did not consider the…

Cited by 0SourceScholar
2023

Win-Win: A Privacy-Preserving Federated Framework for Dual-Target Cross-Domain Recommendation

AAAI 2023technical

Cross-domain recommendation (CDR) aims to alleviate the data sparsity by transferring knowledge from an informative source domain to the target domain, which inevitably proposes stern challenges to data privacy and transferability during the transfer process. A small amount of recent CDR works have…

Cited by 38SourcePDFScholar
2022

Pseudo Strong Labels for Large Scale Weakly Supervised Audio Tagging

ICASSP 2022accepted

Large-scale audio tagging datasets inevitably contain imperfect labels, such as clip-wise annotated (temporally weak) tags with no exact on- and offsets, due to a high manual labeling cost. This work proposes pseudo strong labels (PSL), a simple label augmentation framework that enhances the supervi…

Cited by 8SourceScholar