← Search

Jizhong Zhao

8 accepted papers

2025

Enhancing Multimodal Model Robustness Under Missing Modalities via Memory-Driven Prompt Learning

IJCAI 2025

Existing multimodal models typically assume the availability of all modalities, leading to significant performance degradation when certain modalities are missing. Recent methods have introduced prompt learning to adapt pretrained models to incomplete data, achieving remarkable performance when the

2025

Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech Recognition

ICASSP 2025accepted

Audio-visual speech recognition (AVSR) aims to enhance the robustness of an automatic speech recognition (ASR) systems by incorporating visual information from lip movements, especially in challenging noisy environments. Nevertheless, most current approaches either involve training from scratch or f…

Cited by 0SourceScholar
2025

Open-Modality Latent Modality Interaction Maximization for Audio-Visual Learning

ICASSP 2025accepted

The utilization of multimodal cues enhances the effectiveness of specific cognitive tasks in audio-visual learning. However, on the one hand, designing a unified model for multimodal learning poses challenges due to the presence of information redundancy and modality noise. On the other hand, existi…

Cited by 0SourceScholar
2025

SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation

NeurIPS 2025poster

Dataset condensation aims to synthesize compact yet informative datasets that retain the training efficacy of full-scale data, offering substantial gains in efficiency. Recent studies reveal that the condensation process can be vulnerable to backdoor attacks, where malicious triggers are injected in…

Cited by 0SourcecodeScholar
2024

Attention Shifting to Pursue Optimal Representation for Adapting Multi-granularity Tasks

IJCAI 2024poster

Object recognition in open environments, e.g., video surveillance, poses significant challenges due to the inclusion of unknown and multi-granularity tasks (MGT). However, recent methods exhibit limitations as they struggle to capture subtle differences between different parts within an object and a…

Cited by 1SourcePDFScholar
2024

UFDA: Universal Federated Domain Adaptation with Practical Assumptions

AAAI 2024technical

Conventional Federated Domain Adaptation (FDA) approaches usually demand an abundance of assumptions, which makes them significantly less feasible for real-world situations and introduces security hazards. This paper relaxes the assumptions from previous FDAs and studies a more practical scenario na…

2023

Knowledge-Graph Augmented Music Representation for Genre Classification

ICASSP 2023accepted

In this paper, we propose KGenre, a knowledge-embedded music representation learning framework for improved genre classification. We construct the knowledge graph from the metadata in the open-source FMA-medium and OpenMIC-2018 datasets, with no extra information/effort required. KGenre then mines t…

Cited by 0SourceScholar