← Search

Zhenbo Li

11 accepted papers

2026

RFM-EDITING: RECTIFIED FLOW MATCHING FOR TEXT-GUIDED AUDIO EDITING

ICASSP 2026poster

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest, thus demanding precise localization and faithful editing ac…

Cited by 0SourcePDFScholar
2026

TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery

CVPR 2026

On-the-fly category discovery (OCD) aims to recognize known categories while simultaneously discovering novel ones from an unlabeled online stream, using a model trained only on labeled data. Existing approaches freeze the feature extractor trained offline and employ a hash-based framework that quan

Cited by 0SourcecodeScholar
2026

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

ICASSP 2026poster

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative learning, but neglected stable segment-level supervision and class…

Cited by 0SourcePDFScholar
2026

U$^3$CF: Unbiased, Unconfounding, and Unified Causal Framework for Multi-Target Domain Adaptation

ICML 2026poster

Multi-target domain adaptation (MTDA) trains a model using a labeled source domain and several unlabeled target domains, aiming to enhance performance across all targets. However, existing methods lack a principled causal formulation and often rely on empirical domain-invariance enforcement, which c…

Cited by 0SourceScholar
2026

When Trackers Date Fish: A Benchmark and Framework for Underwater Multiple Fish Tracking

AAAI 2026technical

Multiple object tracking (MOT) technology has made significant progress in terrestrial applications, but underwater tracking scenarios remain underexplored despite their importance to marine ecology and aquaculture. In this paper, we present Multiple Fish Tracking Dataset 2025 (MFT25), a comprehensi

Cited by 0SourcePDFScholar
2025

PDTrack: Progressive Distance Association for Multiple Object Tracking

ICASSP 2025accepted

Multiple Object Tracking (MOT) has made significant progress in recent years. However, it still faces challenges such as frequent ID switches, trajectory fragmentation, and tracking losses in high-density pedestrian scenarios. To address these issues, we optimized the association algorithm based on…

Cited by 0SourceScholar
2025

PlantPCC: Dual Sampling and Multi-level Geometry-aware Contrastive Regularization for Plant Point Cloud Completion

ICASSP 2025accepted

Plant point cloud completion is essential for tasks like segmentation and surface reconstruction in plant phenotyping. Unlike the relatively simpler Computer-Aided Design models found in datasets like ShapeNet, plant point clouds are characterized by their rich geometric shapes and intricate edge fe…

Cited by 0SourceScholar
2024

CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video Parsing

ICASSP 2024accepted

Audio-visual video parsing is the task of categorizing a video with weak labels at the segment level, and predicting them as audible or visible events. Recent methods have leveraged the attention mechanism to capture the semantic correlations among the whole video across the audio-visual modalities.…

Cited by 0SourceScholar
2023

Multi-Frequency Representation Enhancement with Privilege Information for Video Super-Resolution

ICCV 2023poster

CNN's limited receptive field restricts its ability to capture long-range spatial-temporal dependencies, leading to unsatisfactory performance in video super-resolution. To tackle this challenge, this paper presents a novel multi-frequency representation enhancement module (MFE) that performs spatia…

Cited by 20PDFScholar