← Search

Kun Wei

33 accepted papers

2026

Backtrace Mamba: Reviving Critical Temporal Contexts via Hierarchical Memory Compression for Online Action Detection

AAAI 2026technical

Online Action Detection (OAD) requires real-time prediction of ongoing actions without access to future frames, posing challenges in balancing computational efficiency and long-term dependencies modeling.Existing methods either suffer from slow training and limited temporal receptive fields, or face

Cited by 0SourcePDFScholar
2026

DAIEN-TTS: DISENTANGLED AUDIO INFILLING FOR ENVIRONMENT-AWARE TEXT-TO-SPEECH SYNTHESIS

ICASSP 2026poster

This paper presents DAIEN-TTS, a zero-shot text-to-speech (TTS) framework that enables ENvironment-aware synthesis through Disentangled Audio Infilling. By leveraging separate speaker and environment prompts, DAIEN-TTS allows independent control over the timbre and the background environment of the…

Cited by 0SourcePDFScholar
2026

Trajectory-Stabilized Inference for Diffusion-Based Video Inpainting

ICML 2026poster

Video inpainting aims to restore missing regions while preserving spatial and temporal coherence. Diffusion-based methods achieve strong per-frame reconstruction, but their sampling implicitly generates temporally coupled latent trajectories whose long-horizon stability is not explicitly modeled, le…

Cited by 0SourceScholar
2025

Compress to One Point: Neural Collapse for Pre-Trained Model-Based Class-Incremental Learning

AAAI 2025technical

Class-Incremental Learning (CIL) requires an artificial intelligence system to learn different tasks without class overlaps continually. To achieve CIL, some methods introduce the Pre-Trained Model (PTM) and leverage the generalized feature representation of PTM to learn downstream incremental tasks…

2025

Dual-Space Semantic Synergy Distillation for Continual Learning of Unlabeled Streams

NeurIPS 2025poster

Continual learning from unlabeled data streams while effectively combating catastrophic forgetting poses an intractable challenge. Traditional methods predominantly rely on visual clustering techniques to generate pseudo labels, which are frequently plagued by problems such as noise and suboptimal q…

Cited by 0SourceScholar
2025

Energy vs. Noise: Towards Robust Temporal Action Localization in Open-World

AAAI 2025technical

Temporal Action Localization (TAL) aims to accurately identify the start and end times of actions in untrimmed videos and classify them according to specific labels. However, the complexity and imbalance between target actions and background in video data make this task particularly challenging. Alt…

2025

HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models

ICASSP 2025accepted

Recent advancements in integrating Large Language Models (LLM) with automatic speech recognition (ASR) have performed remarkably in general domains. While supervised fine-tuning (SFT) of all model parameters is often employed to adapt pre-trained LLM-based ASR models to specific domains, it imposes…

Cited by 0SourceScholar
2025

Q-MiniSAM2: A Quantization-based Benchmark for Resource-Efficient Video Segmentation

IJCAI 2025

Segment Anything Model 2 (SAM2) is a new-generation, high-precision model for image and video segmentation, offering extensive application prospects across numerous computer vision fields. However, as a large-scale model, its huge memory demands and expansive computing costs pose challenges for prac

Cited by 0SourcePDFScholar
2024

A Versatile Framework for Continual Test-Time Domain Adaptation: Balancing Discriminability and Generalizability

CVPR 2024poster

Continual test-time domain adaptation (CTTA) aims to adapt the source pre-trained model to a continually changing target domain without additional data acquisition or labeling costs. This issue necessitates an initial performance enhancement within the present domain without labels while concurrentl…

Cited by 3SourcePDFScholar
2024

Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video Localization

ACL 2024long

Weakly supervised natural language video localization (WS-NLVL) aims to retrieve the moment corresponding to a language query in a video with only video-language pairs utilized during training. Despite great success, existing WS-NLVL methods seldomly consider the complex temporal relations enclosing…

2024

Long-Tail Class Incremental Learning via Independent Sub-prototype Construction

CVPR 2024poster

Long-tail class incremental learning (LT-CIL) is designed to perpetually acquire novel knowledge from an imbalanced and perpetually evolving data stream while ensuring the retention of previously acquired knowledge. The existing method only re-balances data distribution and ignores exploring the pot…

Cited by 5SourcePDFScholar
2024

Navigating Continual Test-time Adaptation with Symbiosis Knowledge

IJCAI 2024poster

Continual test-time domain adaptation seeks to adapt the source pre-trained model to a continually changing target domain without incurring additional data acquisition or labeling costs. Unfortunately, existing mainstream methods may result in a detrimental cycle. This is attributed to noisy pseudo-…

Cited by 0SourcePDFScholar
2023

Hierarchical Prompt Learning for Compositional Zero-Shot Recognition

IJCAI 2023poster

Compositional Zero-Shot Learning (CZSL) aims to imitate the powerful generalization ability of human beings to recognize novel compositions of known primitive concepts that correspond to a state and an object, e.g., purple apple. To fully capture the intra- and inter-class correlations between compo…

Cited by 23SourcePDFScholar
2023

Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation

ICASSP 2023accepted

Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from the speech of the source language to the speech of the target language are very rare. To add…

Cited by 0SourceScholar
2022

Conversational Speech Recognition by Learning Conversation-Level Characteristics

ICASSP 2022accepted

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can naturally take advantages from specific characteristics of conversation, such as role preference and topical coherence. This p…

Cited by 0SourceScholar
2022

Learning Universal Adversarial Perturbation by Adversarial Example

AAAI 2022technical

Deep learning models have shown to be susceptible to universal adversarial perturbation (UAP), which has aroused wide concerns in the community. Compared with the conventional adversarial attacks that generate adversarial samples at the instance level, UAP can fool the target model for different ins…

2022

ME-GAN: Learning Panoptic Electrocardio Representations for Multi-view ECG Synthesis Conditioned on Heart Diseases

ICML 2022spotlight

Electrocardiogram (ECG) is a widely used non-invasive diagnostic tool for heart diseases. Many studies have devised ECG analysis models (e.g., classifiers) to assist diagnosis. As an upstream task, researches have built generative models to synthesize ECG data, which are beneficial to providing trai…

Cited by 30SourcePDFScholar
2022

Not Just Selection, but Exploration: Online Class-Incremental Continual Learning via Dual View Consistency

CVPR 2022poster

Online class-incremental continual learning aims to learn new classes continually from a never-ending and single-pass data stream, while not forgetting the learned knowledge of old classes. Existing replay-based methods have shown promising performance by storing a subset of old class data. Unfortun…

Cited by 104PDFcodeScholar
2021

Recent Developments on Espnet Toolkit Boosted By Conformer

ICASSP 2021accepted

In this study, we present recent developments on ESPnet: End-to- End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end- to-end speech processing applications, suc…

Cited by 0SourceScholar
2021

SelfSAGCN: Self-Supervised Semantic Alignment for Graph Convolution Network

CVPR 2021poster

Graph convolution networks (GCNs) are a powerful deep learning approach and have been successfully applied to representation learning on graphs in a variety of real-world applications. Despite their success, two fundamental weaknesses of GCNs limit their ability to represent graph-structured data: p…

Cited by 42PDFcodeScholar
2020

Lifelong Zero-Shot Learning

IJCAI 2020poster

Zero-Shot Learning (ZSL) handles the problem that some testing classes never appear in training set. Existing ZSL methods are designed for learning from a fixed training set, which do not have the ability to capture and accumulate the knowledge of multiple training sets, causing them infeasible to m…

Cited by 0SourcePDFScholar
2019

Adversarial Fine-Grained Composition Learning for Unseen Attribute-Object Recognition

ICCV 2019poster

Recognizing unseen attribute-object pairs never appearing in the training data is a challenging task, since an object often refers to a specific entity while an attribute is an abstract semantic description. Besides, attributes are highly correlated to objects, i.e., an attribute tends to describe d…

Cited by 116PDFScholar