← Search

Fei Yang

27 accepted papers

2026

LONGSPEECH: A SCALABLE BENCHMARK FOR TRANSCRIPTION, TRANSLATION AND UNDERSTANDING IN LONG SPEECH

ICASSP 2026poster

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis require robust models capable of processing and reasoning ove…

Cited by 0SourcePDFScholar
2026

Linguistic Steganography via Self-Adjusting Asymmetric Number System (Abstract Reprint)

AAAI 2026technical

Linguistic steganography (stego) seeks to conceal secret information within natural language text. However, existing methods often struggle to balance stego text quality with embedding efficiency, largely due to limitations in generation strategies and coding mechanisms. We propose SA-ANS, a self-ad

Cited by 0SourcePDFScholar
2026

Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search

ICLR 2026poster

As Large Language Models (LLMs) are increasingly used, their security risks have drawn increasing attention. Existing research reveals that LLMs are highly susceptible to jailbreak attacks, with effectiveness varying across language contexts. This paper investigates the role of classical Chinese in…

Cited by 0SourcecodeScholar
2025

BSLoRA: Enhancing the Parameter Efficiency of LoRA with Intra-Layer and Inter-Layer Sharing

ICML 2025poster

Low-Rank Adaptation (LoRA) is a widely adopted parameter-efficient fine-tuning method for large language models (LLMs) to adapt to downstream tasks. However, in scenarios where multiple LoRA models are deployed simultaneously, standard LoRA introduces substantial trainable parameters, resulting in s…

2025

GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery

CVPR 2025poster

Given unlabelled datasets containing both old and new categories, generalized category discovery (GCD) aims to accurately discover new classes while correctly classifying old classes. Current GCD methods only use a single visual modality of information, resulting in poor classification of visually s…

2025

Improving Continual Learning Performance and Efficiency with Auxiliary Classifiers

ICML 2025poster

Continual learning is crucial for applying machine learning in challenging, dynamic, and often resource-constrained environments. However, catastrophic forgetting — overwriting previously learned knowledge when new information is acquired — remains a major challenge. In this work, we examine the int…

Cited by 0SourcePDFScholar
2025

Improving Video Generation with Human Feedback

NeurIPS 2025poster

Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video gene…

Cited by 0SourceScholar
2025

KAC: Kolmogorov-Arnold Classifier for Continual Learning

CVPR 2025highlight

Continual learning requires models to train continuously across consecutive tasks without forgetting. Most existing methods utilize linear classifiers, which struggle to maintain a stable classification space while learning new tasks. Inspired by the success of Kolmogorov-Arnold Networks (KAN) in pr…

2025

Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning

NeurIPS 2025poster

Continual learning in computer vision faces the critical challenge of catastrophic forgetting, where models struggle to retain prior knowledge while adapting to new tasks. Although recent studies have attempted to leverage the generalization capabilities of pre-trained models to mitigate overfitting…

Cited by 0SourceScholar
2025

Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

CVPR 2025poster

With the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the performance of video generation models. We assert that temporal splitting, detailed captions, and video quality filtering are…

2025

Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning

ICCV 2025poster

Continual learning aims to enable models to learn sequentially from continuously incoming data while retaining performance on previously learned tasks. With the Contrastive Language-Image Pre-trained model (CLIP) exhibiting strong capabilities across various downstream tasks, there has been growing…

2025

Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability

CVPR 2025poster

The diffusion models, in early stages focus on constructing basic image structures, while the refined details, including local features and textures, are generated in later stages. Thus the same network layers are forced to learn both structural and textural information simultaneously, significant…

2024

FlattenQuant: Breaking through the Inference Compute-bound for Large Language Models with Per-tensor Quantization

COLING 2024main

Large language models (LLMs) have demonstrated state-of-the-art accuracies across various tasks. However, the latency of inference and the large GPU memory consumption of LLMs restrict their deployment performance. Recently, there have been some efficient attempts to quantize LLMs, yet inference wit…

Cited by 3SourcePDFScholar
2023

Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image Editing

NeurIPS 2023poster

Large-scale text-to-image generative models have been a ground-breaking development in generative AI, with diffusion models showing their astounding ability to synthesize convincing images following an input text prompt. The goal of image editing research is to give users control over the generated…

2023

Efficient Super-Resolution for Compression Of Gaming Videos

ICASSP 2023accepted

Due to the increasing demand for game-streaming services, efficient compression of computer-generated video is more critical than ever, especially when the available bandwidth is low. This paper proposes a super-resolution framework that improves the coding efficiency of computer-generated gaming vi…

Cited by 0SourceScholar
2023

Semantic Preprocessor for Image Compression for Machines

ICASSP 2023accepted

Visual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for hu…

Cited by 0SourceScholar
2022

A Ball Joint With Continuously Adjustable Load Capacity Based on Positive Pressure Method

RA-L 2022

Research on the variable load capacity method is crucial in promoting the application of soft robots. However, the effects on the performance of current structures are still limited. This letter proposes a ball joint with continuously adjustable load capacity based on the principle of friction with

Cited by 3SourceScholar
2022

Exploring evolution-aware & -free protein language models as protein function predictors

NeurIPS 2022accept

Large-scale Protein Language Models (PLMs) have improved performance in protein prediction tasks, ranging from 3D structure prediction to various function predictions. In particular, AlphaFold, a ground-breaking AI system, could potentially reshape structural biology. However, the utility of the PLM…

2022

Learning Unbiased Transferability for Domain Adaptation by Uncertainty Modeling

ECCV 2022poster

"Domain adaptation (DA) aims to transfer knowledge learned from a labeled source domain to an unlabeled or a less labeled but related target domain. Ideally, the source and target distributions should be aligned to each other equally to achieve unbiased knowledge transfer. However, due to the signif…

2022

Towards Unified Prompt Tuning for Few-shot Text Classification

EMNLP 2022finding

Prompt-based fine-tuning has boosted the performance of Pre-trained Language Models (PLMs) on few-shot text classification by employing task-specific prompts. Yet, PLMs are unfamiliar with prompt-style expressions during pre-training, which limits the few-shot learning performance on downstream task…

2021

Meta Distant Transfer Learning for Pre-trained Language Models

EMNLP 2021main

With the wide availability of Pre-trained Language Models (PLMs), multi-task fine-tuning across domains has been extensively applied. For tasks related to distant domains with different class label sets, PLMs may memorize non-transferable knowledge for the target domain and suffer from negative tran…

2021

Slimmable Compressive Autoencoders for Practical Neural Image Compression

CVPR 2021poster

Neural image compression leverages deep neural networks to outperform traditional image codecs in rate-distortion performance. However, the resulting models are also heavy, computationally demanding and generally optimized for a single rate, limiting their practical use. Focusing on practical image…

Cited by 94PDFcodeScholar
2019

Efficient Segmentation: Learning Downsampling Near Semantic Boundaries

ICCV 2019poster

Many automated processes such as auto-piloting rely on a good semantic segmentation as a critical component. To speed up performance, it is common to downsample the input frame. However, this comes at the cost of missed small objects and reduced accuracy at semantic boundaries. To address this probl…

Cited by 108PDFScholar
2018

Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation

CVPR 2018poster

Random data augmentation is a critical technique to avoid overfitting in training deep models. Yet, data augmentation and network training are often two isolated processes in most settings, yielding to a suboptimal training. Why not jointly optimize the two? We propose adversarial data augmentation…

Cited by 284SourcePDFScholar
2015

Web Scale Photo Hash Clustering on A Single Machine

CVPR 2015poster

This paper addresses the problem of clustering a very large number of photos (i.e. hundreds of millions a day) in a stream into millions of clusters. This is particularly important as the popularity of photo sharing websites, such as Facebook, Google, and Instagram. Given large number of photos avai…

Cited by 120SourcePDFScholar