← Search

hui xue

70 accepted papers

2026

Adaptive Hyperbolic Kernels: Modulated Embedding in de Branges-Rovnyak Spaces

AAAI 2026technical

Hierarchical data pervades diverse machine learning applications, including natural language processing, computer vision, and social network analysis. Hyperbolic space, characterized by its negative curvature, has demonstrated strong potential in such tasks due to its capacity to embed hierarchical

Cited by 0SourcePDFScholar
2026

Diffusion Probe: Generated Image Result Prediction Using CNN Probes

CVPR 2026

Text-to-image (T2I) diffusion models currently lack an efficient mechanism for early quality assessment, forcing costly random trial-and-error in scenarios requiring multiple generations (e.g., iterating on prompts, agent-based image generation, flow-grpo). To address this, we first reveal a strong

Cited by 0SourceScholar
2026

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration

CVPR 2026

Ancient inscriptions, as repositories of cultural memory, have suffered from centuries of environmental and human-induced degradation. Restoring their intertwined visual and textual integrity poses one of the most demanding challenges in digital heritage preservation. However, existing AI-based appr

Cited by 0SourcecodeScholar
2026

Intra-Image Mining and Symmetric Maximum Concept Matching for Few Shot Out-of-Distribution Detection

AAAI 2026technical

Recent vision-language model (VLM)-based methods have achieved promising results in zero-shot out-of-distribution (OOD) detection by effectively leveraging the local patch features. However, the zero-shot nature inherently comes with two limitations: 1) imperfect local feature prototypes; 2) lack of

Cited by 0SourcePDFScholar
2026

MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical Documents

CVPR 2026

Reading Chinese historical documents across diverse carriers is central to understanding the evolution of Chinese civilization, yet remains labor-intensive and dependent on scarce expert knowledge. Although recent large-scale models show promise on isolated historical collections, they do not system

Cited by 0SourceScholar
2026

SIMPLEPOSTER: A SIMPLE BASELINE FOR PRODUCT POSTER GENERATION

CVPR 2026

Product poster generation presents unique challenges beyond general-purpose de-sign: it demands not only aesthetic composition and accurate text rendering, butalso strict preservation of the product subject and precise control over dense,multi-line text layouts. While general image editing models st

Cited by 0SourcecodeScholar
2026

TC-Pade: Trajectory-Consistent Pade Approximation for Diffusion Acceleration

CVPR 2026

Despite achieving state-of-the-art generation quality, diffusion models are hindered by the substantial computational burden of their iterative sampling process. While feature caching techniques achieve effective acceleration at higher step counts (e.g., 50 steps), they exhibit critical limitations

Cited by 0SourceScholar
2026

Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrieval

CVPR 2026

This paper studies unsupervised cross-domain image retrieval (UCDIR), which aims to retrieve images of the same category across different domains without relying on labeled data. Existing methods typically utilize pseudo-labels, derived from clustering algorithms, as supervisory signals for intra-do

Cited by 0SourceScholar
2026

Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge Intelligence

ICLR 2026poster

Large Vision Transformers (ViTs) must often be compressed before they can be deployed on resource-constrained edge devices. However, many edge devices require only part of the *all-classes* knowledge of a pre-trained ViT in their corresponding application scenarios. This is overlooked by existing c…

Cited by 0SourcecodeScholar
2025

A New Model for Prototype-based Continual Learning in Hyperspherical Space

ICASSP 2025accepted

The continuous emergence of new objects in the visual world poses a serious challenge to deep object recognition methods, which sparks the increasing study on continual or incremental learning. However, learning new tasks faces the tough catastrophic forgetting problem, i.e., dramatic performance de…

Cited by 0SourceScholar
2025

Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning

ICML 2025poster

Continual multimodal instruction tuning is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving tasks. However, most existing methods adopt a fixed architecture, struggling with adapting to new tasks due to static model capacity. We propose to evolve the architecture under param…

Cited by 0SourcePDFScholar
2025

Exploring the Relationship Between Samples and Masks for Robust Defect Localization

AAAI 2025technical

Defect detection aims to detect and localize regions out of the normal distribution. The previous approaches often explicitly incorporate the defect detection concept, such as by utilizing self-supervised ground truth or manually defined feature comparison. The aforementioned processes involve model…

Cited by 0SourcePDFScholar
2025

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have achieved impressive performance and have been put into practical use in commercial applications, but they still have potential safety mechanism vulnerabilities. Jailbreak attacks are red teaming methods that aim to bypass safety mechanisms and discover M…

2025

KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration

ICASSP 2025accepted

Zero-shot anomaly detection (ZSAD) identifies anomalies without needing training samples from the target dataset, essential for scenarios with privacy concerns or limited data. Vision-language models like CLIP show potential in ZSAD but have limitations: relying on manually crafted fixed textual des…

Cited by 0SourceScholar
2025

Modularized Self-Reflected Video Reasoner for Multimodal LLM with Application to Video Question Answering

ICML 2025poster

Multimodal Large Language Models (Multimodal LLMs) have shown their strength in Video Question Answering (VideoQA). However, due to the black-box nature of end-to-end training strategies, existing approaches based on Multimodal LLMs suffer from the lack of interpretability for VideoQA: they can neit…

Cited by 0SourcePDFScholar
2025

Multi-Modal Interactive Agent Layer for Few-Shot Universal Cross-Domain Retrieval and Beyond

NeurIPS 2025poster

This paper firstly addresses the challenge of few-shot universal cross-domain retrieval (FS-UCDR), enabling machines trained with limited data to generalize to novel retrieval scenarios, with queries from entirely unknown domains and categories. To achieve this, we first formally define the FS-UCDR…

Cited by 0SourceScholar
2025

PEARL: Input-Agnostic Prompt Enhancement with Negative Feedback Regulation for Class-Incremental Learning

AAAI 2025technical

Class-incremental learning (CIL) aims to continuously introduce novel categories into a classification system without forgetting previously learned ones, thus adapting to evolving data distributions. Researchers are currently focusing on leveraging the rich semantic information of pre-trained models…

2025

Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models

NeurIPS 2025poster

Recent advancements, such as DeepSeek-Prover-V2-671B and Kimina-Prover-Preview-72B, demonstrate a prevailing trend in leveraging reinforcement learning (RL)-based large-scale training for automated theorem proving. Surprisingly, we discover that even without any training, careful neuro-symbolic coor…

Cited by 0SourceScholar
2025

SVasP: Self-Versatility Adversarial Style Perturbation for Cross-Domain Few-Shot Learning

AAAI 2025technical

Cross-Domain Few-Shot Learning (CD-FSL) aims to transfer knowledge from seen source domains to unseen target domains, which is crucial for evaluating the generalization and robustness of models. Recent studies focus on utilizing visual styles to bridge the domain gap between different domains. Howev…

2025

The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning

ACL 2025finding

Understanding alignment techniques begins with comprehending zero-shot generalization brought by instruction tuning, but little of the mechanism has been understood. Existing work has largely been confined to the task level, without considering that tasks are artificially defined and, to LLMs, merel…

2025

Towards the Resistance of Neural Network Fingerprinting to Fine-tuning

NeurIPS 2025poster

This paper proves a new fingerprinting method to embed the ownership information into a deep neural network (DNN) with theoretically guaranteed robustness to fine-tuning. Specifically, we prove that when the input feature of a convolutional layer only contains low-frequency components, specific freq…

Cited by 0SourceScholar
2024

KEHRL: Learning Knowledge-Enhanced Language Representations with Hierarchical Reinforcement Learning

COLING 2024main

Knowledge-enhanced pre-trained language models (KEPLMs) leverage relation triples from knowledge graphs (KGs) and integrate these external data sources into language models via self-supervised learning. Previous works treat knowledge enhancement as two independent operations, i.e., knowledge injecti…

2024

One-dimensional Adapter to Rule Them All: Concepts Diffusion Models and Erasing Applications

CVPR 2024highlight

The prevalent use of commercial and open-source diffusion models (DMs) for text-to-image generation prompts risk mitigation to prevent undesired behaviors. Existing concept erasing methods in academia are all based on full parameter or specification-based fine-tuning from which we observe the follow…

2024

TRELM: Towards Robust and Efficient Pre-training for Knowledge-Enhanced Language Models

COLING 2024main

KEPLMs are pre-trained models that utilize external knowledge to enhance language understanding. Previous language models facilitated knowledge acquisition by incorporating knowledge-related pre-training tasks learned from relation triples in knowledge graphs. However, these models do not prioritize…

2024

Text Image Inpainting via Global Structure-Guided Diffusion Models

AAAI 2024technical

Real-world text can be damaged by corrosion issues caused by environmental or human factors, which hinder the preservation of the complete styles of texts, e.g., texture and structure. These corrosion issues, such as graffiti signs and incomplete signatures, bring difficulties in understanding the t…

2024

UniPSDA: Unsupervised Pseudo Semantic Data Augmentation for Zero-Shot Cross-Lingual Natural Language Understanding

COLING 2024main

Cross-lingual representation learning transfers knowledge from resource-rich data to resource-scarce ones to improve the semantic understanding abilities of different languages. However, previous works rely on shallow unsupervised data generated by token surface matching, regardless of the global co…

2023

An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies

ICASSP 2023accepted

In this paper, we propose an effective anomalous sound detection (ASD) method based on representation learning with simulated anomalies. Recently, ASD systems have used Outlier Exposure (OE) strategy to achieve promising performance in DCASE challenges. These exploit deep Convolutional Neural Networ…

Cited by 0SourceScholar
2023

COCO-O: A Benchmark for Object Detectors under Natural Distribution Shifts

ICCV 2023poster

Practical object detection application can lose its effectiveness on image inputs with natural distribution shifts. This problem leads the research community to pay more attention on the robustness of detectors under Out-Of-Distribution (OOD) inputs. Existing works construct datasets to benchmark th…

Cited by 23PDFcodeScholar
2023

Expanding the Hyperbolic Kernels: A Curvature-aware Isometric Embedding View

IJCAI 2023poster

Modeling data relation as a hierarchical structure has proven beneficial for many learning scenarios, and the hyperbolic space, with negative curvature, can encode such data hierarchy without distortion. Several recent studies also show that the representation power of the hyperbolic space can be fu…

2023

From Adversarial Arms Race to Model-centric Evaluation: Motivating a Unified Automatic Robustness Evaluation Framework

ACL 2023findings

Textual adversarial attacks can discover models’ weaknesses by adding semantic-preserved but misleading perturbations to the inputs. The long-lasting adversarial attack-and-defense arms race in Natural Language Processing (NLP) is algorithm-centric, providing valuable techniques for automatic robust…

2023

ImageNet-E: Benchmarking Neural Network Robustness via Attribute Editing

CVPR 2023poster

Recent studies have shown that higher accuracy on ImageNet usually leads to better robustness against different corruptions. In this paper, instead of following the traditional research paradigm that investigates new out-of-distribution corruptions or perturbations deep models may encounter, we cond…

2023

Improving Scene Text Image Super-resolution via Dual Prior Modulation Network

AAAI 2023technical

Scene text image super-resolution (STISR) aims to simultaneously increase the resolution and legibility of the text images, and the resulting images will significantly affect the performance of downstream tasks. Although numerous progress has been made, existing approaches raise two crucial issues:…

2023

Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection

ICASSP 2023accepted

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together wi…

Cited by 28SourceScholar
2023

Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning

ACL 2023findings

Large language models (LLMs) have recently shown great potential for in-context learning, where LLMs learn a new task simply by conditioning on a few input-label pairs (prompts). Despite their potential, our understanding of the factors influencing end-task performance and the robustness of in-conte…

2023

Open-Vocabulary Object Detection With an Open Corpus

ICCV 2023poster

Existing open vocabulary object detection (OVD) works expand the object detector toward open categories by replacing the classifier with the category text embeddings and optimizing the region-text alignment on data of the base categories. However, both the class-agnostic proposal generator and the c…

Cited by 22PDFScholar
2023

Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to…

Cited by 0SourceScholar
2023

Transaudio: Towards the Transferable Adversarial Audio Attack Via Learning Contextualized Perturbations

ICASSP 2023accepted

In a transfer-based attack against Automatic Speech Recognition (ASR) systems, attacks are unable to access the architecture and parameters of the target model. Existing attack methods are mostly investigated in voice assistant scenarios with restricted voice commands, prohibiting their applicabilit…

Cited by 0SourceScholar
2022

Automatically Gating Multi-Frequency Patterns through Rectified Continuous Bernoulli Units with Theoretical Principles

IJCAI 2022poster

Different nonlinearities are only suitable for responding to different frequency signals. The locally-responding ReLU is incapable of modeling high-frequency features due to the spectral bias, whereas the globally-responding sinusoidal function is intractable to represent low-frequency concepts chea…

Cited by 0SourcePDFScholar
2022

LightPose: A Lightweight and Efficient Model with Transformer for Human Pose Estimation

ICASSP 2022accepted

The prediction of keypoints by generating high-resolution heatmaps has become a popular solution in human pose estimation. While this kind of method requires up-sampling or deconvolution operations, which would bring a great challenge to the acceleration of model inference. If performing keypoint pr…

Cited by 0SourceScholar
2022

Multiple Instance Learning for Offensive Language Detection

EMNLP 2022finding

Automatic offensive language detection has become a crucial issue in recent years. Existing researches on this topic are usually based on a large amount of data annotated at sentence level to train a robust model. However, sentence-level annotations are expensive in practice as the scenario expands,…

Cited by 5SourcePDFScholar
2022

RoChBert: Towards Robust BERT Fine-tuning for Chinese

EMNLP 2022finding

Despite of the superb performance on a wide range of tasks, pre-trained language models (e.g., BERT) have been proved vulnerable to adversarial texts. In this paper, we present RoChBERT, a framework to build more Robust BERT-based models by utilizing a more comprehensive adversarial graph to fuse Ch…

2022

Supervised Prototypical Contrastive Learning for Emotion Recognition in Conversation

EMNLP 2022main

Capturing emotions within a conversation plays an essential role in modern dialogue systems. However, the weak correlation between emotions and semantics brings many challenges to emotion recognition in conversation (ERC). Even semantically similar utterances, the emotion may vary drastically depend…

2021

Cut out the annotator, keep the cutout: better segmentation with weak supervision

ICLR 2021poster

Constructing large, labeled training datasets for segmentation models is an expensive and labor-intensive process. This is a common challenge in machine learning, addressed by methods that require few or no labeled data points such as few-shot learning (FSL) and weakly-supervised learning (WS). Such…

Cited by 23SourcePDFScholar
2021

Enhancing Model Robustness by Incorporating Adversarial Knowledge into Semantic Representation

ICASSP 2021accepted

Despite that deep neural networks (DNNs) have achieved enormous success in many domains like natural language processing (NLP), they have also been proven to be vulnerable to maliciously generated adversarial examples. Such inherent vulnerability has threatened various real-world deployed DNNs-based…

Cited by 0SourceScholar
2021

QAIR: Practical Query-Efficient Black-Box Attacks for Image Retrieval

CVPR 2021poster

We study the query-based attack against image retrieval to evaluate its robustness against adversarial examples under the black-box setting, where the adversary only has query access to the top-k ranked unlabeled images from the database. Compared with query attacks in image classification, which pr…

Cited by 64PDFcodeScholar
2021

Self-Supervised Learning for Few-Shot Image Classification

ICASSP 2021accepted

Few-shot image classification aims to classify unseen classes with limited labelled samples. Recent works benefit from the meta-learning process with episodic tasks and can fast adapt to class from training to testing. Due to the limited number of samples for each task, the initial embedding network…

Cited by 0SourceScholar
2021

Spatial-Phase Shallow Learning: Rethinking Face Forgery Detection in Frequency Domain

CVPR 2021poster

The remarkable success in face forgery techniques has received considerable attention in computer vision due to security concerns. We observe that up-sampling is a necessary step of most face forgery techniques, and cumulative up-sampling will result in obvious changes in the frequency domain, espec…

Cited by 517PDFScholar
2021

Towards Face Encryption by Generating Adversarial Identity Masks

ICCV 2021poster

As billions of personal data being shared through social media and network, the data privacy and security have drawn an increasing attention. Several attempts have been made to alleviate the leakage of identity information from face photos, with the aid of, e.g., image obfuscation techniques. Howeve…

Cited by 120PDFcodeScholar
2020

Design-Gan: Cross-Category Fashion Translation Driven By Landmark Attention

ICASSP 2020accepted

The rise of generative adversarial networks has boosted a vast interest in the field of fashion image-to-image translation. However, previous methods do not perform well in cross-category translation tasks, e.g., translating jeans to skirts in fashion images. The translated skirts are easier to lose…

Cited by 0SourceScholar
2020

Hierarchical Sequence Representation with Graph Network

ICASSP 2020accepted

Video classification problem is a challenging task in computer vision. The performance of this task is highly relied on the scale of training data and the effectiveness of video embedding via a robust embedding network. Unsupervised solutions such as feature average pooling technique, as a simple la…

Cited by 0SourceScholar
2020

The Open Brands Dataset: Unified Brand Detection and Recognition at Scale

ICASSP 2020accepted

Intellectual property protection(IPP) have received more and more attention recently due to the development of the global e-commerce platforms. brand recognition plays a significant role in IPP. Recent studies for brand recognition and detection are based on small-scale datasets that are not compreh…

Cited by 0SourceScholar
2020

Which Is Plagiarism: Fashion Image Retrieval Based on Regional Representation for Design Protection

CVPR 2020oral

With the rapid growth of e-commerce and the popularity of online shopping, fashion retrieval has received considerable attention in the computer vision community. Different from the existing works that mainly focus on identical or similar fashion item retrieval, in this paper, we aim to study the pl…

Cited by 47PDFScholar
2019

Bilinear Representation for Language-based Image Editing Using Conditional Generative Adversarial Networks

ICASSP 2019accepted

The task of Language-Based Image Editing (LBIE) aims at generating a target image by editing the source image based on the given language description. The main challenge of LBIE is to disentangle the semantics in image and text and then combine them to generate realistic images. Therefore, the editi…

Cited by 0SourceScholar