← Search

Xiaolin Hu

77 accepted papers

2026

A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation

ICML 2026poster

Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods continue to suffer from residual interference in complex acoustic scenes. This performance limitation stems largely from …

Cited by 0SourcecodeScholar
2026

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

ICLR 2026poster

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) require a rigorous assessment of their trustworthiness. However, existing evaluation frameworks, primarily designed for text, are not equipped to handle the unique vulnerabilities introduced by audio’s acoustic prop…

Cited by 0SourcecodeScholar
2026

CoPE: Continual Probe-guided Expansion for Large Vision-Language Models

ICML 2026poster

Mixture of Experts architectures have recently advanced the scalability and adaptability of Large Language Models for continual multimodal learning. However, extending these models to accommodate sequential tasks remains challenging. As new tasks arrive, naive model expansion leads to rapid paramete…

Cited by 0SourceScholar
2026

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

ICLR 2026poster

Audio-visual speech separation (AVSS) methods leverage visual cues to extract target speech and have demonstrated strong separation quality in noisy acoustic environments. However, these methods usually involve a large number of parameters and require high computational cost, which is unacceptable i…

Cited by 0SourcecodeScholar
2026

FGNet: Leveraging Feature-Guided Attention to Refine SAM2 for 3D EM Neuron Segmentation

AAAI 2026technical

Accurate segmentation of neural structures in Electron Microscopy (EM) images is paramount for neuroscience. However, this task is challenged by intricate morphologies, low signal-to-noise ratios, and scarce annotations, limiting the accuracy and generalization of existing methods. To address these

Cited by 0SourcePDFScholar
2026

MiVE: Multiscale Vision-language features for reference-guided video Editing

ICML 2026poster

Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content. Existing methods fall into two paradigms, each with inherent limitations: deco…

Cited by 0SourceScholar
2026

Mirror Illusion Art

CVPR 2026

Mirror Illusion Art is a novel reflection-conditioned 3D illusion where one object yields two target appearances (front and mirror). The task is formulated as inverse design from two target 2D images (front and mirror) to a printable 3D object with geometry and texture. Prior topology-driven and sha

Cited by 0SourceScholar
2026

Physical Adversarial Clothing Evades Visible-Thermal Detectors via Non-Overlapping RGB-T Pattern

CVPR 2026

Visible-thermal (RGB-T) object detection is a crucial technology for applications such as autonomous driving, where multimodal fusion enhances performance in challenging conditions like low light. However, the security of RGB-T detectors, particularly in the physical world, has been largely overlook

Cited by 0SourcecodeScholar
2026

Put the Space of LoRA Initialization to the Extreme to Preserve Pre-trained Knowledge

AAAI 2026technical

Low-Rank Adaptation (LoRA) is the leading parameter-efficient fine-tuning method for Large Language Models (LLMs), but it still suffers from catastrophic forgetting. Recent work has shown that specialized LoRA initialization can alleviate catastrophic forgetting. There are currently two approaches t

Cited by 0SourcePDFScholar
2026

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

ICLR 2026poster

Large language and vision-language models have inspired end-to-end vision-language-action (VLA) systems in robotics, yet existing robot datasets remain costly, embodiment-specific, and insufficient, limiting robustness and generalization. Recent approaches address this by adopting a plan-then-execut…

Cited by 0SourcecodeScholar
2026

VaccineRAG: Boosting Multimodal Large Language Models’ Immunity to Harmful RAG Samples

AAAI 2026technical

Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating particular strength in real-time queries and Visual Question Answering tasks. However, the effectiveness of RAG is fre

Cited by 0SourcePDFScholar
2025

ADBM: Adversarial Diffusion Bridge Model for Reliable Adversarial Purification

ICLR 2025poster

Recently Diffusion-based Purification (DiffPure) has been recognized as an effective defense method against adversarial examples. However, we find DiffPure which directly employs the original pre-trained diffusion models for adversarial purification, to be suboptimal. This is due to an inherent trad…

2025

ADePT: Adaptive Decomposed Prompt Tuning for Parameter-Efficient Fine-tuning

ICLR 2025poster

Prompt Tuning (PT) enables the adaptation of Pre-trained Large Language Models (PLMs) to downstream tasks by optimizing a small amount of soft virtual tokens, which are prepended to the input token embeddings. Recently, Decomposed Prompt Tuning (DePT) has demonstrated superior adaptation capabilitie…

2025

Connectome-Based Modelling Reveals Orientation Maps in the Drosophila Optic Lobe

NeurIPS 2025poster

The ability to extract oriented edges from visual input is a core computation across animal vision systems. Orientation maps, long associated with the layered architecture of the mammalian visual cortex, systematically organise neurons by their preferred edge orientation. Despite lacking cortical st…

Cited by 0SourceScholar
2025

Efficient Neuron Segmentation in Electron Microscopy by Affinity-Guided Queries

ICLR 2025poster

Accurate segmentation of neurons in electron microscopy (EM) images plays a crucial role in understanding the intricate wiring patterns of the brain. Existing automatic neuron segmentation methods rely on traditional clustering algorithms, where affinities are predicted first, and then watershed and…

2025

GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing

CVPR 2025poster

Scene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promise in text generation, they still struggle to produce high-quality results. Thes…

2025

Improving Accuracy and Calibration via Differentiated Deep Mutual Learning

CVPR 2025poster

Deep Neural Networks (DNNs) have achieved remarkable success in a variety of tasks, particularly in terms of prediction accuracy. However, in real-world scenarios, especially in safety-critical applications, accuracy alone is insufficient; reliable uncertainty estimates are essential. Modern DNNs, o…

Cited by 0SourcePDFScholar
2025

MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting

CVPR 2025poster

Advancements in generative models have enabled image inpainting models to generate content within specific regions of an image based on provided prompts and masks. However, existing inpainting methods often suffer from problems such as semantic misalignment, structural distortion, and style inconsis…

Cited by 0SourcePDFScholar
2025

PBCAT: Patch-Based Composite Adversarial Training against Physically Realizable Attacks on Object Detection

ICCV 2025poster

Object detection plays a crucial role in many security-sensitive applications, such as autonomous driving and video surveillance. However, several recent studies have shown that object detectors can be easily fooled by physically realizable attacks, e.g., adversarial patches and recent adversarial t…

Cited by 0SourcePDFScholar
2025

PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning

COLING 2025main

Low-rank adaptation (LoRA) and its variants have recently gained much interest due to their ability to avoid excessive inference costs. However, LoRA still encounters the following challenges: (1) Limitation of low-rank assumption; and (2) Its initialization method may be suboptimal. To this end, we…

Cited by 3SourcePDFScholar
2025

SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Model

CVPR 2025poster

Interactive segmentation is to segment the mask of the target object according to the user's interactive prompts. There are two mainstream strategies: early fusion and late fusion. Current specialist models utilize the early fusion strategy that encodes the combination of images and prompts to targe…

Cited by 0SourcePDFScholar
2025

SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

ICLR 2025poster

Systematic evaluation of speech separation and enhancement models under moving sound source conditions requires extensive and diverse data. However, real-world datasets often lack sufficient data for training and evaluation, and synthetic datasets, while larger, lack acoustic realism. Consequently,…

2025

Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology

AAAI 2025technical

Zeroth-order (ZO) optimization as the gradient-free method has become a powerful tool when the first-order gradient is unavailable or expensive to obtain, especially in decentralized learning scenarios where data and computational resources are distributed across multiple clients. There have been ma…

Cited by 0SourcePDFScholar
2025

TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation

ICLR 2025poster

In recent years, much speech separation research has focused primarily on improving model performance. However, for low-latency speech processing systems, high efficiency is equally important. Therefore, we propose a speech separation model with significantly reduced parameters and computational cos…

Cited by 0SourcePDFScholar
2025

Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and Optimization

IJCAI 2025

Large Language Models (LLMs), built on Transformer architectures, exhibit remarkable generalization across a wide range of tasks. However, fine-tuning these models for specific tasks remains resource-intensive due to their extensive parameterization. In this paper, we explore two remarkable phenomen

Cited by 0SourcePDFScholar
2025

Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization

ICLR 2025poster

Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: \textbf{(a) Limited \textit{i.i.d.} Setting.} Most studies focus on supervised function learning tasks where prompts are co…

Cited by 0SourcePDFScholar
2024

Bridging the Sim-to-Real Gap from the Information Bottleneck Perspective

CoRL 2024poster

Reinforcement Learning (RL) has recently achieved remarkable success in robotic control. However, most works in RL operate in simulated environments where privileged knowledge (e.g., dynamics, surroundings, terrains) is readily available. Conversely, in real-world scenarios, robot agents usually rel…

Cited by 9SourcecodeScholar
2024

Controllable Navigation Instruction Generation with Chain of Thought Prompting

ECCV 2024poster

"Instruction generation is a vital and multidisciplinary research area with broad applications. Existing instruction generation models are limited to generating instructions in a single style from a particular dataset, and the style and content of generated instructions cannot be controlled. Moreove…

2024

CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics

NeurIPS 2024spotlight

Enabling humanoid robots to clean rooms has long been a pursued dream within humanoid research communities. However, many tasks require multi-humanoid collaboration, such as carrying large and heavy furniture together. Given the scarcity of motion capture data on multi-humanoid collaboration and the…

Cited by 9SourcePDFScholar
2024

Enhancing In-Context Learning Performance with just SVD-Based Weight Pruning: A Theoretical Perspective

NeurIPS 2024poster

Pre-trained large language models (LLMs) based on Transformer have demonstrated striking in-context learning (ICL) abilities. With a few demonstration input-label pairs, they can predict the label for an unseen input without any parameter updates. In this paper, we show an exciting phenomenon that…

2024

Full-Distance Evasion of Pedestrian Detectors in the Physical World

NeurIPS 2024poster

Many studies have proposed attack methods to generate adversarial patterns for evading pedestrian detection, alarming the computer vision community about the need for more attention to the robustness of detectors. However, adversarial patterns optimized by these methods commonly have limited perform…

2024

IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation

ICML 2024poster

Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without employing selective attention mechanisms, which is in sharp contras…

2024

Language-Driven Anchors for Zero-Shot Adversarial Robustness

CVPR 2024poster

Deep Neural Networks (DNNs) are known to be susceptible to adversarial attacks. Previous researches mainly focus on improving adversarial robustness in the fully supervised setting leaving the challenging domain of zero-shot adversarial robustness an open question. In this work we investigate this d…

2024

PartImageNet++ Dataset: Scaling up Part-based Models for Robust Recognition

ECCV 2024poster

"Deep learning-based object recognition systems can be easily fooled by various adversarial perturbations. One reason for the weak robustness may be that they do not have part-based inductive bias like the human recognition process. Motivated by this, several part-based recognition models have been…

2024

RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation

ICLR 2024poster

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA) models operate in the time domain. However, their overly sim…

2024

SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object Detection

CVPR 2024poster

LiDAR-based 3D object detection plays an essential role in autonomous driving. Existing high-performing 3D object detectors usually build dense feature maps in the backbone network and prediction head. However the computational costs introduced by the dense feature maps grow quadratically as the per…

2023

An efficient encoder-decoder architecture with top-down attention for speech separation

ICLR 2023poster

Deep neural networks have shown excellent prospects in speech separation tasks. However, obtaining good results while keeping a low model complexity remains challenging in real-world applications. In this paper, we provide a bio-inspired efficient encoder-decoder architecture by mimicking the brain’…

2023

Generalization Bounds for Federated Learning: Fast Rates, Unparticipating Clients and Unbounded Losses

ICLR 2023poster

In {federated learning}, the underlying data distributions may be different across clients. This paper provides a theoretical analysis of generalization error of {federated learning}, which captures both heterogeneity and relatedness of the distributions. In particular, we assume that the heterogene…

Cited by 18SourcePDFScholar
2023

HEDNet: A Hierarchical Encoder-Decoder Network for 3D Object Detection in Point Clouds

NeurIPS 2023poster

3D object detection in point clouds is important for autonomous driving systems. A primary challenge in 3D object detection stems from the sparse distribution of points within the 3D scene. Existing high-performance methods typically employ 3D sparse convolutional neural networks with small kernels…

2023

NP-SemiSeg: When Neural Processes meet Semi-Supervised Semantic Segmentation

ICML 2023poster

Semi-supervised semantic segmentation involves assigning pixel-wise labels to unlabeled images at training time. This is useful in a wide range of real-world applications where collecting pixel-wise labels is not feasible in time or cost. Current approaches to semi-supervised semantic segmentation w…

2023

Physically Realizable Natural-Looking Clothing Textures Evade Person Detectors via 3D Modeling

CVPR 2023poster

Recent works have proposed to craft adversarial clothes for evading person detectors, while they are either only effective at limited viewing angles or very conspicuous to humans. We aim to craft adversarial texture for clothes based on 3D modeling, an idea that has been used to craft rigid adversar…

2022

Active Pointly-Supervised Instance Segmentation

ECCV 2022poster

"The requirement of expensive annotations is a major burden for training a well-performed instance segmentation model. In this paper, we present an economic active learning setting, named active pointly-supervised instance segmentation (APIS), which starts with box-level annotations and iteratively…

2022

Adversarial Texture for Fooling Person Detectors in the Physical World

CVPR 2022oral

Nowadays, cameras equipped with AI systems can capture and analyze images to detect people automatically. However, the AI system can make mistakes when receiving deliberately designed patterns in the real world, i.e., physical adversarial examples. Prior works have shown that it is possible to print…

Cited by 147PDFcodeScholar
2022

Infrared Invisible Clothing: Hiding From Infrared Detectors at Multiple Angles in Real World

CVPR 2022oral

Thermal infrared imaging is widely used in body temperature measurement, security monitoring, and so on, but its safety research attracted attention only in recent years. We proposed the infrared adversarial clothing, which could fool infrared pedestrian detectors at different angles. We simulated t…

Cited by 73PDFScholar
2022

NP-Match: When Neural Processes meet Semi-Supervised Learning

ICML 2022spotlight

Semi-supervised learning (SSL) has been widely explored in recent years, and it is an effective way of leveraging unlabeled data to reduce the reliance on labeled data. In this work, we adjust neural processes (NPs) to the semi-supervised image classification task, resulting in a new method named NP…

2021

Attack on Practical Speaker Verification System Using Universal Adversarial Perturbations

ICASSP 2021accepted

In authentication scenarios, applications of practical speaker verification systems usually require a person to read a dynamic authentication text. Previous studies played an audio adversarial example as a digital signal to perform physical attacks, which would be easily rejected by audio replay det…

Cited by 0SourceScholar
2021

CloudAAE: Learning 6D Object Pose Regression with On-line Data Synthesis on Point Clouds

ICRA 2021poster

It is often desired to train 6D pose estimation systems on synthetic data because manual annotation is expensive. However, due to the large domain gap between the synthetic and real images, synthesizing color images is expensive. In contrast, this domain gap is considerably smaller and easier to fil…

Cited by 58SourcecodeScholar
2021

DAM: Discrepancy Alignment Metric for Face Recognition

ICCV 2021poster

The field of face recognition (FR) has witnessed remarkable progress with the surge of deep learning. The effective loss functions play an important role for FR. In this paper, we observe that a majority of loss functions, including the widespread triplet loss and softmax-based cross-entropy loss, e…

Cited by 22PDFScholar
2021

Fooling Thermal Infrared Pedestrian Detectors in Real World Using Small Bulbs

AAAI 2021technical

Thermal infrared detection systems play an important role in many areas such as night security, autonomous driving, and body temperature detection. They have the unique advantages of passive imaging, temperature sensitivity and penetration. But the security of these systems themselves has not been f…

Cited by 95SourcePDFScholar
2021

Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object Detection

CVPR 2021poster

Localization Quality Estimation (LQE) is crucial and popular in the recent advancement of dense object detectors since it can provide accurate ranking scores that benefit the Non-Maximum Suppression processing and improve detection performance. As a common practice, most existing methods predict LQE…

Cited by 324PDFcodeScholar
2021

Look Closer To Segment Better: Boundary Patch Refinement for Instance Segmentation

CVPR 2021poster

Tremendous efforts have been made on instance segmentation but the mask quality is still not satisfactory. The boundaries of predicted instance masks are usually imprecise due to the low spatial resolution of feature maps and the imbalance problem caused by the extremely low proportion of boundary p…

Cited by 114PDFcodeScholar
2021

RSG: A Simple but Effective Module for Learning Imbalanced Datasets

CVPR 2021poster

Imbalanced datasets widely exist in practice and are a great challenge for training deep neural models with a good generalization on infrequent classes. In this work, we propose a new rare-class sample generator (RSG) to solve this problem. RSG aims to generate some new samples for rare classes duri…

Cited by 128PDFcodeScholar
2021

RefineMask: Towards High-Quality Instance Segmentation With Fine-Grained Features

CVPR 2021poster

The two-stage methods for instance segmentation, e.g. Mask R-CNN, have achieved excellent performance recently. However, the segmented masks are still very coarse due to the downsampling operations in both the feature pyramid and the instance-wise pooling process, especially for large objects. In th…

Cited by 154PDFcodeScholar
2021

Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural Network

NeurIPS 2021poster

Recent advances in the design of neural network architectures, in particular those specialized in modeling sequences, have provided significant improvements in speech separation performance. In this work, we propose to use a bio-inspired architecture called Fully Recurrent Convolutional Neural Netwo…

2020

6D Object Pose Regression via Supervised Learning on Point Clouds

ICRA 2020poster

This paper addresses the task of estimating the 6 degrees of freedom pose of a known 3D object from depth information represented by a point cloud. Deep features learned by convolutional neural networks from color information have been the dominant features to be used for inferring object poses, whi…

Cited by 111SourcecodeScholar
2020

Boosting Decision-based Black-box Adversarial Attacks with Random Sign Flip

ECCV 2020poster

Decision-based black-box adversarial attacks (decision-based attack) pose a severe threat to current deep neural networks, as they only need the predicted label of the target model to craft adversarial examples. However, existing decision-based attacks perform poorly on the $ l_\infty $ setting and…

Cited by 83SourcePDFScholar
2020

Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection

NeurIPS 2020poster

One-stage detector basically formulates object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is…

2020

Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement Learning

NeurIPS 2020spotlight

Goal-conditioned hierarchical reinforcement learning (HRL) is a promising approach for scaling up reinforcement learning (RL) techniques. However, it often suffers from training inefficiency as the action space of the high-level, i.e., the goal space, is often large. Searching in a large goal space…

2020

Online Knowledge Distillation via Collaborative Learning

CVPR 2020oral

This work presents an efficient yet effective online Knowledge Distillation method via Collaborative Learning, termed KDCL, which is able to consistently improve the generalization ability of deep neural networks (DNNs) that have different learning capacities. Unlike existing two-stage knowledge dis…

Cited by 395PDFScholar
2020

Rotation Consistent Margin Loss for Efficient Low-Bit Face Recognition

CVPR 2020poster

In this paper, we consider the low-bit quantization problem of face recognition (FR) under the open-set protocol. Different from well explored low-bit quantization on closed-set image classification task, the open-set task is more sensitive to quantization errors (QEs). We redefine the QEs in angula…

Cited by 50PDFScholar
2019

Improving Pedestrian Attribute Recognition With Weakly-Supervised Multi-Scale Attribute-Specific Localization

ICCV 2019poster

Pedestrian attribute recognition has been an emerging research topic in the area of video surveillance. To predict the existence of a particular attribute, it is demanded to localize the regions related to the attribute. However, in this task, the region annotations are not available. How to carve o…

Cited by 172PDFcodeScholar
2019

Knowledge Distillation via Route Constrained Optimization

ICCV 2019oral

Distillation-based learning boosts the performance of the miniaturized neural network based on the hypothesis that the representation of a teacher model can be used as structured and relatively weak supervision, and thus would be easily learned by a miniaturized model. However, we find that the repr…

Cited by 229PDFScholar
2019

Understanding the Disharmony Between Dropout and Batch Normalization by Variance Shift

CVPR 2019poster

This paper first answers the question "why do the two most powerful techniques Dropout and Batch Normalization (BN) often lead to a worse performance when they are combined together in many modern neural networks, but cooperate well sometimes as in Wide ResNet (WRN)?" in both theoretical and empiric…

Cited by 415PDFScholar
2018

Boosting Adversarial Attacks With Momentum

CVPR 2018poster

Deep neural networks are vulnerable to adversarial examples, which poses security concerns on these algorithms due to the potentially severe consequences. Adversarial attacks serve as an important surrogate to evaluate the robustness of deep learning models before they are deployed. However, most of…

Cited by 3515SourcePDFScholar
2018

Defense Against Adversarial Attacks Using High-Level Representation Guided Denoiser

CVPR 2018poster

Neural networks are vulnerable to adversarial examples, which poses a threat to their application in security sensitive systems. We propose high-level representation guided denoiser (HGD) as a defense for image classification. Standard denoiser suffers from the error amplification effect, in which s…

2018

High Performance Visual Tracking With Siamese Region Proposal Network

CVPR 2018poster

Visual object tracking has been a fundamental topic in recent years and many deep learning based trackers have achieved state-of-the-art performance on multiple benchmarks. However, most of these trackers can hardly get top performance with real-time speed. In this paper, we propose the Siamese regi…

Cited by 3306SourcePDFScholar
2018

Interpret Neural Networks by Identifying Critical Data Routing Paths

CVPR 2018poster

Interpretability of a deep neural network aims to explain the rationale behind its decisions and enable the users to understand the intelligent agents, which has become an important issue due to its importance in practical applications. To address this issue, we develop a Distillation Guided Routing…

Cited by 122SourcePDFScholar
2015

Convolutional Neural Networks with Intra-Layer Recurrent Connections for Scene Labeling

NeurIPS 2015poster

Scene labeling is a challenging computer vision task. It requires the use of both local discriminative features and global context information. We adopt a deep recurrent convolutional neural network (RCNN) for this task, which is originally proposed for object recognition. Different from traditional…

Cited by 80SourcePDFScholar