← Search

Bin Xiao

63 accepted papers

2026

Clear Nights Ahead: Towards Multi-Weather Nighttime Image Restoration

AAAI 2026technical

Restoring nighttime images affected by multiple adverse weather conditions is a practical yet under-explored research problem, as multiple weather degradations usually coexist in the real world alongside various lighting effects at night. This paper first explores the challenging multi-weather night

Cited by 0SourcePDFScholar
2026

ClinAlign: Clinical Workflow Aligned Memory Retrieval for Radiology Report Generation

IJCAI 2026

Automated radiology report generation aims to create clear and clinically correct diagnostic reports from medical images. Existing retrieval enhancement methods primarily focus on reusing textual knowledge, neglecting the crucial role of local visual pattern memory in clinical diagnosis. Furthermore

Cited by 0Scholar
2026

Diversity over Uniformity: Rethinking Representation in Generated Image Detection

CVPR 2026

With the rapid advancement of generative models, generated image detection has become an important task in visual forensics. Although existing methods have achieved remarkable progress, they often rely, after training, on only a small subset of highly salient forgery cues, which limits their ability

Cited by 0SourcecodeScholar
2026

PureProof: Diffusion-Resistant Black-box Targeted Attack on Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (VLMs) are increasingly deployed across diverse applications, such as AI agents, yet remain vulnerable to targeted adversarial attacks. However, the practical robustness of such attacks often remains unclear with limited evaluation under defenses. Diffusion-based purific

Cited by 0SourceScholar
2026

SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised Segmentation

AAAI 2026technical

In recent years, Contrastive Language-Image Pretraining (CLIP) has been widely applied to Weakly Supervised Semantic Segmentation (WSSS) tasks due to its powerful cross-modal semantic understanding capabilities. This paper proposes a novel Semantic and Spatial Rectification (SSR) method to address t

Cited by 0SourcePDFScholar
2026

TGDD: Trajectory Guided Dataset Distillation with Balanced Distribution

AAAI 2026technical

Dataset distillation compresses large datasets into compact synthetic ones to reduce storage and computational costs. Among various approaches, distribution matching (DM)-based methods have attracted attention for their high efficiency. However, they often overlook the evolution of feature represent

Cited by 0SourcePDFScholar
2025

Breaking Grid Constraints: Dynamic Graph Reconstruction Network for Multi-organ Segmentation

ICCV 2025poster

Morphological differences and dense spatial relations of organs make multi-organ segmentation challenging. Current segmentation networks, primarily based on CNNs and Transformers, represent organs by aggregating information within fixed regions. However, aggregated representations often fail to accu…

2025

Continual Knowledge Adaptation for Reinforcement Learning

NeurIPS 2025poster

Reinforcement Learning enables agents to learn optimal behaviors through interactions with environments. However, real-world environments are typically non-stationary, requiring agents to continuously adapt to new tasks and changing conditions. Although Continual Reinforcement Learning facilitates l…

Cited by 0SourcecodeScholar
2025

Covert and Potent: A Weather-Camouflaged Backdoor Attacks on Self-Supervised Learning

ICASSP 2025accepted

Self-supervised learning is widely applied across various domains due to its advantage of learning data representations without the need for labels. However, recent research shows that backdoor attacks on self-supervised learning are achievable by coupling benign features with trigger features witho…

Cited by 0SourceScholar
2025

CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training

AAAI 2025technical

Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality cust…

2025

Efficient Dynamic Ensembling for Multiple LLM Experts

IJCAI 2025

LLMs have demonstrated impressive performance across various language tasks. However, the strengths of LLMs can vary due to different architectures, model sizes, areas of training data, etc. Therefore, ensemble reasoning for the strengths of different LLM experts is critical to achieving consistent

2025

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

CVPR 2025poster

We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2, a generative vision foundation model. Unlike the widely used CLIP-style vision transformer trained by contrastive learning, Florence-2 can capture different l…

2025

Improving Transferable Targeted Attacks with Feature Tuning Mixup

CVPR 2025poster

Deep neural networks (DNNs) exhibit vulnerability to adversarial examples that can transfer across different DNN models. A particularly challenging problem is developing transferable targeted attacks that can mislead DNN models into predicting specific target classes. While various methods have been…

2025

LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language Models

NeurIPS 2025poster

Large vision-language models (VLLMs) have driven significant progress in multi-modal systems, enabling a wide range of applications across domains such as healthcare, education, and content generation. Despite the success, the large-scale datasets used to train these models often contain sensitive o…

Cited by 0SourceScholar
2025

Learning Preconditioners in Gates-controlled Deep Unfolding Networks based on Quasi-Newton Methods For Accelerated MRI Reconstruction

ICASSP 2025accepted

Deep unfolding networks (DUNs) have made significant progress in MRI reconstruction, successfully tackling the problem of prolonged imaging time. However, the ill-conditioned nature of MRI reconstruction often causes slow convergence in iterative optimization, potentially compromising the performanc…

Cited by 0SourceScholar
2025

PLA: Prompt Learning Attack against Text-to-Image Generative Models

ICCV 2025poster

Text-to-Image (T2I) models have gained widespread adoption across various applications. Despite the success, the potential misuse of T2I models poses significant risks of generating Not-Safe-For-Work (NSFW) content. To investigate the vulnerability of T2I models, this paper delves into adversarial a…

2025

Power of Diversity: Enhancing Data-Free Black-Box Attack with Domain-Augmented Learning

AAAI 2025technical

Substitute training-based data-free black-box attacks pose a significant threat to enterprise-deployed models. These attacks use a generator to synthesize data and query APIs, then train a substitute model to approximate the target model's decision boundary based on the returned results. However, ex…

Cited by 0SourcePDFScholar
2025

Stacking U-Nets in U-shape: Redesigning the Information Flow in Model-based Networks for MRI Reconstruction

ICASSP 2025accepted

Model-based networks have shown convincing performance in MRI reconstruction. However, the unrolled cascades within the networks are constrained to solely obtain information from the preceding counterpart, resulting in potential error accumulation. Moreover, the linear structure fails to address the…

Cited by 0SourceScholar
2025

StyleGuard: Preventing Text-to-Image-Model-based Style Mimicry Attacks by Style Perturbations

NeurIPS 2025poster

Recently, text-to-image diffusion models have been widely used for style mimicry and personalized customization through methods such as DreamBooth and Textual Inversion. This has raised concerns about intellectual property protection and the generation of deceptive content. Recent studies, such as G…

Cited by 0SourcecodeScholar
2025

Subsampling Decomposition based k-Space Refinement for Accelerated MRI Reconstruction

ICASSP 2025accepted

In accelerated MRI reconstruction problem, directly recovering all the missing k-space data from undersampled measurements is highly ill-posed and often leads to suboptimal performance. To address the problem, we propose a novel deep unfolding network (DUN) with subsampling decomposition (SD) based…

Cited by 0SourceScholar
2025

Test-Time Learning for Large Language Models

ICML 2025poster

While Large Language Models (LLMs) have exhibited remarkable emergent capabilities through extensive pre-training, they still face critical limitations in generalizing to specialized domains and handling diverse linguistic variations, known as distribution shifts. In this paper, we propose a Test-T…

Cited by 0SourcePDFScholar
2025

Towards Universal AI-Generated Image Detection by Variational Information Bottleneck Network

CVPR 2025poster

The rapid advancement of generative models has significantly improved the quality of generated images. Meanwhile, it challenges information authenticity and credibility. Current generated image detection methods based on large-scale pre-trained multimodal models have achieved impressive results. Alt…

2025

UV-Attack: Physical-World Adversarial Attacks on Person Detection via Dynamic-NeRF-based UV Mapping

ICLR 2025poster

Recent works have attacked person detectors using adversarial patches or static-3D-model-based texture modifications. However, these methods suffer from low attack success rates when faced with significant human movements. The primary challenge stems from the highly non-rigid nature of the human bod…

2025

Who Controls the Authorization? Invertible Networks for Copyright Protection in Text-to-Image Synthesis

ICCV 2025poster

To defend against personalized generation, a new form of infringement that is more concealed and destructive, the existing copyright protection methods is to add adversarial perturbations in images. However, these methods focus solely on countering illegal personalization, neglecting the requirement…

Cited by 0SourcePDFScholar
2024

AdvDiff: Generating Unrestricted Adversarial Examples using Diffusion Models

ECCV 2024poster

"Unrestricted adversarial attacks present a serious threat to deep learning models and adversarial defense techniques. They pose severe security problems for deep learning applications because they can effectively bypass defense mechanisms. However, previous attack methods often directly inject Proj…

2024

Efficient Modulation for Vision Networks

ICLR 2024poster

In this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context modeling and feature projection layers, and fuses features via element-wise multiplication and an MLP block. We demonst…

2024

Facial Aesthetic Enhancement Network for Asian Faces Based on Differential Facial Aesthetic Activations

ICASSP 2024accepted

In this paper, we addressed facial aesthetic enhancement (FAE). Although existing methods have made great progress, the beautified images generated by them are highly prone to poor beautification, which limits their application to real-world scenes. To tackle this problem, we proposed a new method c…

Cited by 0SourceScholar
2024

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

CVPR 2024poster

We introduce Florence-2 a novel vision foundation model with a unified prompt-based representation for various computer vision and vision-language tasks. While existing large vision models excel in transfer learning they struggle to perform diverse tasks with simple instructions a capability that im…

2024

Using My Artistic Style? You Must Obtain My Authorization

ECCV 2024poster

"Artistic images typically contain the unique creative styles of artists. However, it is easy to transfer an artist’s style to arbitrary target images using style transfer techniques. To protect styles, some researchers use adversarial attacks to safeguard artists’ artistic style images. Prior metho…

2024

i-Code V2: An Autoregressive Generation Framework over Vision, Language, and Speech Data

NAACL 2024findings

The convergence of text, visual, and audio data is crucial towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encoder-only models that lack generative abilities. We propose closing this gap with i-Code V2, one of the first models capable…

Cited by 3SourcePDFScholar
2023

DAA: A Delta Age AdaIN Operation for Age Estimation via Binary Code Transformer

CVPR 2023poster

Naked eye recognition of age is usually based on comparison with the age of others. However, this idea is ignored by computer tasks because it is difficult to obtain representative contrast images of each age. Inspired by the transfer learning, we designed the Delta Age AdaIN (DAA) operation to obta…

2023

DLBD: A Self-Supervised Direct-Learned Binary Descriptor

CVPR 2023poster

For learning-based binary descriptors, the binarization process has not been well addressed. The reason is that the binarization blocks gradient back-propagation. Existing learning-based binary descriptors learn real-valued output, and then it is converted to binary descriptors by their proposed bin…

2023

MCF: Mutual Correction Framework for Semi-Supervised Medical Image Segmentation

CVPR 2023poster

Semi-supervised learning is a promising method for medical image segmentation under limited annotation. However, the model cognitive bias impairs the segmentation performance, especially for edge regions. Furthermore, current mainstream semi-supervised medical image segmentation (SSMIS) methods lack…

2023

Physical-World Optical Adversarial Attacks on 3D Face Recognition

CVPR 2023poster

The success rate of current adversarial attacks remains low on real-world 3D face recognition tasks because the 3D-printing attacks need to meet the requirement that the generated points should be adjacent to the surface, which limits the adversarial example' searching space. Additionally, they have…

2023

Self-Supervised Image Local Forgery Detection by JPEG Compression Trace

AAAI 2023technical

For image local forgery detection, the existing methods require a large amount of labeled data for training, and most of them cannot detect multiple types of forgery simultaneously. In this paper, we firstly analyzed the JPEG compression traces which are mainly caused by different JPEG compression c…

Cited by 7SourcePDFScholar
2023

TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight Inheritance

ICCV 2023poster

In this paper, we propose a novel cross-modal distillation method, called TinyCLIP, for large-scale language-image pre-trained models. The method introduces two core techniques: affinity mimicking and weight inheritance. Affinity mimicking explores the interaction between modalities during distillat…

Cited by 65PDFcodeScholar
2023

i-Code: An Integrative and Composable Multimodal Learning Framework

AAAI 2023technical

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the m…

2022

DaViT: Dual Attention Vision Transformers

ECCV 2022poster

"In this work, we introduce Dual Attention Vision Transformers (DaViT), a simple yet effective vision transformer architecture that is able to capture global context while maintaining computational efficiency. We propose approaching the problem from an orthogonal angle: exploiting self-attention mec…

2022

Efficient Self-supervised Vision Transformers for Representation Learning

ICLR 2022poster

This paper investigates two techniques for developing efficient self-supervised vision transformers (EsViT) for visual representation learning. First, we show through a comprehensive empirical study that multi-stage architectures with sparse self-attentions can significantly reduce modeling complexi…

2022

Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training

ECCV 2022poster

"Large-scale multi-modal contrastive pre-training has demonstrated great utility to learn transferable features for a range of downstream tasks by mapping multiple modalities into a shared embedding space. Typically, this has employed separate encoders for each modality. However, recent work suggest…

2022

MiniViT: Compressing Vision Transformers With Weight Multiplexing

CVPR 2022poster

Vision Transformer (ViT) models have recently drawn much attention in computer vision due to their high model capability. However, ViT models suffer from huge number of parameters, restricting their applicability on devices with limited computation. To alleviate this problem, we propose MiniViT, a n…

Cited by 162PDFcodeScholar
2022

TinyViT: Fast Pretraining Distillation for Small Vision Transformers

ECCV 2022poster

"Vision transformer (ViT) recently has drawn great attention in computer vision due to its remarkable model capability. However, most prevailing ViT models suffer from huge number of parameters, restricting their applicability on devices with limited resources. To alleviate this issue, we propose Ti…

2022

Unified Contrastive Learning in Image-Text-Label Space

CVPR 2022poster

Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs. While supervised learning may result in a more discriminative representation, language-image pretraining shows unprece…

Cited by 249PDFcodeScholar
2021

Bottom-Up Human Pose Estimation via Disentangled Keypoint Regression

CVPR 2021poster

In this paper, we are interested in the bottom-up paradigm of estimating human poses from an image. We study the dense keypoint regression framework that is previously inferior to the keypoint detection and grouping framework. Our motivation is that regressing keypoint positions accurately needs to…

Cited by 401PDFcodeScholar
2021

CvT: Introducing Convolutions to Vision Transformers

ICCV 2021poster

We present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convolutions into ViT to yield the best of both designs. This is accomplished through two primary modifications: a hierarchy…

Cited by 2598PDFcodeScholar
2021

DTMNet: A Discrete Tchebichef Moments-Based Deep Neural Network for Multi-Focus Image Fusion

ICCV 2021poster

Compared with traditional methods, the deep learning-based multi-focus image fusion methods can effectively improve the performance of image fusion tasks. However, the existing deep learning-based methods encounter a common issue of a large number of parameters, which leads to the deep learning mode…

Cited by 16PDFScholar
2021

Dynamic Head: Unifying Object Detection Heads With Attentions

CVPR 2021poster

The complex nature of combining localization and classification in object detection has resulted in the flourished development of methods. Previous works tried to improve the performance in various object detection heads but failed to present a unified view. In this paper, we present a novel dynamic…

Cited by 870PDFcodeScholar
2021

Focal Attention for Long-Range Interactions in Vision Transformers

NeurIPS 2021spotlight

Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability to capture local and global visual dependencies through self-attention is the key to its success. But it also brings challenges due to quadratic computational overhead, especially for…

Cited by 171SourcePDFScholar
2021

Lite-HRNet: A Lightweight High-Resolution Network

CVPR 2021poster

We present an efficient high-resolution network, Lite-HRNet, for human pose estimation. We start by simply applying the efficient shuffle block in ShuffleNet to HRNet (high-resolution network), yielding stronger performance over popular lightweight networks, such as MobileNet, ShuffleNet, and Small…

Cited by 503PDFcodeScholar
2021

Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding

ICCV 2021poster

This paper presents a new Vision Transformer (ViT) architecture Multi-Scale Vision Longformer, which significantly enhances the ViT of [??] for encoding high-resolution images using two techniques. The first is the multi-scale model structure, which provides image encodings at multiple scales with…

Cited by 419PDFcodeScholar
2021

Reality Transform Adversarial Generators for Image Splicing Forgery Detection and Localization

ICCV 2021poster

When many forged images become more and more realistic with the help of image editing tools and deep learning techniques, authenticators need to improve their ability to verify these forged images. The process of generating and detecting forged images is thus similar to the principle of Generative A…

Cited by 33PDFScholar
2020

HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation

CVPR 2020poster

Bottom-up human pose estimation methods have difficulties in predicting the correct pose for small persons due to challenges in scale variation. In this paper, we present HigherHRNet: a novel bottom-up human pose estimation method for learning scale-aware representations using high-resolution featur…

Cited by 1075PDFcodeScholar
2019

Deep High-Resolution Representation Learning for Human Pose Estimation

CVPR 2019poster

In this paper, we are interested in the human pose estimation problem with a focus on learning reliable high-resolution representations. Most existing methods recover high-resolution representations from low-resolution representations produced by a high-to-low resolution network. Instead, our propos…

Cited by 6183PDFcodeScholar