← Search

Han Xu

31 accepted papers

2026

Circular-DPO: Aligning Multi-Stage 3D Generative Models via Preference Feedback Loop

CVPR 2026

Multi-stage generative models have shown great promise in 3D content creation due to focused generation of structure or texture in different stages, but their outputs often fail to align with human preferences. The key bottleneck to apply alignment methods is the presence of non-differentiable opera

Cited by 0SourceScholar
2026

ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents

ICLR 2026poster

We introduce ComputerRL, a framework for autonomous desktop intelligence that enables agents to operate complex digital workspaces skillfully. ComputerRL features the API-GUI paradigm, which unifies programmatic API calls and direct GUI interaction to address the inherent mismatch between machine ag…

Cited by 0SourcecodeScholar
2026

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

CVPR 2026

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due to the scarcity of large-scale multi-sensor video datasets, l

Cited by 0SourcecodeScholar
2026

WaveC2R: Wavelet-Driven Coarse-to-Refined Hierarchical Learning for Radar Retrieval

AAAI 2026technical

Satellite-based radar retrieval methods are widely employed to fill coverage gaps in ground-based radar systems, especially in remote areas affected by terrain blockage and limited detection range. Existing methods predominantly rely on overly simplistic spatial-domain architectures constructed from

Cited by 0SourcePDFScholar
2025

Deno-IF: Unsupervised Noisy Visible and Infrared Image Fusion Method

NeurIPS 2025spotlight

Most image fusion methods are designed for ideal scenarios and struggle to handle noise. Existing noise-aware fusion methods are supervised and heavily rely on constructed paired data, limiting performance and generalization. This paper proposes a novel unsupervised noisy visible and infrared image…

Cited by 0SourcecodeScholar
2025

Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution

ICCV 2025poster

Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image super-resolution (SR). Despite some DWT-based methods improving SR by capturing fine-grained frequency signals, most existing approaches neglect the interrelations among multi-scale frequency sub-bands, res…

Cited by 0SourcePDFScholar
2025

LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up Tables

ICCV 2025poster

Current advanced research on infrared and visible image fusion primarily focuses on improving fusion performance, often neglecting the applicability on real-time fusion devices. In this paper, we propose a novel approach that towards extremely fast fusion via distillation to learnable lookup tables…

2025

Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Data

EMNLP 2025

Retrieval-augmented generation (RAG) enhances the outputs of language models by integrating relevant information retrieved from external knowledge sources. However, when the retrieval process involves private data, RAG systems may face severe privacy risks, potentially leading to the leakage of sens

2025

Red-Teaming LLM Multi-Agent Systems via Communication Attacks

ACL 2025finding

Large Language Model-based Multi-Agent Systems (LLM-MAS) have revolutionized complex problem-solving capability by enabling sophisticated agent collaboration through message-based communications. While the communication framework is crucial for agent coordination, it also introduces a critical yet u…

Cited by 0SourcePDFScholar
2024

A Robust Semantics-based Watermark for Large Language Model against Paraphrasing

NAACL 2024findings

Large language models (LLMs) have show their remarkable ability in various natural language tasks. However, there are concerns that LLMs are possible to be used improperly or even illegally. To prevent the malicious usage of LLMs, detecting LLM-generated text becomes crucial in the deployment of LLM…

2024

Certified Robustness for Deep Equilibrium Models via Serialized Random Smoothing

NeurIPS 2024poster

Implicit models such as Deep Equilibrium Models (DEQs) have emerged as promising alternative approaches for building deep neural networks. Their certified robustness has gained increasing research attention due to security concerns. Existing certified defenses for DEQs employing interval bound propa…

2024

Encoding Hierarchical Schema via Concept Flow for Multifaceted Ideology Detection

ACL 2024findings

Multifaceted ideology detection (MID) aims to detect the ideological leanings of texts towards multiple facets. Previous studies on ideology detection mainly focus on one generic facet and ignore label semantics and explanatory descriptions of ideologies, which are a kind of instructive information…

2024

Exploring Memorization in Fine-tuned Language Models

ACL 2024long

Large language models (LLMs) have shown great capabilities in various tasks but also exhibited memorization of training data, raising tremendous privacy and copyright concerns. While prior works have studied memorization during pre-training, the exploration of memorization during fine-tuning is rath…

Cited by 26SourcePDFScholar
2024

On the Generalization of Training-based ChatGPT Detection Methods

EMNLP 2024finding

Large language models, such as ChatGPT, achieve amazing performance on various language processing tasks. However, they can also be exploited for improper purposes such as plagiarism or misinformation dissemination. Thus, there is an urgent need to detect the texts generated by LLMs. One type of mos…

2024

RESTful-Llama: Connecting User Queries to RESTful APIs

EMNLP 2024industry

Recent advancements in Large Language Models (LLMs) have showcased exceptional performance in zero-shot learning and reasoning tasks. However, integrating these models with external tools - a crucial need for real-world applications - remains a significant challenge. We propose RESTful-Llama, a nove…

2024

Self-playing Adversarial Language Game Enhances LLM Reasoning

NeurIPS 2024poster

We explore the potential of self-play training for large language models (LLMs) in a two-player adversarial language game called Adversarial Taboo. In this game, an attacker and a defender communicate around a target word only visible to the attacker. The attacker aims to induce the defender to spea…

2024

Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

CVPR 2024poster

Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve th…

2024

The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)

ACL 2024findings

Retrieval-augmented generation (RAG) is a powerful technique to facilitate language model generation with proprietary and private data, where data privacy is a pivotal concern. Whereas extensive research has demonstrated the privacy risks of large language models (LLMs), the RAG technique could pote…

2024

Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis

EMNLP 2024main

Large language models (LLMs) are susceptible to a type of attack known as jailbreaking, which misleads LLMs to output harmful contents. Although there are diverse jailbreak attack strategies, there is no unified understanding on why some methods succeed and others fail. This paper explores the behav…

2024

Unveiling and Mitigating Memorization in Text-to-image Diffusion Models through Cross Attention

ECCV 2024poster

"Recent advancements in text-to-image (T2I) diffusion models have demonstrated their remarkable capability to generate high-quality images from textual prompts. However, increasing research indicates that these models memorize and replicate images from their training data, raising concerns about pot…

2023

Diff-Retinex: Rethinking Low-light Image Enhancement with A Generative Diffusion Model

ICCV 2023poster

In this paper, we rethink the low-light image enhancement task and propose a physically explainable and generative diffusion model for low-light image enhancement, termed as Diff-Retinex. We aim to integrate the advantages of the physical model and the generative network. Furthermore, we hope to sup…

Cited by 141PDFScholar
2023

Probabilistic Categorical Adversarial Attack and Adversarial Training

ICML 2023poster

The studies on adversarial attacks and defenses have greatly improved the robustness of Deep Neural Networks (DNNs). Most advanced approaches have been overwhelmingly designed for continuous data such as images. However, these achievements are still hard to be generalized to categorical data. To bri…

Cited by 14SourcePDFScholar
2023

Unsupervised Multi-Exposure Image Fusion Breaking Exposure Limits via Contrastive Learning

AAAI 2023technical

This paper proposes an unsupervised multi-exposure image fusion (MEF) method via contrastive learning, termed as MEF-CL. It breaks exposure limits and performance bottleneck faced by existing methods. MEF-CL firstly designs similarity constraints to preserve contents in source images. It eliminates…

2022

RFNet: Unsupervised Network for Mutually Reinforcing Multi-Modal Image Registration and Fusion

CVPR 2022poster

In this paper, we propose a novel method to realize multi-modal image registration and fusion in a mutually reinforcing framework, termed as RFNet. We handle the registration in a coarse-to-fine fashion. For the first time, we exploit the feedback of image fusion to promote the registration accuracy…

Cited by 126PDFcodeScholar
2021

Graph Neural Networks with Adaptive Residual

NeurIPS 2021poster

Graph neural networks (GNNs) have shown the power in graph representation learning for numerous tasks. In this work, we discover an interesting phenomenon that although residual connections in the message passing of GNNs help improve the performance, they immensely amplify GNNs' vulnerability agains…

2021

To be Robust or to be Fair: Towards Fairness in Adversarial Training

ICML 2021spotlight

Adversarial training algorithms have been proved to be reliable to improve machine learning models’ robustness against adversarial examples. However, we find that adversarial training algorithms tend to introduce severe disparity of accuracy and robustness between different groups of data. For insta…

Cited by 225SourcePDFScholar
2019

Effective and Stable Neuron Model Optimization Based on Aggregated CMA-ES

ICASSP 2019accepted

Computer simulations have facilitated our understanding of the dynamic behavior of the brain and the effect of the medical treatment such as deep brain stimulation. For improving the simulation model, it is essential to develop a method for optimizing parameters of a neuron model from available expe…

Cited by 0SourceScholar