← Search

Xi Li

72 accepted papers

2026

All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning

ICLR 2026poster

The rapid proliferation of AI-generated images (AIGIs) highlights the pressing demand for generalizable detection methods. In this paper, we establish two key principles for AIGI detection task through systematic analysis: **(1) All Patches Matter**, since the uniform generation process ensures that…

Cited by 0SourceScholar
2026

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

ICML 2026poster

As cyber threats continue to grow in scale and sophistication, blue team defenders increasingly require advanced tools to proactively detect and mitigate risks. Large Language Models (LLMs) offer promising capabilities for enhancing threat analysis. However, their effectiveness in real-world blue te…

Cited by 0SourceScholar
2026

DiscoX: Benchmarking Discourse-Level Translation in Expert Domains

ICLR 2026poster

The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. While these translations demand discourse-level coherence and strict terminological precision, current evaluation methods p…

Cited by 0SourcecodeScholar
2026

IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion

AAAI 2026technical

Reconstructing complete and interactive 3D scenes remains a fundamental challenge in computer vision and robotics, particularly due to persistent object occlusions and limited sensor coverage. Even multi-view observations from a single scene scan often fail to capture the full structural details. Ex

Cited by 0SourcePDFScholar
2026

InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution

ICML 2026poster

Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long sequences: (1) inefficiency due to the heavy cost of multi-step denoising for full-length sequences; and (2) poor consist…

Cited by 0SourceScholar
2026

Innovative Design of Multi-Functional Supernumerary Robotic Limbs with Ellipsoid Workspace Optimization

ICRA 2026poster

Supernumerary robotic limbs (SRL) offer substantial potential in both the rehabilitation of hemiplegic patients and the enhancement of functional capabilities for healthy individuals. Designing a general-purpose SRL device is inherently challenging, particularly when developing a unified theoretical…

2026

MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment

CVPR 2026

Multi-subject image generation aims to synthesize user-provided subjects in a single image while preserving subject fidelity, ensuring prompt consistency, and aligning with human aesthetic preferences. Existing In-Context-Learning based methods are limited by their highly coupled training paradigm.

Cited by 0SourceScholar
2026

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

ICLR 2026poster

Multimodal large language models are progressively advancing toward multimodal agents that can proactively execute tasks. Existing research on multimodal agents primarily targets either GUI or embodied scenarios, corresponding to interactions within 2D virtual world and 3D physical world, respective…

Cited by 0SourceScholar
2026

REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion

CVPR 2026

As an important and challenging problem in computer vision, Panoramic Semantic Segmentation (PASS) aims to provide complete scene perception based on an ultra-wide angle of view. Most PASS methods often focus on spherical geometry with RGB input or use the depth information in original or HHA format

Cited by 0SourceScholar
2026

UniScene-MoTion: Unified Scene & Motion-aware Diffusion Transition Framework

AAAI 2026technical

Video transitions are critical for ensuring temporal coherence in edited media, yet existing methods often rely on handcrafted effects or relative-scale trajectories that fail to capture the physical structure of real-world scenes. In this work, we introduce a scale-aware video transition framework

Cited by 0SourcePDFScholar
2025

AAAR-1.0: Assessing AI’s Potential to Assist Research

ICML 2025poster

Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face unique challenges and opportunities in leveraging LLMs for the…

Cited by 0SourcePDFScholar
2025

Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training

ICCV 2025poster

Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a high SNR presents a semantic density gap, leading to visual alignment bias. In…

2025

Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models

ACL 2025finding

Large Language Models (LLMs), especially those accessed via APIs, have demonstrated impressive capabilities across various domains. However, users without technical expertise often turn to (untrustworthy) third-party services, such as prompt engineering, to enhance their LLM experience, creating vul…

2025

Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and Algorithm

NeurIPS 2025poster

Existing multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three com…

Cited by 0SourcecodeScholar
2025

CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities

AAAI 2025technical

Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of video diffusion models (VDMs) to combine concepts and generate…

2025

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection

ICCV 2025poster

The Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models, remains unexplored. This work investigates this domain, reveal…

2025

Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models

AAAI 2025technical

The rapid advancement of pretrained text-driven diffusion models has significantly enriched applications in image generation and editing. However, as the demand for personalized content editing increases, new challenges emerge especially when dealing with arbitrary objects and complex scenes. Existi…

2025

Envisioning Class Entity Reasoning by Large Language Models for Few-shot Learning

AAAI 2025technical

Few-shot learning (FSL) aims to recognize new concepts using a limited number of visual samples. Existing methods attempt to incorporate semantic information into the limited visual data for category understanding. However, these methods often enrich class-level feature representations with abstract…

Cited by 9SourcePDFScholar
2025

Experimental Evaluation of Radio-aware Semantic Map with 5G-Enabled Mobile Robots

IROS 2025

With the rapid development of 5G technology and the increasing demand for autonomous mobile robots, there is a trend to leverage the ultra-low latency, high data rates, and reliable wireless connectivity offered by 5G to improve the perception and navigation of robots in unknown environments. This p

Cited by 0SourceScholar
2025

Exploring Unbiased Deepfake Detection via Token-Level Shuffling and Mixing

AAAI 2025technical

The generalization problem is broadly recognized as a critical challenge in detecting deepfakes. Most previous work believes that the generalization gap is caused by the differences among various forgery methods. However, our investigation reveals that the generalization issue can still occur when f…

Cited by 2SourcePDFScholar
2025

Incorporating Review-missing Interactions for Generative Explainable Recommendation

COLING 2025main

Explainable recommendation has attracted much attention from the academic and industry communities. Traditional models usually leverage user reviews as ground truths for model training, and the interactions without reviews are totally ignored. However, in practice, a large amount of users may not le…

Cited by 0SourcePDFScholar
2025

Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language Model

CVPR 2025poster

Large Vision-Language Models (LVLMs) have achieved significant progress in recent years. However, the expensive inference cost limits the realistic deployment of LVLMs. Some works find that visual tokens are redundant and compress tokens to reduce the inference cost. These works identify important n…

2025

PiD: Generalized AI-Generated Images Detection with Pixelwise Decomposition Residuals

ICML 2025poster

Fake images, created by recently advanced generative models, have become increasingly indistinguishable from real ones, making their detection crucial, urgent, and challenging. This paper introduces PiD (Pixelwise Decomposition Residuals), a novel detection method that focuses on residual signals wi…

Cited by 0SourcePDFScholar
2025

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

ICCV 2025poster

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges, as users often struggle to provide precise camera parameter…

Cited by 0SourcePDFScholar
2025

Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model

ICLR 2025poster

The Mixture-of-Experts (MoE) has gained increasing attention in studying Large Vision-Language Models (LVLMs). It uses a sparse model to replace the dense model, achieving comparable performance while activating fewer parameters during inference, thus significantly reducing the inference cost. Exist…

2024

BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-based Roadside 3D Object Detection

CVPR 2024poster

Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain since it encompasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping igno…

2024

Enhancing Text-to-SQL Capabilities of Large Language Models through Tailored Promptings

COLING 2024main

Large language models (LLMs) with prompting have achieved encouraging results on many natural language processing (NLP) tasks based on task-tailored promptings. Text-to-SQL is a critical task that generates SQL queries from natural language questions. However, prompting on LLMs haven’t show superior…

Cited by 10SourcePDFScholar
2024

Improving Zero-Shot Generalization for CLIP with Variational Adapter

ECCV 2024poster

"The excellent generalization capability of pre-trained Vision-Language Models (VLMs) makes fine-tuning VLMs for downstream zero-shot tasks a popular choice. Despite achieving promising performance in the professionality of base classes, most existing fine-tuned methods suffer from feature confusion…

Cited by 8SourcePDFScholar
2024

Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

AAAI 2024technical

Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under the demanding zero-shot and few-shot scenarios. To achieve…

2024

Rethinking Normalization Layers for Domain Generalizable Person Re-identification

ECCV 2024poster

"Domain Generalizable Person Re-Identification (DG-ReID) strives to transfer learned feature representation from source domains to unseen target domains, despite significant distribution shifts. While most existing methods enhance model generalization and discriminative feature extraction capability…

2024

ScanFormer: Referring Expression Comprehension by Iteratively Scanning

CVPR 2024poster

Referring Expression Comprehension (REC) aims to localize the target objects specified by free-form natural language descriptions in images. While state-of-the-art methods achieve impressive performance they perform a dense perception of images which incorporates redundant visual regions unrelated t…

Cited by 10SourcePDFScholar
2024

SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model

AAAI 2024technical

Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduc…

Cited by 8SourcePDFScholar
2024

Temporal-Distributed Backdoor Attack against Video Based Action Recognition

AAAI 2024technical

Deep neural networks (DNNs) have achieved tremendous success in various applications including video action recognition, yet remain vulnerable to backdoor attacks (Trojans). The backdoor-compromised model will mis-classify to the target class chosen by the attacker when a test instance (from a non-t…

Cited by 8SourcePDFScholar
2024

Virtual Immunohistochemistry Staining for Histological Images Assisted by Weakly-supervised Learning

CVPR 2024poster

Recently virtual staining technology has greatly promoted the advancement of histopathology. Despite the practical successes achieved the outstanding performance of most virtual staining methods relies on hard-to-obtain paired images in training. In this paper we propose a method for virtual immunoh…

2023

Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object Detection

ICCV 2023poster

Knowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to t…

Cited by 31PDFcodeScholar
2023

DeSTSeg: Segmentation Guided Denoising Student-Teacher for Anomaly Detection

CVPR 2023poster

Visual anomaly detection, an important problem in computer vision, is usually formulated as a one-class classification and segmentation task. The student-teacher (S-T) framework has proved to be effective in solving this challenge. However, previous works based on S-T only empirically applied constr…

2023

DenseDINO: Boosting Dense Self-Supervised Learning with Token-Based Point-Level Consistency

IJCAI 2023poster

In this paper, we propose a simple yet effective transformer framework for self-supervised learning called DenseDINO to learn dense visual representations. To exploit the spatial information that the dense prediction tasks require but neglected by the existing self-supervised transformers, we introd…

Cited by 4SourcePDFScholar
2023

Enhancing 5G-Enabled Robots Autonomy by Radio-Aware Semantic Maps

IROS 2023poster

Future robotics systems aiming for true autonomy must be robust against dynamic and unstructured environments. The 5th generation (5G) mobile network is expected to provide ubiquitous, reliable and low-latency wireless communications to ground robots, especially in outdoor scenarios. Empowered by 5G…

Cited by 3SourceScholar
2023

GaitGCI: Generative Counterfactual Intervention for Gait Recognition

CVPR 2023poster

Gait is one of the most promising biometrics that aims to identify pedestrians from their walking patterns. However, prevailing methods are susceptible to confounders, resulting in the networks hardly focusing on the regions that reflect effective walking patterns. To address this fundamental proble…

Cited by 59SourcePDFScholar
2023

Language Adaptive Weight Generation for Multi-Task Visual Grounding

CVPR 2023poster

Although the impressive performance in visual grounding, the prevailing approaches usually exploit the visual backbone in a passive way, i.e., the visual backbone extracts features with fixed weights without expression-related hints. The passive perception may lead to mismatches (e.g., redundant and…

2023

LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image Generation

CVPR 2023poster

Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging…

2023

PUPS: Point Cloud Unified Panoptic Segmentation

AAAI 2023technical

Point cloud panoptic segmentation is a challenging task that seeks a holistic solution for both semantic and instance segmentation to predict groupings of coherent points. Previous approaches treat semantic and instance segmentation as surrogate tasks, and they either use clustering methods or bound…

Cited by 25SourcePDFScholar
2023

RWSC-Fusion: Region-Wise Style-Controlled Fusion Network for the Prohibited X-Ray Security Image Synthesis

CVPR 2023poster

Automatic prohibited item detection in security inspection X-ray images is necessary for transportation.The abundance and diversity of the X-ray security images with prohibited item, termed as prohibited X-ray security images, are essential for training the detection model. In order to solve the dat…

Cited by 4SourcePDFScholar
2023

Referring Expression Comprehension Using Language Adaptive Inference

AAAI 2023technical

Different from universal object detection, referring expression comprehension (REC) aims to locate specific objects referred to by natural language expressions. The expression provides high-level concepts of relevant visual and contextual patterns, which vary significantly with different expressions…

Cited by 20SourcePDFScholar
2023

SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation

IJCAI 2023poster

As an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D p…

2023

UniFusion: Unified Multi-View Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-View

ICCV 2023poster

Bird's eye view (BEV) representation is a new perception formulation for autonomous driving, which is based on spatial fusion. Further, temporal fusion is also introduced in BEV representation and gains great success. In this work, we propose a new method that unifies both spatial and temporal fusio…

Cited by 53PDFScholar
2022

Adaptive Cross-Domain Learning for Generalizable Person Re-identification

ECCV 2022poster

"Domain Generalizable Person Re-Identification (DG-ReID) is a more practical ReID task that is trained from multiple source domains and tested on the unseen target domains. Most existing methods are challenged for dealing with the shared and specific characteristics among different domains, which is…

2022

Detecting Backdoor Attacks against Point Cloud Classifiers

ICASSP 2022accepted

Backdoor attacks (BA) are an emerging threat to deep neural network classifiers. A classifier being attacked will predict to the attacker’s target class when a test sample from a source class is embedded with the backdoor pattern (BP). Recently, the first BA against point cloud (PC) classifiers was…

Cited by 0SourceScholar
2022

Dynamic Low-Resolution Distillation for Cost-Efficient End-to-End Text Spotting

ECCV 2022poster

"End-to-end text spotting has attached great attention recently due to its benefits on global optimization and high maintainability for real applications. However, the input scale has always been a tough trade-off since recognizing a small text instance usually requires enlarging the whole image, wh…

2022

Entropy-Driven Sampling and Training Scheme for Conditional Diffusion Generation

ECCV 2022poster

"Denoising Diffusion Probabilistic Model (DDPM) is able to make flexible conditional image generation from prior noise to real data, by introducing an independent noise-aware classifier to provide conditional gradient guidance at each time step of denoising process. However, due to the ability of th…

2022

Explicitly Modeling Importance and Coherence for Timeline Summarization

ICASSP 2022accepted

Timeline summarization (TLS) identifies major events and generates short summaries on how the event evolves in a period of time. Existing timeline summarization methods generate summaries by considering the coverage and diversity of the content and temporized information but ignore the importance an…

Cited by 0SourceScholar
2022

MetaGait: Learning to Learn an Omni Sample Adaptive Representation for Gait Recognition

ECCV 2022poster

"Gait recognition, which aims at identifying individuals by their walking patterns, has recently drawn increasing research attention. However, gait recognition still suffers from the conflicts between the limited binary visual clues of the silhouette and numerous covariates with diverse scales, whic…

Cited by 46SourcePDFScholar
2022

RBC: Rectifying the Biased Context in Continual Semantic Segmentation

ECCV 2022poster

"Recent years have witnessed a great development of Convolutional Neural Networks in semantic segmentation, where all classes of training images are simultaneously available. In practice, new images are usually made available in a consecutive manner, leading to a problem called Continual Semantic Se…

2022

SP-Net: Slowly Progressing Dynamic Inference Networks

ECCV 2022poster

"Dynamic inference networks improve computational efficiency by executing a subset of network components, i.e., executing path, conditioned on input sample. Prevalent methods typically assign routers to computational blocks so that a computational block can be skipped or executed. However, such infe…

2022

Saliency Hierarchy Modeling via Generative Kernels for Salient Object Detection

ECCV 2022poster

"Salient Object Detection (SOD) is a challenging problem that aims to precisely recognize and segment the salient objects. In ground-truth maps, all pixels belonging to the salient objects are positively annotated with the same value. However, the saliency level should be a relative quantity, which…

Cited by 17SourcePDFScholar
2022

Test-Time Detection of Backdoor Triggers for Poisoned Deep Neural Networks

ICASSP 2022accepted

Backdoor (Trojan) attacks are emerging threats against deep neural networks (DNN). A DNN being attacked will predict to an attacker-desired target class whenever a test sample from any source class is embedded with a backdoor pattern, while correctly classifying clean (attack-free) test samples. Exi…

Cited by 0SourceScholar
2021

A Backdoor Attack Against 3D Point Cloud Classifiers

ICCV 2021poster

Vulnerability of 3D point cloud (PC) classifiers has become a grave concern due to the popularity of 3D sensors in safety-critical applications. Existing adversarial attacks against 3D PC classifiers are all test-time evasion (TTE) attacks that aim to induce test-time misclassifications using knowle…

Cited by 94PDFcodeScholar
2021

Automatic Translation of Music-to-Dance for In-Game Characters

IJCAI 2021poster

Music-to-dance translation is an emerging and powerful feature in recent role-playing games. Previous works of this topic consider music-to-dance as a supervised motion generation problem based on time-series data. However, these methods require a large amount of training data pairs and may suffer f…

2021

Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion

CVPR 2021poster

RGB-D salient object detection (SOD) is usually formulated as a problem of classification or regression over two modalities, i.e., RGB and depth. Hence, effective RGB-D feature modeling and multi-modal feature fusion both play a vital role in RGB-D SOD. In this paper, we propose a depth-sensitive RG…

Cited by 214PDFcodeScholar
2020

BANet: Bidirectional Aggregation Network With Occlusion Handling for Panoptic Segmentation

CVPR 2020oral

Panoptic segmentation aims to perform instance segmentation for foreground instances and semantic segmentation for background stuff simultaneously. The typical top-down pipeline concentrates on two key issues: 1) how to effectively model the intrinsic interaction between semantic segmentation and in…

Cited by 92PDFcodeScholar
2020

Graph-Guided Architecture Search for Real-Time Semantic Segmentation

CVPR 2020poster

Designing a lightweight semantic segmentation network often requires researchers to find a trade-off between performance and speed, which is always empirical due to the limited interpretability of neural networks. In order to release researchers from these tedious mechanical trials, we propose a Gra…

Cited by 120PDFScholar
2020

Multi-Way Multi-View Deep Autoencoder for Image Feature Learning with Multi-Level Graph Regularization

ICASSP 2020accepted

Multi-view feature learning has garnered much attention recently since many real world data are comprised of different representations or views. How to explore the consensus structure and eliminate the inconsistency noise in different views remains a challenging problem in multi-view feature learnin…

Cited by 0SourceScholar
2020

Stacked Pooling for Boosting Scale Invariance of Crowd Counting

ICASSP 2020accepted

In this work, we take insight into the dense crowd counting problem by exploring the phenomenon of cross-scale visual similarity caused by perspective distortions. It is a quite common phenomenon in crowd scenarios, suggesting the crowd counting model to enable a good performance of scale invariance…

Cited by 0SourceScholar
2018

Geometry-Aware Scene Text Detection With Instance Transformation Network

CVPR 2018poster

Localizing text in the wild is challenging in the situations of complicated geometric layout of the targets like random orientation and large aspect ratio. In this paper, we propose a geometry-aware modeling approach tailored for scene text representation with an end-to-end learning scheme. In our a…

Cited by 111SourcePDFScholar
2018

Weakly-Supervised Semantic Segmentation by Iteratively Mining Common Object Features

CVPR 2018poster

Weakly-supervised semantic segmentation under image tags supervision is a challenging task as it directly associates high-level semantic to low-level appearance. To bridge this gap, in this paper, we propose an iterative bottom-up and top-down framework which alternatively expands object regions and…

Cited by 375SourcePDFScholar
2017

Deeply-Learned Part-Aligned Representations for Person Re-Identification

ICCV 2017poster

In this paper, we address the problem of person re-identification, which refers to associating the persons captured from different cameras. We propose a simple yet effective human part-aligned representation for handling the body part misalignment problem. Our approach decomposes the human body into…

Cited by 951PDFScholar
2015

3D Hand Pose Estimation Using Randomized Decision Forest With Segmentation Index Points

ICCV 2015poster

In this paper, we propose a real-time 3D hand pose estimation algorithm using the randomized decision forest framework. Our algorithm takes a depth image as input and generates a set of skeletal joints as output. Previous decision forest-based methods often give labels to all points in a point cloud…

Cited by 70PDFScholar