← Search

Fan Li

48 accepted papers

2026

CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions

CVPR 2026

Instruction-based multimodal image manipulation has recently made rapid progress. However, existing evaluation methods lack a systematic and human-aligned framework for assessing model performance on complex and creative editing tasks. To address this gap, we propose CREval, a fully automated questi

Cited by 0SourcecodeScholar
2026

ColorFLUX: A Structure-Color Decoupling Framework for Old Photo Colorization

CVPR 2026

Old photos preserve invaluable historical memories, making their restoration and colorization highly desirable. While existing restoration models can address some degradation issues like denoising and scratch removal, they often struggle with accurate colorization.This limitation arises from the uni

Cited by 0SourcecodeScholar
2026

Concept Bottleneck Models for Explainable Decision Making: A Survey of Progress, Taxonomy, and Future Directions

IJCAI 2026

Deep neural networks deliver strong performance but remain opaque, limiting their use in high-stakes domains that require transparency and human oversight. Concept Bottleneck Models (CBMs) address this gap by introducing a human-interpretable concept layer that mediates inputs and decisions, enablin

Cited by 0Scholar
2026

DHG-Bench: A Comprehensive Benchmark for Deep Hypergraph Learning

ICLR 2026poster

Deep graph models have achieved great success in network representation learning. However, their focus on pairwise relationships restricts their ability to learn pervasive higher-order interactions in real-world systems, which can be naturally modeled as hypergraphs. To tackle this issue, Hypergraph…

Cited by 0SourcecodeScholar
2026

FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUs

AAAI 2026technical

Recent years have witnessed the wide adoption of deep learning recommendation models (DLRMs) for many online services. Unlike traditional DNN training, DLRMs leverage massive embeddings to represent sparse features, which are stored in distributed GPUs following the model parallel paradigm. Existing

Cited by 0SourcePDFScholar
2026

Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning

ICML 2026poster

Vector Quantization (VQ) has recently emerged as a promising approach for learning discrete representations of graph-structured data. However, a fundamental challenge, i.e., codebook collapse, remains underexplored in the graph domain, significantly limiting the expressiveness and generalization of …

Cited by 0SourceScholar
2026

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

CVPR 2026

Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (RL) methods such as Diffusion-DPO and Flow-GRPO have further improved generation quality, efficiently applying Reinforce

Cited by 0SourceScholar
2026

HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection

AAAI 2026technical

Video Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal Large Language Models (MLLMs) offer a promising alternative

Cited by 0SourcePDFScholar
2026

Influence without Confounding: Causal Discovery from Temporal Data with Long-term Carry-over Effects

ICLR 2026poster

Learning causal structures from temporal data is fundamental to many practical tasks, such as physical laws discovery and root causes localization. Real-world systems often exhibit long-term carry-over effects, where the value of a variable at the current time can be influenced by distant past va…

Cited by 0SourceScholar
2026

Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous Driving

AAAI 2026technical

Modern autonomous driving (AD) systems leverage 3D object detection to perceive foreground objects in 3D environments for subsequent prediction and planning. Visual 3D detection based on RGB cameras provides a cost-effective solution compared to the LiDAR paradigm. While achieving promising detectio

Cited by 0SourcePDFScholar
2026

RefSTAR: Blind Face Image Restoration with Reference Selection, Transfer, and Reconstruction

AAAI 2026technical

Introducing high-quality references can largely alleviate the uncertainty in blind face image restoration tasks, yet the equivocal utilization of reference priors makes it still a struggle to well preserve the human identity. We attribute the identity inconsistency to two deficiencies of existing re

Cited by 0SourcePDFScholar
2026

Refine3D: Scene-Adaptive Reference Point Refinement for Sparse 3D Object Detection

AAAI 2026technical

Sparse query-based detectors have emerged as the dominant paradigm in camera-only 3D object detection, owing to their exceptional performance and computational efficiency. A central component of these approaches is the use of reference points, which serve as learnable spatial anchors to guide queri

Cited by 0SourcePDFScholar
2026

Steering and Rectifying Latent representation manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection

ICLR 2026poster

Video anomaly detection (VAD) aims to identify abnormal events in videos. Traditional VAD methods generally suffer from the high costs of labeled data and full training, thus some recent works have explored leveraging frozen multi-modal large language models (MLLMs) in a tuning-free manner to perfor…

Cited by 0SourceScholar
2026

Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection

ICML 2026poster

Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlike in the RGB domain, where attacks rely on color texture, infrared attacks must manipulate thermal signatures, making the geometry shape of heat-blo…

Cited by 0SourceScholar
2026

What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs

CVPR 2026

Split DNNs enable edge devices by offloading intensive computation to a cloud server, but this paradigm exposes privacy vulnerabilities, as the intermediate features can be exploited to reconstruct the private inputs via Feature Inversion Attack (FIA). Existing FIA methods often produce limited reco

Cited by 0SourceScholar
2026

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal

CVPR 2026

Recent advances in Diffusion Transformer (DiT)-based video generation technologies have shown impressive results for video object removal. However, these methods still suffer from substantial inference latency. For instance, although MiniMax Remover achieves state-of-the-art visual quality, it opera

Cited by 0SourcecodeScholar
2025

A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision

ICASSP 2025accepted

Extracting singing melody from polyphonic music is an important topic in the field of music information retrieval. In this paper, we propose a singing melody extraction network consisting of five stacked multi-scale feature time-frequency aggregation (MF-TFA) modules. In the same network, deeper lay…

Cited by 0SourceScholar
2025

A Task-Oriented Real-Time and Robust Feature Compression and Selection Method in Collaborative Intelligence System

ICASSP 2025accepted

The emerging autonomous driving has stringent requirements for latency and reliability. In this paper, we propose a task-oriented real-time and robust feature compression and selection method in collaborative intelligence system. Our design, consisting of a three-dimensional channel compression (TDC…

Cited by 0SourceScholar
2025

ACE: Anti-Editing Concept Erasure in Text-to-Image Models

CVPR 2025poster

Recent advance in text-to-image diffusion models have significantly facilitated the generation of high-quality images, but also raising concerns about the illegal creation of harmful content, such as copyrighted images. Existing concept erasure methods achieve superior results in preventing the prod…

2025

Adaptive Gradient Masking for Balancing ID and MLLM-based Representations in Recommendation

NeurIPS 2025poster

In large-scale recommendation systems, multimodal (MM) content is increasingly introduced to enhance the generalization of ID features. The rise of Multimodal Large Language Models (MLLMs) enables the construction of unified user and item representations. However, the semantic distribution gap betwe…

Cited by 0SourceScholar
2025

Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense Knowledge

ACL 2025long

This paper focuses on a new task of answering geographic reasoning questions based on the given image (called GeoVQA). Unlike traditional VQA tasks, GeoVQA asks for details about the image-related culture, landscape, etc. This requires not only the identification of the objects in the image, their p…

Cited by 0SourcePDFScholar
2025

Better to Teach than to Give: Domain Generalized Semantic Segmentation via Agent Queries with Diffusion Model Guidance

ICML 2025spotlight

Domain Generalized Semantic Segmentation (DGSS) trains a model on a labeled source domain to generalize to unseen target domains with consistent contextual distribution and varying visual appearance. Most existing methods rely on domain randomization or data generation but struggle to capture the un…

2025

Calibrating Video Watch-time Predictions with Credible Prototype Alignment

ICML 2025poster

Accurately predicting user watch-time is crucial for enhancing user stickiness and retention in video recommendation systems. Existing watch-time prediction approaches typically involve transformations of watch-time labels for prediction and subsequent reversal, ignoring both the natural distributio…

Cited by 0SourcePDFScholar
2025

CamEdit: Continuous Camera Parameter Control for Photorealistic Image Editing

NeurIPS 2025poster

Recent advances in diffusion models have substantially improved text-driven image editing. However, existing frameworks based on discrete textual tokens struggle to support continuous control over camera parameters and smooth transitions in visual effects. These limitations hinder their applications…

Cited by 0SourceScholar
2025

Consistent Feature Alignment for Cross-Modal Knowledge Distillation in Monocular 3D Object Detection

IROS 2025

Cross-modal knowledge distillation (CMKD) in monocular 3D object detection transfers LiDAR’s accurate depth information to compensate for the limitations of camera model. However, current methods directly align the intermediate features of the teacher and student networks, in which the modality gap

Cited by 0SourceScholar
2025

Decoupled Feature Matching for Few-shot Counting and Localization

ICASSP 2025accepted

Few-shot counting (FSC) aims to train a generalized visual counting model that can count any novel category given a small number of support samples. Current prevalent approaches treat FSC as a feature-matching task, leveraging attention to aggregate information from all other query patches or suppor…

Cited by 0SourceScholar
2025

Dual Prompting Image Restoration with Diffusion Transformers

CVPR 2025poster

Recent state-of-the-art image restoration methods mostly adopt latent diffusion models with U-Net backbones, yet still facing challenges in achieving high-quality restoration due to their limited capabilities. Diffusion transformers (DiTs), like SD3, are emerging as a promising alternative because o…

Cited by 1SourcePDFScholar
2025

Fast Image Super-Resolution via Consistency Rectified Flow

ICCV 2025poster

Diffusion models (DMs) have demonstrated remarkable success in real-world image super-resolution (SR), yet their reliance on time-consuming multi-step sampling largely hinders their practical applications. While recent efforts have introduced few- or single-step solutions, existing methods either in…

Cited by 0SourcePDFScholar
2025

Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Recently, open-vocabulary semantic segmentation has garnered growing attention. Most current methods leverage vision-language models like CLIP to recognize unseen categories through their zero-shot capabilities. However, CLIP struggles to establish potential spatial dependencies among scene objects…

Cited by 0SourcePDFScholar
2025

MC^2: Multi-concept Guidance for Customized Multi-concept Generation

CVPR 2025poster

Customized text-to-image generation, which synthesizes images based on user-specified concepts, has made significant progress in handling individual concepts. However, when extended to multiple concepts, existing methods often struggle with properly integrating different models and avoiding the unin…

2025

No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion Models

NeurIPS 2025poster

Enhancing the cross-domain generalization of 3D semantic segmentation is a pivotal task in computer vision that has recently gained increasing attention. Most existing methods, whether using consistency regularization or cross-modal feature fusion, focus solely on individual objects while overlookin…

Cited by 0SourcecodeScholar
2025

PCAN: A Pandemic-Compatible Attentive Neural Network for Retail Sales Forecasting

IJCAI 2025

The outbreak of pandemic has a huge impact on production and consumption in the business world, especially for the retail sector. As a crucial component of decision-support technology in the retail industry, sales forecasting is significant for production planning and optimizing the supply of essent

2025

PocketSR: The Super-Resolution Expert in Your Pocket Mobiles

NeurIPS 2025poster

Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for ed…

Cited by 0SourceScholar
2025

Priority Guided Explanation for Knowledge Tracing with Dual Ranking and Similarity Consistency

IJCAI 2025

Knowledge tracing plays a pivotal role in enabling personalized learning on online platforms. While deep learning-based approaches have achieved impressive predictive performance, their limited interpretability poses a significant barrier to practical adoption. Existing explanation methods primarily

Cited by 0SourcePDFScholar
2024

DyFADet: Dynamic Feature Aggregation for Temporal Action Detection

ECCV 2024poster

"Recent proposed neural network-based Temporal Action Detection (TAD) models are inherently limited to extracting the discriminative representations and modeling action instances with various lengths from complex scenes by shared-weights detection heads. Inspired by the successes in dynamic neural n…

2024

Hypergraph Self-supervised Learning with Sampling-efficient Signals

IJCAI 2024poster

Self-supervised learning (SSL) provides a promising alternative for representation learning on hypergraphs without costly labels. However, existing hypergraph SSL models are mostly based on contrastive methods with the instance-level discrimination strategy, suffering from two significant limitation…

2024

LLMs Can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought

IJCAI 2024poster

Self-correction is emerging as a promising approach to mitigate the issue of hallucination in Large Language Models (LLMs). To facilitate effective self-correction, recent research has proposed mistake detection as its initial step. However, current literature suggests that LLMs often struggle with…

2024

MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors

CVPR 2024poster

Foot contact is an important cue for human motion capture understanding and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. However these approaches either suffer from low accuracy or are only designed for s…

2023

Fighting against Organized Fraudsters Using Risk Diffusion-based Parallel Graph Neural Network

IJCAI 2023poster

Medical insurance plays a vital role in modern society, yet organized healthcare fraud causes billions of dollars in annual losses, severely harming the sustainability of the social welfare system. Existing works mostly focus on detecting individual fraud entities or claims, ignoring hidden conspira…

Cited by 14SourcePDFScholar
2022

Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization

NeurIPS 2022accept

Successful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning . While featuring superior stability, these estimators crucially depend on costly large-batch trai…

2021

Counterfactual Representation Learning with Balancing Weights

AISTATS 2021poster

A key to causal inference with observational data is achieving balance in predictive features associated with each treatment type. Recent literature has explored representation learning to achieve this goal. In this work, we discuss the pitfalls of these strategies – such as a steep trade-off betwee…

Cited by 89SourcePDFScholar
2020

Learning Consistency Pursued Correlation Filters for Real-Time UAV Tracking

IROS 2020poster

Correlation filter (CF)-based methods have demonstrated exceptional performance in visual object tracking for unmanned aerial vehicle (UAV) applications, but suffer from the undesirable boundary effect. To solve this issue, spatially regularized correlation filters (SRDCF) proposes the spatial regul…

Cited by 11SourceScholar
2020

Reconsidering Generative Objectives For Counterfactual Reasoning

NeurIPS 2020poster

There has been recent interest in exploring generative goals for counterfactual reasoning, such as individualized treatment effect (ITE) estimation. However, existing solutions often fail to address issues that are unique to causal inference, such as covariate balancing and (infeasible) counterfactu…

2020

Training-Set Distillation for Real-Time UAV Object Tracking

ICRA 2020poster

Correlation filter (CF) has recently exhibited promising performance in visual object tracking for unmanned aerial vehicle (UAV). Such online learning method heavily depends on the quality of the training-set, yet complicated aerial scenarios like occlusion or out of view can reduce its reliability.…

Cited by 34SourcecodeScholar