← Search

Xiaofan Li

18 accepted papers

2026

Artemis: Structured Visual Reasoning for Perception Policy Learning

ICML 2026poster

Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indicate that such purely linguistic intermediate reasoning often reduces performance on perception tasks. We argue that the c…

Cited by 0SourceScholar
2026

Benchmarking Dense and Indiscernible Object Counting with Blueberries

ICML 2026poster

Real-world agricultural counting often operates in the extreme regime of \textbf{Dense and Indiscernible Object Counting (DIOC)}, where targets are tiny, clustered, and highly camouflaged. To facilitate research in this domain, we introduce \textbf{DIOCblueberry}, a large-scale benchmark that pushes…

Cited by 0SourceScholar
2026

Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception

CVPR 2026

Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods for generating such captions often rely on distilling the captions from pretrained LMMs, constructing them from publicly available internet images, or even generating the

Cited by 0SourcecodeScholar
2026

FM-Steer: Enhance Generalist Policies with Value-Guided Cascaded Denoising

CVPR 2026

Humans naturally allocate more time before acting when handling complex tasks in the physical world. This paradigm has recently led to remarkable advances in boosting Large Language Models (LLMs) on complex tasks in digital domains. However, the potential of test-time computing remains largely unexp

Cited by 0SourcecodeScholar
2026

FVAR: Next-Focus Prediction for Visual Autoregressive Modeling

CVPR 2026

Visual autoregressive models achieve remarkable generation quality through next-scale predictions across multi-scale token pyramids. However, the conventional method uses uniform scale downsampling to build these pyramids, leading to aliasing artifacts that compromise fine details and introduce unwa

Cited by 0SourceScholar
2026

FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain

CVPR 2026

In controllable driving-scene reconstruction and 3D scene generation, maintaining geometric fidelity while synthesizing visually plausible appearance under large viewpoint shifts is crucial. However, effective fusion of geometry-based 3DGS and appearance-driven diffusion models faces inherent challe

Cited by 0SourceScholar
2026

From Prompts to Printable Models: Support-Effective 3D Generation via Offset Direct Preference Optimization

RA-L 2026

Current text-to-3D models prioritize visual fidelity but often neglect physical fabricability, resulting in geometries requiring excessive support structures. This paper introduces SEG (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><u>S</u>upport-<

Cited by 0SourceScholar
2026

UniFuture: A 4D Driving World Model for Future Generation and Perception

ICRA 2026poster

We present UniFuture, a unified 4D Driving World Model designed to simulate the dynamic evolution of the 3D physical world. Unlike existing driving world models that focus solely on 2D pixel-level video generation (lacking geometry) or static perception (lacking temporal dynamics), our approach brid…

2026

ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

CVPR 2026

MLLMs exhibit strong reasoning on isolated queries, yet they operate de novo--solving each problem independently and often repeating the same mistakes. Existing memory-augmented agents mainly store past trajectories for reuse. However, trajectory-based memory suffers from brevity bias, gradually los

Cited by 0SourcecodeScholar
2026

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

CVPR 2026

Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA, a training-free identify-then-guide framework for improved numerical alignment. NUMINA identifies prompt-layout inconsi

Cited by 0SourcecodeScholar
2025

CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential Perception

ICCV 2025poster

Cooperative perception aims to address the inherent limitations of single-vehicle autonomous driving systems through information exchange among multiple agents. Previous research has primarily focused on single-frame perception tasks. However, the more challenging cooperative sequential perception t…

2025

Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving

ICRA 2025

The autoregressive world model exhibits robust generalization capabilities in vectorized scene understanding but encounters difficulties in deriving actions due to insufficient uncertainty modeling and self-delusion. In this paper, we explore the feasibility of deriving decisions from an autoregres-

Cited by 7SourcecodeScholar
2025

One-for-More: Continual Diffusion Model for Anomaly Detection

CVPR 2025poster

With the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also fall into this scope and utilize diffusion models to generate or reconstruct normal samples when given arbitrary anomaly images. However, our study foun…

2025

U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration

ICCV 2025poster

Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and construction sites significantly degrade GNSS (Global Navigation Satellite System) signal quality. This issue underscores the importance of visual localizati…

Cited by 0SourcePDFScholar
2024

One-Stage Training Generative Paradigm for Generalized Zero-Shot Learning

ICASSP 2024accepted

Zero-shot learning image classification aims to identify unseen classes not present during training. Generalized zero-shot learning (GZSL) is more in line with realistic scenarios due to its ability of recognizing both seen and unseen classes. Current GZSL methods mostly utilize generative adversari…

Cited by 0SourceScholar
2024

PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection

CVPR 2024poster

The vision-language model has brought great improvement to few-shot industrial anomaly detection which usually needs to design of hundreds of prompts through prompt engineering. For automated scenarios we first use conventional prompt learning with many-class paradigm as the baseline to automaticall…

2023

VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot Learning

IJCAI 2023poster

Unlike conventional zero-shot learning (CZSL) which only focuses on the recognition of unseen classes by using the classifier trained on seen classes and semantic embeddings, generalized zero-shot learning (GZSL) aims at recognizing both the seen and unseen classes, so it is more challenging due to…

Cited by 17SourcePDFScholar
2022

En-Compactness: Self-Distillation Embedding & Contrastive Generation for Generalized Zero-Shot Learning

CVPR 2022poster

Generalized zero-shot learning (GZSL) requires a classifier trained on seen classes that can recognize objects from both seen and unseen classes. Due to the absence of unseen training samples, the classifier tends to bias towards seen classes. To mitigate this problem, feature generation based model…

Cited by 89PDFScholar