← Search

Di Xu

18 accepted papers

2026

AutoMetrics: Approximate Human Judgments with Automatically Generated Evaluators

ICLR 2026poster

Evaluating user-facing AI applications remains a central challenge, especially in open-ended domains such as travel planning, clinical note generation, or dialogue. The gold standard is user feedback (e.g., thumbs up/down) or behavioral signals (e.g., retention), but these are often scarce in protot…

Cited by 0SourcecodeScholar
2026

Do Large Language Models Reason About Uncertainty Like Humans? A Benchmark on Hurricane Forecast Visualization Comprehension

AAAI 2026technical

Uncertainty visualizations, such as hurricane cones and ensemble tracks, are essential for risk communication but are often misinterpreted, leading to harmful decisions. As AI assistants like large language models (LLMs) increasingly support understanding of graphics and decision-making, they offer

Cited by 0SourcePDFScholar
2026

DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance

AAAI 2026technical

Infrared imaging plays a critical role in low-light and adverse weather conditions. However, due to the distinct characteristics of infrared images, existing foundation models such as Masked Autoencoder (MAE) trained on visible data perform suboptimal in infrared image interpretation tasks. To bridg

Cited by 0SourcePDFScholar
2026

Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning

AAAI 2026technical

Continual learning (CL) aims to equip models with the ability to learn from a stream of tasks without forgetting previous knowledge. With the progress of vision-language models like Contrastive Language-Image Pre-training (CLIP), their promise for CL has attracted increasing attention due to their s

Cited by 0SourcePDFScholar
2026

Interference-Isolated Elastic Weight Consolidation and Knowledge Calibration for Incremental Object Detection

ICLR 2026poster

Incremental Object Detection (IOD) enables AI systems to continuously learn new object classes over time while retaining knowledge of previously learned categories. This capability is essential for adapting to dynamic environments without forgetting prior information. Although existing IOD methods h…

Cited by 0SourceScholar
2026

Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition Model

ICML 2026poster

Open-world object detection (OWOD) requires incrementally detecting known categories while reliably identifying unknown objects. Existing methods primarily focus on improving unknown recall, yet overlook interpretability, often leading to known–unknown confusion and reduced prediction reliability. T…

Cited by 0SourceScholar
2026

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

ICLR 2026poster

Composed video retrieval presents a complex challenge: retrieving a target video based on a source video and a textual modification instruction. This task demands fine-grained reasoning over multimodal transformations. However, existing benchmarks predominantly focus on vision–text alignment, largel…

Cited by 0SourceScholar
2026

Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection

ICML 2026poster

Accurately localizing and segmenting small targets in low signal-to-noise ratio (SNR) infrared sequences remains a challenging task. Since targets are often indistinguishable from the background in individual frames, existing methods, even when equipped with advanced foundation model and powerful in…

Cited by 0SourceScholar
2026

YOLO-IOD: Towards Real Time Incremental Object Detection

AAAI 2026technical

Current methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to c

Cited by 0SourcePDFScholar
2025

Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector

ICML 2025poster

Catastrophic forgetting is a critical chanllenge for incremental object detection (IOD). Most existing methods treat the detector monolithically, relying on instance replay or knowledge distillation without analyzing component-specific forgetting. Through dissection of Faster R-CNN, we reveal a key…

Cited by 0SourcePDFScholar
2025

Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening

CVPR 2025poster

Pansharpening aims at integrating complementary information from panchromatic and multispectral images. Available deep-learning based pansharpening methods typically perform exceptionally with particular satellite datasets. At the same time, it has been observed that these models also exhibit scene…

2025

Dual-Process Watermarked Diffusion: Integrating Watermarking With Denoising in Point Clouds

ICASSP 2025accepted

The integration of depth sensing and laser scanning technologies has propelled point cloud data to the forefront of 3D graphical modeling. This paper addresses a critical gap in the literature: the protection of intellectual property in generating point clouds using Diffusion Models (DMs). We introd…

Cited by 0SourceScholar
2025

Implanting Robust Watermarks in Latent Diffusion Models for Video Generation

ICASSP 2025accepted

In the dynamic realm of digital media, latent diffusion models (LDM) have revolutionized the generation of videos, surpassing the capabilities of traditional generative models. This paper presents Stable Video Signature, a pioneering watermarking framework for LDM in video generation. Addressing the…

Cited by 0SourceScholar
2025

Revisiting Generative Replay for Class Incremental Object Detection

CVPR 2025poster

Generative replay has gained significant attention in class-incremental learning; however, its application to Class Incremental Object Detection (CIOD) remains limited due to the challenges in generating complex images with precise spatial arrangements. In this study, motivated by the observation th…

2024

3D-Aware Face Editing via Warping-Guided Latent Direction Learning

CVPR 2024poster

3D facial editing a longstanding task in computer vision with broad applications is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness generalization and efficiency. To overcome these cha…

2024

Topo4D: Topology-Preserving Gaussian Splatting for High-Fidelity 4D Head Capture

ECCV 2024poster

"Recent significant advances in high-quality face reconstruction have been made, but challenges remain in 4D face asset reconstruction. 4D head capture aims to generate dynamic topological meshes and corresponding texture maps from videos, which is widely utilized in movies and games for its ability…

2023

PARF: Primitive-Aware Radiance Fusion for Indoor Scene Novel View Synthesis

ICCV 2023poster

This paper proposes a method for fast scene radiance field reconstruction with strong novel view synthesis performance and convenient scene editing functionality. The key idea is to fully utilize semantic parsing and primitive extraction for constraining and accelerating the radiance field reconstru…

Cited by 7PDFScholar
2015

Semi-supervised training in low-resource ASR and KWS

ICASSP 2015accepted

In particular for “low resource” Keyword Search (KWS) and Speech-to-Text (STT) tasks, more untranscribed test data may be available than training data. Several approaches have been proposed to make this data useful during system development, even when initial systems have Word Error Rates (WER) abov…

Cited by 0SourceScholar