← Search

Xudong Liu

20 accepted papers

2026

AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

CVPR 2026

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning, tracking who speaks, maintaining roles, and grounding events across time. These scenarios are central to multim

Cited by 0SourceScholar
2026

Multi-level Causal LLM-based Text-to-Motion Generation with Human Alignment

CVPR 2026

Although progress has been made in LLM-based text-driven motion generation, it still has the limitations of generating fine-grained and semantically consistent motions. These limitations stem from: 1) fine-grained motion quantization errors; 2) mismatches between causal reasoning language and non-ca

Cited by 0SourceScholar
2026

WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation

AAAI 2026technical

Vision modeling has advanced rapidly with Transformers, whose attention mechanisms capture visual dependencies but lack a principled account of how semantic information propagates spatially. We revisit this problem from a wave-based perspective: feature maps are treated as spatial signals whose evol

Cited by 0SourcePDFScholar
2025

DFM: Differentiable Feature Matching for Anomaly Detection

CVPR 2025poster

Feature matching methods for unsupervised anomaly detection have demonstrated impressive performance. Existing methods primarily rely on self-supervised training and handcrafted matching schemes for task adaptation. However, they can only achieve an inferior feature representation for anomaly detect…

Cited by 0SourcePDFScholar
2025

Glance2Gaze: Efficient Vision-Language Models from Glance Fusion to Gaze Compression

NeurIPS 2025poster

Vision-language models heavily rely on visual representations, yet ensuring its efficiency remains a critical challenge. Most existing approaches focus on reducing visual tokens either at the visual encoder phase or during the LLM decoder stage. Inspired by human visual cognition, where an initial g…

Cited by 0SourceScholar
2025

IOP: An Idempotent-Like Optimization Method on the Pareto Front of Hypernetwork

AAAI 2025technical

Pareto Front Learning (PFL) has been one of the effective means to resolve multi-objective optimization problems through exploring all optimal solutions to learn the entire Pareto front. Pareto Hypernetwork (PHN) is a new promising way to generate the sequence of Pareto-optimal solutions that can be…

Cited by 0SourcePDFScholar
2024

Code-Style In-Context Learning for Knowledge-Based Question Answering

AAAI 2024technical

Current methods for Knowledge-Based Question Answering (KBQA) usually rely on complex training techniques and model frameworks, leading to many limitations in practical applications. Recently, the emergence of In-Context Learning (ICL) capabilities in Large Language Models (LLMs) provides a simple a…

2024

LEOD: Label-Efficient Object Detection for Event Cameras

CVPR 2024poster

Object detection with event cameras benefits from the sensor's low latency and high dynamic range. However it is costly to fully label event streams for supervised training due to their high temporal resolution. To reduce this cost we present LEOD the first method for label-efficient event-based det…

2024

SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark Estimation

CVPR 2024poster

Self-supervised landmark estimation is a challenging task that demands the formation of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To tackle this task existing state-of-the-art (SOTA) methods (1) extract coarse features from backbon…

Cited by 1SourcePDFScholar
2023

Cross-view Semantic Alignment for Livestreaming Product Recognition

ICCV 2023poster

Live commerce is the act of selling products online through livestreaming. The customer's diverse demands for online products introduces more challenges to Livestreaming Product Recognition. Previous works are either focus on fashion clothing data or subject to single-modal input, thus inconsistent…

Cited by 4PDFcodeScholar
2023

Deep Autoencoding One-Class time Series Anomaly Detection

ICASSP 2023accepted

Time-series Anomaly Detection(AD) is widely used in monitoring and security applications in various industries and has become a hot spot in the field of deep learning. Normality-representation-based methods perform well in certain scenarios but may ignore some aspects of the overall normality. Featu…

Cited by 0SourceScholar
2023

High-Fidelity Clothed Avatar Reconstruction From a Single Image

CVPR 2023poster

This paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR)…

2022

A Transformational Biencoder with In-Domain Negative Sampling for Zero-Shot Entity Linking

ACL 2022findings

Recent interest in entity linking has focused in the zero-shot scenario, where at test time the entity mention to be labelled is never seen during training, or may belong to a different domain from the source domain. Current work leverage pre-trained BERT with the implicit assumption that it bridges…

2022

Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives

AAAI 2022technical

Unsupervised sentence representation learning is a fundamental problem in natural language processing. Recently, contrastive learning has made great success on this task. Existing constrastive learning based models usually apply random sampling to select negative examples for training. Previous work…

Cited by 65SourcePDFScholar
2021

Arrhythmia Classification with Heartbeat-Aware Transformer

ICASSP 2021accepted

Electrocardiography (ECG) is a conventional method in arrhythmia diagnosis. In this paper, we proposed a novel neural network model which treats typical heartbeat classification task as ‘Translation’ problem. By introducing Transformer structure into model, and adding heartbeat-aware attention mecha…

Cited by 0SourceScholar
2021

Progressive Multi-task Learning with Controlled Information Flow for Joint Entity and Relation Extraction

AAAI 2021technical

Multitask learning has shown promising performance in learning multiple related tasks simultaneously, and variants of model architectures have been proposed, especially for supervised classification problems. One goal of multitask learning is to extract a good representation that sufficiently captur…

2020

IMRAM: Iterative Matching With Recurrent Attention Memory for Cross-Modal Image-Text Retrieval

CVPR 2020poster

Enabling bi-directional retrieval of images and texts is important for understanding the correspondence between vision and language. Existing methods leverage the attention mechanism to explore such correspondence in a fine-grained manner. However, most of them consider all semantics equally and thu…

Cited by 461PDFcodeScholar
2020

Structured Probabilistic End-to-End Learning from Crowds

IJCAI 2020poster

End-to-end learning from crowds has recently been introduced as an EM-free approach to training deep neural networks directly from noisy crowdsourced annotations. It models the relationship between true labels and annotations with a specific type of neural layer, termed as the crowd layer, which can…

Cited by 0SourcePDFScholar