← Search

Changsheng Xu

66 accepted papers

2026

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

ICML 2026poster

Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordinates, which violates the principle of general covariance. Theoretically, this conflates the intrinsic task geometry with rigid execution patterns, …

Cited by 0SourceScholar
2026

I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-Disentangle

AAAI 2026technical

Compositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives,

Cited by 0SourcePDFScholar
2026

LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment

ICML 2026poster

We formulate the learning of generalist Vision-Language-Action (VLA) models as a Gromov-Wasserstein alignment problem, aiming to map semantically similar VL embeddings to physically similar motion primitives. However, solving this is challenging due to the mathematical heterogeneity between the doma…

Cited by 0SourceScholar
2026

QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response

ICLR 2026poster

The increasing demand for real-time interaction in online video scenarios necessitates a new class of efficient streaming video understanding models. However, existing approaches often rely on a flawed, query-agnostic ``change-is-important'' principle, which conflates visual dynamics with semantic r…

Cited by 0SourceScholar
2026

SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

ICML 2026poster

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We pro…

Cited by 0SourceScholar
2026

SoMe: A Realistic Benchmark for LLM-based Social Media Agents

AAAI 2026technical

Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM agents are reshaping the ecology of social media, there exists a current gap in conducting a comprehensive evaluation of

Cited by 0SourcePDFScholar
2025

Language Guided Concept Bottleneck Models for Interpretable Continual Learning

CVPR 2025poster

Continual learning (CL) aims to enable learning systems to acquire new knowledge constantly without forgetting previously learned information. CL faces the challenge of mitigating catastrophic forgetting while maintaining interpretability across tasks.Most existing CL methods focus primarily on pres…

2025

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

NeurIPS 2025poster

Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and determine optimal response timing, often compromising real-time responsiveness and na…

Cited by 0SourcecodeScholar
2025

Locality Preserving Markovian Transition for Instance Retrieval

ICML 2025poster

Diffusion-based re-ranking methods are effective in modeling the data manifolds through similarity propagation in affinity graphs. However, positive signals tend to diminish over several steps away from the source, reducing discriminative power beyond local regions. To address this issue, we introdu…

Cited by 0SourcePDFScholar
2025

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

NeurIPS 2025poster

In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-…

Cited by 0SourcecodeScholar
2025

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments

ICCV 2025poster

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often struggle with generalizing to novel environments and adapting to ongoing changes durin…

2025

Overcoming Dual Drift for Continual Long-Tailed Visual Question Answering

ICCV 2025poster

Visual Question Answering (VQA) is a widely explored multimodal task aimed at answering questions based on images. Recently, a few studies have started to investigate continual learning in VQA to cope with evolving multimodal data streams. However, these studies fall short of tackling another critic…

Cited by 0SourcePDFScholar
2025

Pilot: Building the Federated Multimodal Instruction Tuning Framework

AAAI 2025technical

In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices. To solve the new task, we propose a federated multimodal instruction tuning fra…

Cited by 1SourcePDFScholar
2025

Pseudo Informative Episode Construction for Few-Shot Class-Incremental Learning

AAAI 2025technical

Few-Shot Class-Incremental Learning (FSCIL) studies how to empower the machine learning system to learn novel classes with only a few annotated examples continually. To tackle the FSCIL task, recent state-of-the-art methods propose to employ the meta-learning mechanism, which constructs the pseudo i…

Cited by 0SourcePDFScholar
2025

Rethinking the Temperature for Federated Heterogeneous Distillation

ICML 2025poster

Federated Distillation (FedKD) relies on lightweight knowledge carriers like logits for efficient client-server communication. Although logit-based methods have demonstrated promise in addressing statistical and architectural heterogeneity in federated learning (FL), current approaches remain const…

Cited by 0SourcePDFScholar
2025

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

ICLR 2025spotlight

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding ty…

2025

When Open-Vocabulary Visual Question Answering Meets Causal Adapter: Benchmark and Approach

AAAI 2025technical

Visual Question Answering (VQA) is a multifaceted task that integrates computer vision and natural language processing to produce textual answers from images and questions. Existing VQA benchmarks predominantly adhere to a closed-set paradigm, limiting their ability to address arbitrary, unseen answ…

Cited by 0SourcePDFScholar
2024

Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language Models

NeurIPS 2024poster

A straightforward pipeline for zero-shot out-of-distribution (OOD) detection involves selecting potential OOD labels from an extensive semantic pool and then leveraging a pre-trained vision-language model to perform classification on both in-distribution (ID) and OOD labels. In this paper, we theori…

2024

Enhancing Storage and Computational Efficiency in Federated Multimodal Learning for Large-Scale Models

ICML 2024poster

The remarkable generalization of large-scale models has recently gained significant attention in multimodal research. However, deploying heterogeneous large-scale models with different modalities under Federated Learning (FL) to protect data privacy imposes tremendous challenges on clients' limited…

2024

Fast-Slow Test-Time Adaptation for Online Vision-and-Language Navigation

ICML 2024poster

The ability to accurately comprehend natural language instructions and navigate to the target location is essential for an embodied agent. Such agents are typically required to execute user instructions in an online manner, leading us to explore the use of unlabeled test samples for effective online…

2024

Libra: Building Decoupled Vision System on Large Language Models

ICML 2024poster

In this work, we introduce **Libra**, a prototype model with a decoupled vision system on a large language model (LLM). The decoupled vision system decouples inner-modal modeling and cross-modal interaction, yielding unique visual information modeling and effective cross-modal comprehension. Libra i…

2024

Modality-Collaborative Test-Time Adaptation for Action Recognition

CVPR 2024poster

Video-based Unsupervised Domain Adaptation (VUDA) method improves the generalization of the video model enabling it to be applied to action recognition tasks in different environments. However these methods require continuous access to source data during the adaptation process which are impractical…

Cited by 6SourcePDFScholar
2024

Music Style Transfer with Time-Varying Inversion of Diffusion Models

AAAI 2024technical

With the development of diffusion models, text-guided image style transfer has demonstrated great controllable and high-quality results. However, the utilization of text for diverse music style transfer poses significant challenges, primarily due to the limited availability of matched audio-text dat…

2024

OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling

NeurIPS 2024poster

Constrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a variety of early-stage interaction technologies. Simultaneously, the current mask visual language modeling (MVLM) fails t…

2024

SignGen: End-to-End Sign Language Video Generation with Latent Diffusion

ECCV 2024poster

"The seamless transformation of textual input into natural and expressive sign language holds profound societal significance. Sign language is not solely about hand gestures. It encompasses vital facial expressions and mouth movements essential for nuanced communication. Achieving both semantic prec…

2024

StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion

ECCV 2024poster

"Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trained text-to-image model into an auto-regressive manner. Although these models have shown notable progress, there are st…

2024

Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models

NeurIPS 2024poster

Due to the impressive zero-shot capabilities, pre-trained vision-language models (e.g. CLIP), have attracted widespread attention and adoption across various domains. Nonetheless, CLIP has been observed to be susceptible to adversarial examples. Through experimental analysis, we have observed a phen…

2024

Three Heads Are Better than One: Complementary Experts for Long-Tailed Semi-supervised Learning

AAAI 2024technical

We address the challenging problem of Long-Tailed Semi-Supervised Learning (LTSSL) where labeled data exhibit imbalanced class distribution and unlabeled data follow an unknown distribution. Unlike in balanced SSL, the generated pseudo-labels are skewed towards head classes, intensifying the trainin…

2023

Active Exploration of Multimodal Complementarity for Few-Shot Action Recognition

CVPR 2023poster

Recently, few-shot action recognition receives increasing attention and achieves remarkable progress. However, previous methods mainly rely on limited unimodal data (e.g., RGB frames) while the multimodal information remains relatively underexplored. In this paper, we propose a novel Active Multimod…

Cited by 42SourcePDFScholar
2023

Cascade Evidential Learning for Open-World Weakly-Supervised Temporal Action Localization

CVPR 2023poster

Targeting at recognizing and localizing action instances with only video-level labels during training, Weakly-supervised Temporal Action Localization (WTAL) has achieved significant progress in recent years. However, living in the dynamically changing open world where unknown actions constantly spri…

Cited by 21SourcePDFScholar
2023

Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio-Visual Event Perception

CVPR 2023poster

With only video-level event labels, this paper targets at the task of weakly-supervised audio-visual event perception (WS-AVEP), which aims to temporally localize and categorize events belonging to each modality. Despite the recent progress, most existing approaches either ignore the unsynchronized…

2023

GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis

CVPR 2023poster

Synthesizing high-fidelity complex images from text is challenging. Based on large pretraining, the autoregressive and diffusion models can synthesize photo-realistic images. Although these large models have shown notable progress, there remain three flaws. 1) These models require tremendous trainin…

2023

Inversion-Based Style Transfer With Diffusion Models

CVPR 2023poster

The artistic style within a painting is the means of expression, which includes not only the painting material, colors, and brushstrokes, but also the high-level attributes, including semantic elements and object shapes. Previous arbitrary example-guided artistic image generation methods often fail…

2023

Multi-modal Queried Object Detection in the Wild

NeurIPS 2023poster

We introduce MQ-Det, an efficient architecture and pre-training strategy design to utilize both textual description with open-set generalization and visual exemplars with rich description granularity as category queries, namely, Multi-modal Queried object Detection, for real-world detection with bot…

2023

UTM: A Unified Multiple Object Tracking Model With Identity-Aware Feature Enhancement

CVPR 2023poster

Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. H…

Cited by 66SourcePDFScholar
2023

Unlearnable Clusters: Towards Label-Agnostic Unlearnable Examples

CVPR 2023poster

There is a growing interest in developing unlearnable examples (UEs) against visual privacy leaks on the Internet. UEs are training samples added with invisible but unlearnable noise, which have been found can prevent unauthorized training of machine learning models. UEs typically are generated via…

2023

VQACL: A Novel Visual Question Answering Continual Learning Setting

CVPR 2023poster

Research on continual learning has recently led to a variety of work in unimodal community, however little attention has been paid to multimodal tasks like visual question answering (VQA). In this paper, we establish a novel VQA Continual Learning setting named VQACL, which contains two key componen…

2023

Variational Causal Inference Network for Explanatory Visual Question Answering

ICCV 2023poster

Explanatory Visual Question Answering (EVQA) is a recently proposed multimodal reasoning task that requires answering visual questions and generating multimodal explanations for the reasoning processes. Unlike traditional Visual Question Answering (VQA) which focuses solely on answering, EVQA aims t…

Cited by 16PDFcodeScholar
2022

Cross-Modal Federated Human Activity Recognition via Modality-Agnostic and Modality-Specific Representation Learning

AAAI 2022technical

In this paper, we propose a new task of cross-modal federated human activity recognition (CMF-HAR), which is conducive to promote the large-scale use of the HAR model on more local devices. To address the new task, we propose a feature-disentangled activity recognition network (FDARN), which has fiv…

Cited by 30SourcePDFScholar
2022

DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis

CVPR 2022oral

Synthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between g…

Cited by 349PDFcodeScholar
2022

Dual-Evidential Learning for Weakly-Supervised Temporal Action Localization

ECCV 2022poster

"Weakly-supervised temporal action localization (WS-TAL) aims to localize the action instances and recognize their categories with only video-level labels. Despite great progress, existing methods suffer from severe action-background ambiguity, which mainly comes from background noise introduced by…

2022

Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer

AAAI 2022technical

Vision transformers (ViTs) have recently received explosive popularity, but the huge computational cost is still a severe issue. Since the computation complexity of ViT is quadratic with respect to the input sequence length, a mainstream paradigm for computation reduction is to reduce the number of…

2022

Fine-Grained Temporal Contrastive Learning for Weakly-Supervised Temporal Action Localization

CVPR 2022poster

We target at the task of weakly-supervised action localization (WSAL), where only video-level action labels are available during model training. Despite the recent progress, existing methods mainly embrace a localization-by-classification paradigm and overlook the fruitful fine-grained temporal dist…

Cited by 104PDFcodeScholar
2022

StyTr2: Image Style Transfer With Transformers

CVPR 2022poster

The goal of image style transfer is to render an image with artistic features guided by a style reference while maintaining the original content. Owing to the locality in convolutional neural networks (CNNs), extracting and maintaining the global information of input images is difficult. Therefore,…

Cited by 379PDFcodeScholar
2021

Arbitrary Video Style Transfer via Multi-Channel Correlation

AAAI 2021technical

Video style transfer is attracting increasing attention from the artificial intelligence community because of its numerous applications, such as augmented reality and animation production. Relative to traditional image style transfer, video style transfer presents new challenges, including how to ef…

2021

Dual Adversarial Graph Neural Networks for Multi-label Cross-modal Retrieval

AAAI 2021technical

Cross-modal retrieval has become an active study field with the expanding scale of multimodal data. To date, most existing methods transform multimodal data into a common representation space where semantic similarities between items can be directly measured across different modalities. However, the…

Cited by 70SourcePDFScholar
2021

ECKPN: Explicit Class Knowledge Propagation Network for Transductive Few-Shot Learning

CVPR 2021poster

Recently, the transductive graph-based methods have achieved great success in the few-shot classification task. However, most existing methods ignore exploring the class-level knowledge that can be easily learned by humans from just a handful of samples. In this paper, we propose an Explicit Class K…

Cited by 77PDFScholar
2021

Fast Video Moment Retrieval

ICCV 2021poster

This paper targets at fast video moment retrieval (fast VMR), aiming to localize the target moment efficiently and accurately as queried by a given natural language sentence. We argue that most existing VMR approaches can be divided into three modules namely video encoder, text encoder, and cross-mo…

Cited by 137PDFScholar
2021

Unveiling the Potential of Structure Preserving for Weakly Supervised Object Localization

CVPR 2021poster

Weakly supervised object localization (WSOL) remains an open problem due to the deficiency of finding object extent information using a classification network. While prior works struggle to localize objects by various spatial regularization strategies, we argue that how to extract object structural…

Cited by 110PDFcodeScholar
2020

Dynamic Refinement Network for Oriented and Densely Packed Object Detection

CVPR 2020oral

Object detection has achieved remarkable progress in the past decade. However, the detection of oriented and densely packed objects remains challenging because of following inherent reasons: (1) receptive fields of neurons are all axis-aligned and of the same shape, whereas objects are usually of di…

Cited by 411PDFcodeScholar
2019

GCAN: Graph Convolutional Adversarial Network for Unsupervised Domain Adaptation

CVPR 2019poster

To bridge source and target domains for domain adaptation, there are three important types of information including data structure, domain label, and class label. Most existing domain adaptation approaches exploit only one or two types of this information and cannot make them complement and enhance…

Cited by 169PDFcodeScholar
2018

Joint Pose and Expression Modeling for Facial Expression Recognition

CVPR 2018poster

Facial expression recognition (FER) is a challenging task due to different expressions under arbitrary poses. Most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifiers for each pose. Different from existing methods, in this paper, we…

2015

Matching-CNN Meets KNN: Quasi-Parametric Human Parsing

CVPR 2015poster

Both parametric and non-parametric approaches have demonstrated encouraging performances in the human parsing task, namely segmenting a human image into several semantic regions (e.g., hat, bag, left arm, face). In this work, we aim to develop a new solution with the advantages of both methodologie…

Cited by 203SourcePDFScholar
2015

Structural Sparse Tracking

CVPR 2015poster

Sparse representation has been applied to visual tracking by finding the best target candidate with minimal reconstruction error by use of target templates. However, most sparse representation based trackers only consider holistic or local representations and do not make full use of the intrinsic st…

Cited by 216SourcePDFScholar