← Search

Xiaojun Chang

73 accepted papers

2026

BiOTPrompt: Bidirectional Optimal Transport Guided Prompting for Disease Evolution-aware Radiology Report Generation

CVPR 2026

Radiology report generation (RRG) aims to automatically describe medical images via free-text reports. In clinical practice, comparing current and prior chest X-rays is essential for assessing disease progression, motivating the development of longitudinal RRG methods. However, most existing approac

Cited by 0SourcecodeScholar
2026

CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning

CVPR 2026

Group-relative reinforcement learning with verifiable rewards (RLVR) often wastes the most informative data it already has--the failures. When all rollouts are wrong, gradients stall; when one happens to be correct, the update usually ignores why the others are close-but-wrong, and credit can be mis

Cited by 0SourcecodeScholar
2026

Correspondence Coverage Matters for Multi-Modal Dataset Distillation

AAAI 2026technical

Multi-modal dataset distillation (DD) condenses large datasets into compact ones that retain task efficacy by capturing correspondence patterns, i.e., shared semantics between paired modalities. However, such patterns rely on cross-modal similarity and cannot be faithfully captured by intra-modal si

Cited by 0SourcePDFScholar
2026

Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events

CVPR 2026

Multimodal Summarization (MMS) aims to generate concise textual summaries by understanding and integrating information across videos, transcripts, and images. However, existing approaches still suffer from three main challenges: (1) reliance on domain-specific supervision, (2) implicit fusion with w

Cited by 0SourcecodeScholar
2026

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

CVPR 2026

Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training these models on high-resolution, multi-frame data pose significant challenges. In this paper, we propose Entropy-Guided Pr

Cited by 0SourcecodeScholar
2026

GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning

ICML 2026poster

Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Models (MLLMs), where geometry information is essential. Existing methods often addre…

Cited by 0SourceScholar
2026

MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models

AAAI 2026technical

Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured kn

Cited by 0SourcePDFScholar
2026

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

ICLR 2026poster

Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challen…

Cited by 0SourceScholar
2026

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

AAAI 2026technical

Referring Audio-Visual Segmentation (Ref-AVS) aims to segment target objects in audible videos based on given reference expressions. Prior works typically rely on learning latent embeddings via multimodal fusion to prompt a tunable SAM/SAM2 decoder for segmentation, which requires strong pixel-level

Cited by 0SourcePDFScholar
2026

Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models

AAAI 2026technical

Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant, their property of modeling the entire image in latent space makes it challenging for the results to align well with prom

Cited by 0SourcePDFScholar
2025

Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

AAAI 2025technical

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for longer, untrimmed videos. This task seeks to identify and temporal…

2025

HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation

AAAI 2025technical

Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long sequence dependencies when incorporating historical informatio…

2025

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified mode…

2025

RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation

CVPR 2025poster

Vision-and-Language Navigation (VLN) suffers from the limited diversity and scale of training data, primarily constrained by the manual curation of existing simulators.To address this, we introduce RoomTour3D, a video-instruction dataset derived from web-based room tour videos that capture real-worl…

Cited by 3SourcePDFScholar
2025

Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

ICLR 2025poster

A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark \dataset with de…

Cited by 22SourcePDFScholar
2025

Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes

ICLR 2025poster

Recent advancements in human motion synthesis have focused on specific types of motions, such as human-scene interaction, locomotion or human-human interaction, however, there is a lack of a unified system capable of generating a diverse combination of motion types. In response, we introduce *Sitcom…

2025

Towards Efficient General Feature Prediction in Masked Skeleton Modeling

ICCV 2025poster

Recent advances in the masked autoencoder (MAE) paradigm have significantly propelled self-supervised skeleton-based action recognition. However, most existing approaches limit reconstruction targets to raw joint coordinates or their simple variants, resulting in computational redundancy and limited…

Cited by 0SourcePDFScholar
2025

Towards Open-Vocabulary Audio-Visual Event Localization

CVPR 2025poster

The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible.Most research in this field assumes a closed-set setting, which restricts these models' ability to handle test data containing event categories absent (unseen) during…

2024

Label-anticipated Event Disentanglement for Audio-Visual Video Parsing

ECCV 2024poster

"Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identification challenging. While traditional methods usually focus on improving the early audio-visual encoders to embed more e…

Cited by 15SourcePDFScholar
2024

Learning with Counterfactual Explanations for Radiology Report Generation

ECCV 2024poster

"Due to the common content of anatomy, radiology images with their corresponding reports exhibit high similarity. Such inherent data bias can predispose automatic report generation models to learn entangled and spurious representations resulting in misdiagnostic reports. To tackle these, we propose…

2024

LongVLM: Efficient Long Video Understanding via Large Language Models

ECCV 2024oral

"Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video representations through pooling or query aggregation over a vast number of visual tokens, making computational and memo…

2024

MLP Can Be A Good Transformer Learner

CVPR 2024poster

Self-attention mechanism is the key of the Transformer but often criticized for its computation demands. Previous token pruning works motivate their methods from the view of computation redundancy but still need to load the full network and require same memory costs. This paper introduces a novel st…

2024

Masked Distillation Advances Self-Supervised Transformer Architecture Search

ICLR 2024poster

Transformer architecture search (TAS) has achieved remarkable progress in automating the neural architecture design process of vision transformers. Recent TAS advancements have discovered outstanding transformer architectures while saving tremendous labor from human experts. However, it is still cum…

Cited by 2SourcePDFScholar
2024

Maximum Entropy Heterogeneous-Agent Reinforcement Learning

ICLR 2024spotlight

*Multi-agent reinforcement learning* (MARL) has been shown effective for cooperative games in recent years. However, existing state-of-the-art methods face challenges related to sample complexity, training instability, and the risk of converging to a suboptimal Nash Equilibrium. In this paper, we pr…

Cited by 18SourcePDFScholar
2024

Noisy Correspondence Learning with Self-Reinforcing Errors Mitigation

AAAI 2024technical

Cross-modal retrieval relies on well-matched large-scale datasets that are laborious in practice. Recently, to alleviate expensive data collection, co-occurring pairs from the Internet are automatically harvested for training. However, it inevitably includes mismatched pairs, i.e., noisy corresponde…

Cited by 7SourcePDFScholar
2024

ProAgent: Building Proactive Cooperative Agents with Large Language Models

AAAI 2024technical

Building agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they int…

2024

SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image Generation

AAAI 2024technical

Despite significant progress in Text-to-Image (T2I) generative models, even lengthy and complex text descriptions still struggle to convey detailed controls. In contrast, Layout-to-Image (L2I) generation, aiming to generate realistic and complex scene images from user-specified layouts, has risen to…

Cited by 16SourcePDFScholar
2024

SWAP-NAS: Sample-Wise Activation Patterns for Ultra-fast NAS

ICLR 2024spotlight

Training-free metrics (a.k.a. zero-cost proxies) are widely used to avoid resource-intensive neural network training, especially in Neural Architecture Search (NAS). Recent studies show that existing training-free metrics have several limitations, such as limited correlation and poor generalisation…

2024

Video Recognition in Portrait Mode

CVPR 2024poster

The creation of new datasets often presents new challenges for video recognition and can inspire novel ideas while addressing these challenges. While existing datasets mainly comprise landscape mode videos our paper seeks to introduce portrait mode videos to the research community and highlight the…

2023

Dynamic Graph Enhanced Contrastive Learning for Chest X-Ray Report Generation

CVPR 2023poster

Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced data-driven neural networks with medical knowledge graphs to eliminate the severe visual and textual bias in this task. Th…

2023

FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration

ICCV 2023poster

Multi-modality fusion and multi-task learning are becoming trendy in 3D autonomous driving scenario, considering robust prediction and computation budget. However, naively extending the existing framework to the domain of multi-modality multi-task learning remains ineffective and even poisonous due…

Cited by 10PDFScholar
2023

HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object Segmentation

ICCV 2023poster

Referring Video Object Segmentation (RVOS) is to segment the object instance from a given video, according to the textual description of this object. However, in the open world, the object descriptions are often diversified in contents and flexible in lengths. This leads to the key difficulty in RVO…

Cited by 30PDFScholar
2023

Mask Propagation for Efficient Video Semantic Segmentation

NeurIPS 2023poster

Video Semantic Segmentation (VSS) involves assigning a semantic label to each pixel in a video sequence. Prior work in this field has demonstrated promising results by extending image semantic segmentation models to exploit temporal relationships across video frames; however, these approaches often…

2023

Towards Real-Time Person Search with Invariant Feature Learning

ICASSP 2023accepted

Person search aims to locate a query person in a gallery of unconstrained scene images, which has many real-world applications. However, existing methods directly build off of advances in object detection for better performance rather than efficiency. Complex designs in heavy-weight detectors are re…

Cited by 0SourceScholar
2023

ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

ICLR 2023poster

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic alignment, while ignoring the correspondence among multiple au…

2023

Vision Language Navigation with Knowledge-driven Environmental Dreamer

IJCAI 2023poster

Vision-language navigation (VLN) requires an agent to perceive visual observation in a house scene and navigate step-by-step following natural language instruction. Due to the high cost of data annotation and data collection, current VLN datasets provide limited instruction-trajectory data samples.…

Cited by 2SourcePDFScholar
2022

AdverFacial: Privacy-Preserving Universal Adversarial Perturbation Against Facial Micro-Expression Leakages

ICASSP 2022accepted

Privacy safeguards are crucial, notably now with increased virtual conferencing usage during the Covid pandemic. In contrast to conventional facial expressions that are visually obvious to humans, micro-expressions are involuntary and transient facial expressions, commonly manifested involuntarily w…

Cited by 0SourceScholar
2022

An Efficient Spatio-Temporal Pyramid Transformer for Action Detection

ECCV 2022poster

"The task of action detection aims at deducing both the action category and localization of the start and end moment for each action instance in a long, untrimmed video. While vision Transformers have driven the recent advances in video understanding, it is non-trivial to design an efficient archite…

2022

Automated Progressive Learning for Efficient Training of Vision Transformers

CVPR 2022poster

Recent advances in vision Transformers (ViTs) have come with a voracious appetite for computing power, high-lighting the urgent need to develop efficient training methods for ViTs. Progressive learning, a training scheme where the model capacity grows progressively during training, has started showi…

Cited by 49PDFcodeScholar
2022

BaLeNAS: Differentiable Architecture Search via the Bayesian Learning Rule

CVPR 2022poster

Differentiable Architecture Search (DARTS) has received massive attention in recent years, mainly because it significantly reduces the computational cost through weight sharing and continuous relaxation. However, more recent works find that existing differentiable NAS techniques struggle to outperfo…

Cited by 24PDFScholar
2022

Beyond Fixation: Dynamic Window Visual Transformer

CVPR 2022poster

Recently, a surge of interest in visual transformers is to reduce the computational cost by limiting the calculation of self-attention to a local window. Most current work uses a fixed single-scale window for modeling by default, ignoring the impact of window size on model performance. However, this…

Cited by 41PDFcodeScholar
2022

Cross-Modal Clinical Graph Transformer for Ophthalmic Report Generation

CVPR 2022poster

Automatic generation of ophthalmic reports using data-driven neural networks has great potential in clinical practice. When writing a report, ophthalmologists make inferences with prior clinical knowledge. This knowledge has been neglected in prior medical report generation methods. To endow models…

Cited by 55PDFcodeScholar
2022

Dual-AI: Dual-Path Actor Interaction Learning for Group Activity Recognition

CVPR 2022oral

Learning spatial-temporal relation among multiple actors is crucial for group activity recognition. Different group activities often show the diversified interactions between actors in the video. Hence, it is often difficult to model complex group activities from a single view of spatial-temporal ac…

Cited by 79PDFScholar
2022

Knowledge Distillation via the Target-Aware Transformer

CVPR 2022oral

Knowledge distillation becomes a de facto standard to improve the performance of small neural networks. Most of the previous works propose to regress the representational features from the teacher to the student in a one-to-one spatial matching fashion. However, people tend to overlook the fact that…

Cited by 150PDFcodeScholar
2022

PAR: Political Actor Representation Learning with Social Context and Expert Knowledge

EMNLP 2022main

Modeling the ideological perspectives of political actors is an essential task in computational political science with applications in many downstream tasks. Existing approaches are generally limited to textual data and voting records, while they neglect the rich social context and valuable expert k…

2022

Policy Diagnosis via Measuring Role Diversity in Cooperative Multi-agent RL

ICML 2022spotlight

Cooperative multi-agent reinforcement learning (MARL) is making rapid progress for solving tasks in a grid world and real-world scenarios, in which agents are given different attributes and goals, resulting in different behavior through the whole multi-agent task. In this study, we quantify the agen…

Cited by 34SourcePDFScholar
2022

Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation With Reliable Voted Pseudo Labels

CVPR 2022poster

In this paper, we propose an unsupervised domain adaptation method for deep point cloud representation learning. To model the internal structures in target point clouds, we first propose to learn the global representations of unlabeled data by scaling up or down point clouds and then predicting the…

Cited by 68PDFScholar
2021

BossNAS: Exploring Hybrid CNN-Transformers With Block-Wisely Self-Supervised Neural Architecture Search

ICCV 2021poster

A myriad of recent breakthroughs in hand-crafted neural architectures for visual recognition have highlighted the urgent need to explore hybrid architectures consisting of diversified building blocks. Meanwhile, neural architecture search methods are surging with an expectation to reduce human effor…

Cited by 142PDFcodeScholar
2021

Exploring Inter-Channel Correlation for Diversity-Preserved Knowledge Distillation

ICCV 2021poster

Knowledge Distillation has shown very promising ability in transferring learned representation from the larger model (teacher) to the smaller one (student). Despite many efforts, prior methods ignore the important role of retaining inter-channel correlation of features, leading to the lack of captur…

Cited by 125PDFcodeScholar
2021

FFA-IR: Towards an Explainable and Reliable Medical Report Generation Benchmark

NeurIPS 2021poster

The automatic generation of long and coherent medical reports given medical images (e.g. Chest X-ray and Fundus Fluorescein Angiography (FFA)) has great potential to support clinical practice. Researchers have explored advanced methods from computer vision and natural language processing to incorpor…

Cited by 48SourcecodeScholar
2021

Person Search Challenges and Solutions: A Survey

IJCAI 2021poster

Person search has drawn increasing attention due to its real-world applications and research significance. Person search aims to find a probe person in a gallery of scene images with a wide range of applications, such as criminals search, multicamera tracking, missing person search, etc. Early perso…

Cited by 17SourcePDFScholar
2021

SOON: Scenario Oriented Object Navigation With Graph-Based Exploration

CVPR 2021poster

The ability to navigate like a human towards a language-guided target from anywhere in a 3D embodied environment is one of the 'holy grail' goals of intelligent robots. Most visual navigation benchmarks, however, focus on navigating toward a target from a fixed starting point, guided by an elaborate…

Cited by 131PDFcodeScholar
2021

UPDeT: Universal Multi-agent RL via Policy Decoupling with Transformers

ICLR 2021spotlight

Recent advances in multi-agent reinforcement learning have been largely limited in training one model from scratch for every new task. The limitation is due to the restricted model architecture related to fixed input and output dimensions. This hinders the experience accumulation and transfer of the…

Cited by 0SourcePDFScholar
2021

Vision-Language Navigation With Random Environmental Mixup

ICCV 2021poster

Vision-language Navigation (VLN) task requires an agent to perceive both the visual scene and natural language and navigate step-by-step. Large data bias makes the VLN task challenging, which is caused by the disparity ratio between small data scale and large navigation space. Previous works have pr…

Cited by 99PDFcodeScholar
2021

iDARTS: Differentiable Architecture Search with Stochastic Implicit Gradients

ICML 2021spotlight

Differentiable ARchiTecture Search(DARTS) has recently become the mainstream in the neural architecture search (NAS) due to its efficiency and simplicity. With a gradient-based bi-level optimization, DARTS alternately optimizes the inner model weights and the outer architecture parameter in a weight…

2020

Block-Wisely Supervised Neural Architecture Search With Knowledge Distillation

CVPR 2020poster

Neural Architecture Search (NAS), aiming at automatically designing network architectures by machines, is expected to bring about a new revolution in machine learning. Despite these high expectation, the effectiveness and efficiency of existing NAS solutions are unclear, with some recent works going…

Cited by 244PDFcodeScholar
2020

Differentiable Neural Architecture Search in Equivalent Space with Exploration Enhancement

NeurIPS 2020poster

Recent works on One-Shot Neural Architecture Search (NAS) mostly adopt a bilevel optimization scheme to alternatively optimize the supernet weights and architecture parameters after relaxing the discrete search space into a differentiable space. However, the non-negligible incongruence in their rela…

Cited by 42SourcePDFScholar
2020

Hierarchical Neural Architecture Search for Deep Stereo Matching

NeurIPS 2020poster

To reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the netw…

2020

Mining Inter-Video Proposal Relations for Video Object Detection

ECCV 2020poster

Recent studies have shown that, context aggregating information from proposals in different frames can clearly enhance the performance of video object detection. However, these approaches mainly exploit the intra-proposal relation within single video, while ignoring the intra-proposal relation among…

2020

Overcoming Multi-Model Forgetting in One-Shot NAS With Diversity Maximization

CVPR 2020poster

One-Shot Neural Architecture Search (NAS) significantly improves the computational efficiency through weight sharing. However, this approach also introduces multi-model forgetting during the supernet training (architecture search phase), where the performance of previous architectures degrade when s…

Cited by 103PDFcodeScholar
2020

Quadratic Sparse Gaussian Graphical Model Estimation Method for Massive Variables

IJCAI 2020poster

We consider the problem of estimating a sparse Gaussian Graphical Model with a special graph topological structure and more than a million variables. Most previous scalable estimators still contain expensive calculation steps (e.g., matrix inversion or Hessian matrix calculation) and become infeasib…

Cited by 0SourcePDFScholar
2020

Vision-Dialog Navigation by Exploring Cross-Modal Memory

CVPR 2020poster

Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language…

Cited by 55PDFcodeScholar
2018

RCAA: Relational Context-Aware Agents for Person Search

ECCV 2018poster

We aim to search for a target person from a gallery of whole scene images for which the annotations of pedestrian bounding boxes are unavailable. Previous approaches to this problem have relied on a pedestrian proposal net, which may generate redundant proposals and increase the computational burden…

Cited by 129SourcePDFScholar
2018

Reinforcement Cutting-Agent Learning for Video Object Segmentation

CVPR 2018poster

Video object segmentation is a fundamental yet challenging task in computer vision community. In this paper, we formulate this problem as a Markov Decision Process, where agents are learned to segment object regions under a deep reinforcement learning framework. Essentially, learning agents for segm…

Cited by 105SourcePDFScholar
2017

Complex Event Detection by Identifying Reliable Shots From Untrimmed Videos

ICCV 2017poster

The goal of complex event detection is to automatically detect whether an event of interest happens in temporally untrimmed long videos which usually consist of multiple video shots. Observing some video shots in positive (resp. negative) videos are irrelevant (resp. relevant) to the given event cla…

Cited by 55PDFScholar
2016

They Are Not Equally Reliable: Semantic Event Search Using Differentiated Concept Classifiers

CVPR 2016poster

Complex event detection on unconstrained Internet videos has seen much progress in recent years. However, state-of-the-art performance degrades dramatically when the number of positive training exemplars falls short. Since label acquisition is costly, laborious, and time-consuming, there is a real n…

Cited by 38PDFScholar
2015

Complex Event Detection using Semantic Saliency and Nearly-Isotonic SVM

ICML 2015poster

We aim to detect complex events in long Internet videos that may last for hours. A major challenge in this setting is that only a few shots in a long video are relevant to the event of interest while others are irrelevant or even misleading. Instead of indifferently pooling the shots, we first defin…

Cited by 83SourcePDFScholar