← Search

Weidi Xie

65 accepted papers

2026

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation

ICML 2026poster

Post-training of flow matching models—aligning the output distribution with a high-quality target—is mathematically equivalent to imitation learning. While Supervised Fine-Tuning mimics expert demonstrations effectively, it cannot correct policy drift in unseen states. Preference optimization method…

Cited by 0SourceScholar
2026

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and streaming scenarios--poses a major challenge to their scalability and real-world

Cited by 0SourcecodeScholar
2026

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

CVPR 2026

Existing evaluations of multimodal large language models (MLLMs) on spatial intelligence are typically fragmented and limited in scope. In this work, we conduct a holistic assessment of the spatial understanding abilities of modern MLLMs and propose complementary data-driven and agent-based solution

Cited by 0SourcecodeScholar
2026

Versatile Vision-Language Model for 3D Computed Tomography

AAAI 2026technical

Representation learning serves as a foundational component of medical vision-language models (MVLMs), enabling cross-modal alignment, semantic consistency, and enhanced generalization capabilities for downstream tasks. As generalist models rapidly evolve, there is a pressing need to unify diverse do

Cited by 0SourcePDFScholar
2026

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

ICLR 2026poster

We introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features: (i) collaboration of omni-modality, we design the evaluation tasks to f…

Cited by 0SourcecodeScholar
2025

A Sanity Check for AI-generated Image Detection

ICLR 2025poster

With the rapid development of generative models, discerning AI-generated content has evoked increasing attention from both industry and academia. In this paper, we conduct a sanity check on whether the task of AI-generated image detection has been solved. To start with, we present Chameleon dataset,…

2025

EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

ICLR 2025poster

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and…

Cited by 0SourcePDFScholar
2025

LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant

CVPR 2025poster

With the rapid advancement of multimodal information retrieval, increasingly complex retrieval tasks have emerged. Existing methods predominately rely on task-specific fine-tuning of vision-language models, often those trained with image-text contrastive learning. In this paper, we explore the possi…

Cited by 8SourcePDFScholar
2025

Learning Streaming Video Representation via Multitask Training

ICCV 2025poster

Understanding continuous video streams plays a fundamental role in real-time applications, including embodied AI and autonomous driving. Unlike offline video processing, streaming video understanding requires the ability to process video streams frame by frame, preserve historical information, and m…

Cited by 0SourcePDFScholar
2025

MRGen: Segmentation Data Engine For Underrepresented MRI Modalities

ICCV 2025poster

Training medical image segmentation models for rare yet clinically important imaging modalities is challenging due to the scarcity of annotated data, and manual mask annotations can be costly and labor-intensive to acquire. This paper investigates leveraging generative models to synthesize data, for…

2025

Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning

ICLR 2025poster

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on aligning video representation with high-level narrations, overlooking the intricate d…

2025

Object-centric Video Question Answering with Visual Grounding and Referring

ICCV 2025poster

Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only responses, restricting the flexibility for object-centric, multi-round interactions…

Cited by 14SourcePDFScholar
2025

Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation

ICCV 2025poster

Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework that leverages "shots" as the fundamental units of video understanding. This includes extending temporal context to neighbo…

Cited by 0SourcePDFScholar
2025

Towards Universal Soccer Video Understanding

CVPR 2025poster

As a globally celebrated sport, soccer has attracted widespread interest from fans over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding.Specifically, we make the following contributions in this paper:(i) we introduce **SoccerReplay-1988**, t…

2025

Track-On: Transformer-based Online Point Tracking with Memory

ICLR 2025poster

In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appearance, lighting, perspective, and occlusions. We target online tracking on a frame-by-frame basis, making it suitable for…

2025

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

NeurIPS 2025poster

This paper presents a computational model for universal video temporal grounding, which accurately localizes temporal moments in videos based on natural language queries (e.g., questions or descriptions). Unlike existing methods that are often limited to specific video domains or durations, we prop…

Cited by 0SourcecodeScholar
2024

A General Protocol to Probe Large Vision Models for 3D Physical Understanding

NeurIPS 2024poster

Our objective in this paper is to probe large vision models to determine to what extent they ‘understand’ different physical properties of the 3D scene depicted in an image. To this end, we make the following contributions: (i) We introduce a general and lightweight protocol to evaluate whether feat…

2024

AutoAD III: The Prequel - Back to the Pixels

CVPR 2024poster

Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently visual language models for AD generation are limited by a lack of suitable training data and also their evaluation is hampe…

Cited by 19SourcePDFScholar
2024

InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

CVPR 2024poster

In this paper we present a novel paradigm to enhance the ability of object detector e.g. expanding categories or improving detection performance by training on syn- thetic dataset generated from diffusion models. Specifically we integrate an instance-level grounding head into a pre- trained generati…

Cited by 13SourcePDFScholar
2024

Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models

CVPR 2024poster

Generative models have recently exhibited exceptional capabilities in text-to-image generation but still struggle to generate image sequences coherently. In this work we focus on a novel yet challenging task of generating a coherent image sequence based on a given storyline denoted as open-ended vis…

2024

Knowledge-enhanced Visual-Language Pretraining for Computational Pathology

ECCV 2024oral

"In this paper, we consider the problem of visual representation learning for computational pathology, by exploiting large-scale image-text pairs gathered from public resources, along with the domain-specific knowledge in pathology. Specifically, we make the following contributions: (i) We curate a…

2024

Made to Order: Discovering monotonic temporal changes via self-supervised video ordering

ECCV 2024oral

"Our objective is to discover and localize monotonic temporal changes in a sequence of images. To achieve this, we exploit a simple proxy task of ordering a shuffled image sequence, with ‘time’ serving as a supervisory signal, since only changes that are monotonic with time can give rise to the corr…

Cited by 2SourcePDFScholar
2024

MatchTime: Towards Automatic Soccer Game Commentary Generation

EMNLP 2024main

Soccer is a globally popular sport with a vast audience, in this paper, we consider constructing an automatic soccer game commentary model to improve the audiences’ viewing experience. In general, we make the following contributions: *First*, observing the prevalent video-text misalignment in existi…

2024

RaTEScore: A Metric for Radiology Report Generation

EMNLP 2024main

This paper introduces a novel, entity-aware metric, termed as Radiological Report (Text) Evaluation (RaTEScore), to assess the quality of medical reports generated by AI models. RaTEScore emphasizes crucial medical entities such as diagnostic outcomes and anatomical details, and is robust against co…

2024

Retrieval-Augmented Egocentric Video Captioning

CVPR 2024poster

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper (1) we develop EgoI…

Cited by 38SourcePDFScholar
2024

VISA: Reasoning Video Object Segmentation via Large Language Model

ECCV 2024poster

"Existing Video Object Segmentation (VOS) relies on explicit user instructions, such as categories, masks, or short phrases, restricting their ability to perform complex video segmentation requiring reasoning with world knowledge. In this paper, we introduce a new task, Reasoning Video Object Segmen…

2023

AutoAD II: The Sequel - Who, When, and What in Movie Audio Description

ICCV 2023poster

Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presents notable challenges -- AD must occur only during existing pauses in dialogue, should refer to characters by name, and…

Cited by 47PDFScholar
2023

AutoAD: Movie Description in Context

CVPR 2023highlight

The objective of this paper is an automatic Audio Description (AD) model that ingests movies and outputs AD in text form. Generating high-quality movie AD is challenging due to the dependency of the descriptions on context, and the limited amount of training data available. In this work, we leverage…

2023

Collaboration Helps Camera Overtake LiDAR in 3D Detection

CVPR 2023poster

Camera-only 3D detection provides an economical solution with a simple configuration for localizing objects in 3D space compared to LiDAR-based detection systems. However, a major challenge lies in precise depth estimation due to the lack of direct 3D measurements in the input. Many previous methods…

2023

Joint-Relation Transformer for Multi-Person Motion Prediction

ICCV 2023poster

Multi-person motion prediction is a challenging problem due to the dependency of motion on both individual past movements and interactions with other people. Transformer-based methods have shown promising resultson this task, but they miss the explicit relation representation between joints, such as…

Cited by 13PDFcodeScholar
2023

Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision

CVPR 2023poster

In this paper, we consider the problem of open-vocabulary semantic segmentation (OVS), which aims to segment objects of arbitrary classes instead of pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor,…

2023

MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray Diagnosis

ICCV 2023poster

In this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the following contributions: First, unlike existing works that directly process the…

Cited by 125PDFcodeScholar
2023

Open-vocabulary Object Segmentation with Diffusion Models

ICCV 2023poster

The goal of this paper is to extract the visual-language correspondence from a pre-trained text-to-image diffusion model, in the form of segmentation map, i.e., simultaneously generating images and segmentation masks for the corresponding visual entities described in the text prompt. We make the fol…

Cited by 60PDFScholar
2023

OvarNet: Towards Open-Vocabulary Object Attribute Recognition

CVPR 2023poster

In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-vocabulary scenario. To achieve this goal, we make the following contributions: (…

2023

Towards Open-Vocabulary Video Instance Segmentation

ICCV 2023oral

Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel categories in real-world videos. To address this limitation, we make the following three contributions. First, we intro…

Cited by 36PDFcodeScholar
2022

Associating Objects and Their Effects in Video through Coordination Games

NeurIPS 2022accept

We explore a feed-forward approach for decomposing a video into layers, where each layer contains an object of interest along with its associated shadows, reflections, and other visual effects. This problem is challenging since associated effects vary widely with the 3D geometry and lighting conditi…

Cited by 5SourcePDFScholar
2022

PromptDet: Towards Open-Vocabulary Detection Using Uncurated Images

ECCV 2022poster

"The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make the following four contributions: (i) in pursuit of generalisation, we propose a two-stage open-vocabulary object detec…

2022

Prompting Visual-Language Models for Efficient Video Understanding

ECCV 2022poster

"Image-based visual-language (I-VL) pre-training has shown great success for learning joint visual-textual representations from large-scale web data, revealing remarkable ability for zero-shot generalisation. This paper presents a simple but strong baseline to efficiently adapt the pre-trained I-VL…

2022

Segmenting Moving Objects via an Object-Centric Layered Representation

NeurIPS 2022accept

The objective of this paper is a model that is able to discover, track and segment multiple moving objects in a video. We make four contributions: First, we introduce an object-centric segmentation model with a depth-ordered layer representation. This is implemented using a variant of the transforme…

2021

Localizing Visual Sounds the Hard Way

CVPR 2021poster

The objective of this work is to localize sound sources that are visible in a video without using manual annotations. Our key technical contribution is to show that, by training the network to explicitly discriminate challenging image fragments, even for images that do contain the object emitting th…

Cited by 234PDFScholar
2021

Self-Supervised Video Object Segmentation by Motion Grouping

ICCV 2021poster

Animals have evolved highly functional visual systems to understand motion, assisting perception even under complex environments. In this paper, we work towards developing a computer vision system able to segment objects by exploiting motion cues, i.e. motion segmentation. To achieve this, we introd…

Cited by 185PDFScholar
2020

Memory-augmented Dense Predictive Coding for Video Representation Learning

ECCV 2020poster

The objective of this paper is self-supervised learning from video, in particular for representations for action recognition. We make the following contributions: (i) We propose a new architecture and learning framework Memory-augmented Dense Predictive Coding (MemDPC) for the task. It is trained wi…

2020

Self-supervised Co-Training for Video Representation Learning

NeurIPS 2020poster

The objective of this paper is visual-only self-supervised video representation learning. We make the following contributions: (i) we investigate the benefit of adding semantic-class positives to instance-based Info Noise Contrastive Estimation (InfoNCE) training, showing that this form of supervise…

2020

Smooth-AP: Smoothing the Path Towards Large-Scale Image Retrieval

ECCV 2020poster

Optimising a ranking-based metric, such as Average Precision (AP), is notoriously challenging due to the fact that it is non-differentiable, and hence cannot be optimised directly using gradient-descent methods. To this end, we introduce an objective that optimises instead a smoothed approximation o…

2019

Utterance-level Aggregation for Speaker Recognition in the Wild

ICASSP 2019accepted

The objective of this paper is speaker recognition `in the wild' - where utterances may be of variable length and also contain irrelevant signals. Crucial elements in the design of deep networks for this task are the type of trunk (frame level) network, and the method of temporal aggregation. We pro…

Cited by 0SourceScholar