← Search

Mingze Xu

19 accepted papers

2025

MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

ICLR 2025poster

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systema…

Cited by 29SourcePDFScholar
2025

StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant

NeurIPS 2025poster

We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in adapting existing models into online scenarios: (1) limited capability for multi-turn real-time understanding, and (2) lac…

Cited by 0SourceScholar
2025

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

NeurIPS 2025poster

We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More imp…

Cited by 0SourceScholar
2023

SkeleTR: Towards Skeleton-based Action Recognition in the Wild

ICCV 2023poster

We present SkeleTR, a new framework for skeleton-based action recognition. In contrast to prior work, which focuses mainly on controlled environments, we target in-the-wild scenarios that typically involve a variable number of people and various forms of interaction between people. SkeleTR works wit…

Cited by 35PDFScholar
2022

An In-depth Study of Stochastic Backpropagation

NeurIPS 2022accept

In this paper, we provide an in-depth study of Stochastic Backpropagation (SBP) when training deep neural networks for standard image classification and object detection tasks. During backward propagation, SBP calculates gradients by using only a subset of feature maps to save GPU memory and computa…

2022

Stochastic Backpropagation: A Memory Efficient Strategy for Training Video Models

CVPR 2022oral

We propose a memory efficient method, named Stochastic Backpropagation (SBP), for training deep neural networks on videos. It is based on the finding that gradients from incomplete execution for backpropagation can still effectively train the models with minimal accuracy loss, which attributes to th…

Cited by 22PDFcodeScholar
2022

TubeR: Tubelet Transformer for Video Action Detection

CVPR 2022oral

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in video by simulta…

Cited by 95PDFScholar
2021

Long Short-Term Transformer for Online Action Detection

NeurIPS 2021spotlight

We present Long Short-term TRansformer (LSTR), a temporal modeling algorithm for online action detection, which employs a long- and short-term memory mechanism to model prolonged sequence data. It consists of an LSTR encoder that dynamically leverages coarse-scale historical information from an exte…

2019

Egocentric Vision-based Future Vehicle Localization for Intelligent Driving Assistance Systems

ICRA 2019poster

Predicting the future location of vehicles is essential for safety-critical applications such as advanced driver assistance systems (ADAS) and autonomous driving. This paper introduces a novel approach to simultaneously predict both the location and scale of target vehicles in the first-person (egoc…

Cited by 172SourceScholar
2019

Embodied Amodal Recognition: Learning to Move to Perceive Objects

ICCV 2019poster

Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment and actively control the viewing angle to better understand object shapes and semantics. In this…

Cited by 75PDFScholar
2019

StartNet: Online Detection of Action Start in Untrimmed Videos

ICCV 2019poster

We propose StartNet to address Online Detection of Action Start (ODAS) where action starts and their associated categories are detected in untrimmed, streaming videos. Previous methods aim to localize action starts by learning feature representations that can directly separate the start point from i…

Cited by 70PDFcodeScholar
2019

Temporal Recurrent Networks for Online Action Detection

ICCV 2019poster

Most work on temporal action detection is formulated as an offline problem, in which the start and end times of actions are determined after the entire video is fully observed. However, important real-time applications including surveillance and driver assistance systems require identifying actions…

Cited by 233PDFcodeScholar
2019

Unsupervised Traffic Accident Detection in First-Person Videos

IROS 2019poster

Recognizing abnormal events such as traffic violations and accidents in natural driving scenes is essential for successful autonomous driving and advanced driver assistance systems. However, most work on video anomaly detection suffers from two crucial drawbacks. First, they assume cameras are fixed…

Cited by 209SourcecodeScholar
2018

Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos

ECCV 2018poster

In a world of pervasive cameras, public spaces are often captured from multiple perspectives by cameras of different types, both fixed and mobile. An important problem is to organize these heterogeneous collections of videos by finding connections between them, such as identifying correspondences be…

Cited by 46SourcePDFScholar
2017

Identifying First-Person Camera Wearers in Third-Person Videos

CVPR 2017poster

We consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establ…

Cited by 77PDFScholar