← Search

Limin Wang

119 accepted papers

2026

Balancing the Experts: Unlocking LoRA-MoE for GRPO via Mechanism-Aware Rewards

ICLR 2026poster

Parameter-efficient Mixture-of-Experts (MoE) architectures, such as LoRA-MoE, enable strong and generalizable fine-tuning. However, a critical problem arises when fine-tuning these architectures with advanced reinforcement learning algorithms such as Group Relative Policy Optimization (GRPO). Tradit…

Cited by 0SourceScholar
2026

CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval

ICLR 2026poster

Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only include short descriptions, limits their ability of detailed video understanding evaluation. To address this problem, w…

Cited by 0SourcecodeScholar
2026

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

CVPR 2026

Unsupervised learning of latent motion from Internet videos is crucial for robot learning. Existing discrete methods generally mitigate the shortcut learning caused by extracting excessive static backgrounds through vector quantization with a small codebook size. However, they suffer from informatio

Cited by 0SourcecodeScholar
2026

ExpVid: A Benchmark for Experiment Video Understanding & Reasoning

ICLR 2026poster

Multimodal Large Language Models (MLLMs) hold promise for accelerating scientific discovery by interpreting complex experimental procedures. However, their true capabilities are poorly understood, as existing benchmarks neglect the fine-grained and long-horizon nature of authentic laboratory work, e…

Cited by 0SourcecodeScholar
2026

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos

ICML 2026poster

The rapid advancement of AI-driven video generation has transformed content creation, while simultaneously increasing the risk of misinformation through localized manipulations in long-form videos. Existing video forensic methods predominantly operate on short, independent clips, and thus fail to ca…

Cited by 0SourceScholar
2026

Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment

AAAI 2026technical

Normalizing Flows (NFs) are a class of generative models distinguished by a mathematically invertible architecture, where the forward pass transforms data into a latent space for density estimation, and the reverse pass generates new samples from this space. This characteristic creates an intrinsic

Cited by 0SourcePDFScholar
2026

InternVideo-Next: Towards World-Understanding Video Models

CVPR 2026

Large-scale video-text pretraining achieves strong performance but depends on noisy, synthetic captions with limited semantic coverage, often overlooking implicit world knowledge such as object motion, 3D geometry, and physical cues. In contrast, masked video modeling (MVM) directly exploits spatiot

Cited by 0SourcecodeScholar
2026

RIVER: Real-time Video Interaction Benchmark

ICLR 2026poster

The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real-time interactivity. Addressing this gap, we introduce the Real-tIme Video intERaction Bench (RIVER Bench), designed for evaluating online…

Cited by 0SourcecodeScholar
2026

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

CVPR 2026

Enhancing temporal understanding of MLLMs is essential for long-form video analysis, supporting tasks such as temporal localization and time-sensitive question answering. While reinforcement learning (RL) has been explored for temporal reasoning, existing approaches are often limited to specific tas

Cited by 0SourceScholar
2026

TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

CVPR 2026

This paper does not introduce a novel method but instead establishes a straightforward, incremental, yet essential baseline for video temporal grounding (VTG), a core capability in video understanding. While multimodal large language models (MLLMs) excel at various video understanding tasks, the rec

Cited by 0SourceScholar
2026

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

CVPR 2026

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a unified framework for human-centric joint audio and video ge

Cited by 0SourceScholar
2026

UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation

ICLR 2026poster

Tokenizer is a crucial component for both visual understanding and generation. To advance toward the ultimate goal of universal modeling, recent research has focused on developing a unified tokenizer. However, existing tokenizers face a significant performance trade-off between understanding and gen…

Cited by 0SourceScholar
2026

VMonarch: Efficient Video Diffusion Transformers with Structured Attention

CVPR 2026

The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal attention patterns exhibited in Video DiTs can be naturally represented by the Monarch matrix. It is a class of structure

Cited by 2SourceScholar
2026

VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning

AAAI 2026technical

Recent advances in video understanding have been driven by MLLMs. But these MLLMs are good at analyzing short videos, while suffering from difficulties in understanding videos with a longer context. To address this difficulty, several agent paradigms have recently been proposed, using MLLMs as agen

Cited by 0SourcePDFScholar
2026

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

ICLR 2026poster

Long-context video modeling is critical for multimodal large language models (MLLMs), enabling them to process movies, online video streams, and so on. Despite its advances, handling long videos remains challenging due to the difficulty in efficiently understanding the extremely long video context.…

Cited by 0SourcecodeScholar
2026

VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric Videos

CVPR 2026

With the great advancement of video generation models, a growing number of content creators and researchers are leveraging these technologies to produce large volumes of human-centric videos for content creation and customized data generation for specific tasks. Although existing video generation mo

Cited by 0SourcecodeScholar
2026

VideoSeeker: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

ICML 2026poster

Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting their ability to identify sparse yet critical evidence amid extensive redundancy. We introduce VideoSeeker, a novel framework that supports iterative disc…

Cited by 13SourceScholar
2025

Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel

ICLR 2025poster

Creating high-quality data for training robust language-instructed agents is a long-lasting challenge in embodied AI. In this paper, we introduce a Self-Refining Data Flywheel (SRDF) that generates high-quality and large-scale navigational instruction-trajectory pairs by iteratively refining the dat…

2025

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

ICLR 2025poster

The existing video understanding benchmarks for multimodal large language models (MLLMs) mainly focus on short videos. The few benchmarks for long video understanding often rely on multiple-choice questions (MCQs). Due to the limitations of MCQ evaluations and the advanced reasoning abilities of MLL…

Cited by 5SourcePDFScholar
2025

Contextual AD Narration with Interleaved Multimodal Sequence

CVPR 2025poster

The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video contents, like movie. With video feature, text, character bank and context information as inputs, the generated ADs are able to correspond to the cha…

2025

Differentiable Solver Search for Fast Diffusion Sampling

ICML 2025poster

Diffusion models have demonstrated remarkable generation quality but at the cost of numerous function evaluations. Recently, advanced ODE-based solvers have been developed to mitigate the substantial computational demands of reverse-diffusion solving under limited sampling steps. However, these solv…

Cited by 0SourcePDFScholar
2025

Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

NeurIPS 2025poster

We introduce Eagle2.5, a frontier vision-language model (VLM) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist framework for both tasks. The proposed training framework incorporates A…

Cited by 0SourceScholar
2025

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

CVPR 2025highlight

The intuitive nature of drag-based interaction has led to its growing adoption for controlling object trajectories in image-to-video synthesis. Still, existing methods that perform dragging in the 2D space usually face ambiguity when handling out-of-plane movements. In this work, we augment the inte…

2025

LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization

NeurIPS 2025poster

We present LongVPO, a novel two‑stage Direct Preference Optimization framework that enables short‑context vision‑language models to robustly understand ultra‑long videos without any long‑video annotations. In Stage 1, we synthesize preference triples by anchoring questions to individual short clips,…

Cited by 0SourceScholar
2025

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

NeurIPS 2025poster

We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed videos containing relevant segments using multi-modal queries. It has the following features: **1) Practical retrieval par…

Cited by 0SourcecodeScholar
2025

Make Your Training Flexible: Towards Deployment-Efficient Video Models

ICCV 2025poster

Current video training methods rely on fixed spatiotemporal sampling grids to extract a predetermined number of tokens, limiting adaptability to diverse computational budgets and resulting in suboptimal accuracy-computation trade-offs. This rigidity constrains high-performance models trained in reso…

2025

MobileViCLIP: An Efficient Video-Text Model for Mobile Devices

ICCV 2025poster

Efficient lightweight neural networks have received increasing attention due to their faster reasoning speed and easier deployment on mobile devices. However, existing video models still focus on the larger ViT architecture, and few works attempt to build efficient architecture. Since many efficient…

2025

Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning

ICLR 2025poster

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on aligning video representation with high-level narrations, overlooking the intricate d…

2025

MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation

NeurIPS 2025poster

Image-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accurately modeling motion, which involves capturing physical constraints, object in…

Cited by 0SourceScholar
2025

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

ICLR 2025spotlight

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains th…

2025

Online Video Understanding: OVBench and VideoChat-Online

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have significantly progressed in offline video understanding. However, applying these models to real-world scenarios, such as autonomous driving and human-computer interaction, presents unique challenges due to the need for real-time processing of continuous…

Cited by 0SourcePDFScholar
2025

SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

ICLR 2025poster

In this paper, we introduce SPA, a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI. Our approach leverages differentiable neural rendering on multi-view images to endow a vanilla Vision Transformer (ViT) with intrinsic spatial understandi…

2025

Scalable Image Tokenization with Index Backpropagation Quantization

ICCV 2025poster

Existing vector quantization (VQ) methods struggle with scalability, largely attributed to the instability of the codebook that undergoes partial updates during training. The codebook is prone to collapse as utilization decreases, due to the progressively widening distribution gap between non-activa…

Cited by 0SourcePDFScholar
2025

Stochastic Layer-Wise Shuffle for Improving Vision Mamba Training

ICML 2025poster

Recent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies and their potential are still not sufficiently explored. In this paper, we investigate strategies for Vim and propose St…

2025

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

NeurIPS 2025spotlight

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To a…

Cited by 0SourceScholar
2025

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

CVPR 2025poster

Current multimodal large language models (MLLMs) struggle with fine-grained or precise understanding of visuals although they give comprehensive perception and reasoning in a spectrum of vision applications. Recent studies either develop tool-using or unify specific visual tasks into the autoregress…

2025

TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for lon…

Cited by 10SourcePDFScholar
2025

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning

CVPR 2025poster

Learning from multiple domains is a primary factor that influences the generalization of a single unified robot system. In this paper, we aim to learn the trajectory prediction model by using broad out-of-domain data to improve its performance and generalization ability. Trajectory model is designed…

2025

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

ICCV 2025poster

We present VRBench, the first long narrative video benchmark crafted for evaluating large models' multi-step reasoning capabilities, addressing limitations in existing evaluations that overlook temporal reasoning and procedural validity. It comprises 960 long videos (with an average duration of 1.6…

Cited by 0SourcePDFScholar
2025

VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception

NeurIPS 2025poster

Inducing reasoning in multimodal large language models (MLLMs) is critical for achieving human-level perception and understanding. Existing methods mainly leverage LLM reasoning to analyze parsed visuals, often limited by static perception stages. This paper introduces Visual Test-Time Scaling (VTTS…

Cited by 0SourceScholar
2025

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

ICCV 2025poster

Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this paper, we propose p-MoD, an efficient MLLM architecture that significantly reduces training and inference costs while ma…

2024

AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation

NeurIPS 2024poster

Pre-trained vision-language models (VLMs) have shown impressive results in various visual classification tasks. However, we often fail to fully unleash their potential when adapting them for new concept understanding due to limited information on new classes. To address this limitation, we introduce…

2024

Adapting Short-Term Transformers for Action Detection in Untrimmed Videos

CVPR 2024poster

Vision Transformer (ViT) has shown high potential in video recognition owing to its flexible design adaptable self-attention mechanisms and the efficacy of masked pre-training. Yet it remains unclear how to adapt these pre-trained short-term ViTs for temporal action detection (TAD) in untrimmed vide…

2024

Asymmetric Masked Distillation for Pre-Training Small Foundation Models

CVPR 2024poster

Self-supervised foundation models have shown great potential in computer vision thanks to the pre-training paradigm of masked autoencoding. Scale is a primary factor influencing the performance of these foundation models. However these large foundation models often result in high computational cost.…

2024

BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models

CVPR 2024poster

Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various downstream image synthesis tasks such as controllable image generation and image editing while downstream video synthesis tasks are less explore…

2024

Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?

NeurIPS 2024poster

Video understanding relies on accurate action detection for temporal analysis. However, existing mainstream methods have limitations in real-world applications due to their offline and closed-set evaluation approaches, as well as their dependence on manual annotations. To address these challenges an…

2024

EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

CVPR 2024poster

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability we introduce EgoExoLearn a large-scale dataset that emulates the human demonstration following process in which ind…

2024

Exploring DCN-like architecture for fast image generation with arbitrary resolution

NeurIPS 2024poster

Arbitrary-resolution image generation still remains a challenging task in AIGC, as it requires handling varying resolutions and aspect ratios while maintaining high visual quality. Existing transformer-based diffusion methods suffer from quadratic computation cost and limited resolution extrapolatio…

Cited by 0SourcePDFScholar
2024

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

ICLR 2024spotlight

This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. InternVid contains over 7 million videos lasting nearly 760K hours, yielding 234M video clips accomp…

2024

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

CVPR 2024highlight

With the rapid development of Multi-modal Large Language Models (MLLMs) a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However most benchmarks predominantly assess spatial understanding in the static image tasks while overlooking t…

2024

Sparse Global Matching for Video Frame Interpolation with Large Motion

CVPR 2024poster

Large motion poses a critical challenge in Video Frame Interpolation (VFI) task. Existing methods are often constrained by limited receptive fields resulting in sub-optimal performance when handling scenarios with large motion. In this paper we introduce a new pipeline for VFI which can effectively…

Cited by 14SourcePDFScholar
2024

SparseFormer: Sparse Visual Recognition via Limited Latent Tokens

ICLR 2024poster

Human visual recognition is a sparse process, where only a few salient visual cues are attended to rather than every detail being traversed uniformly. However, most current vision networks follow a dense paradigm, processing every single visual unit (such as pixels or patches) in a uniform manner. I…

2024

SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos

CVPR 2024poster

Video-based visual relation detection tasks such as video scene graph generation play important roles in fine-grained video understanding. However current video visual relation detection datasets have two main limitations that hinder the progress of research in this area. First they do not explore c…

2024

VBench: Comprehensive Benchmark Suite for Video Generative Models

CVPR 2024highlight

Video generation has witnessed significant advancements yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should pro…

2024

VFIMamba: Video Frame Interpolation with State Space Models

NeurIPS 2024poster

Inter-frame modeling is pivotal in generating intermediate frames for video frame interpolation (VFI). Current approaches predominantly rely on convolution or attention-based models, which often either lack sufficient receptive fields or entail significant computational overheads. Recently, Selectiv…

2023

CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D Datasets

AAAI 2023technical

Current RGB-D scene recognition approaches often train two standalone backbones for RGB and depth modalities with the same Places or ImageNet pre-training. However, the pre-trained depth network is still biased by RGB-based models which may result in a suboptimal solution. In this paper, we present…

2023

Efficient Video Action Detection with Token Dropout and Context Refinement

ICCV 2023poster

Streaming video clips with large-scale video tokens impede vision transformers (ViTs) for efficient recognition, especially in video action detection where sufficient spatiotemporal representations are required for precise actor identification. In this work, we propose an end-to-end framework for ef…

Cited by 26PDFcodeScholar
2023

Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation

CVPR 2023poster

Effectively extracting inter-frame motion and appearance information is important for video frame interpolation (VFI). Previous works either extract both types of information in a mixed way or devise separate modules for each type of information, which lead to representation ambiguity and low effici…

2023

Filter-Recovery Network for Multi-Speaker Audio-Visual Speech Separation

ICLR 2023poster

In this paper, we systematically study the audio-visual speech separation task in a multi-speaker scenario. Given the facial information of each speaker, the goal of this task is to separate the corresponding speech from the mixed speech. The existing works are designed for speech separation in a co…

Cited by 5SourcePDFScholar
2023

LinK: Linear Kernel for LiDAR-Based 3D Perception

CVPR 2023poster

Extending the success of 2D Large Kernel to 3D perception is challenging due to: 1. the cubically-increasing overhead in processing 3D data; 2. the optimization difficulties from data scarcity and sparsity. Previous work has taken the first step to scale up the kernel size from 3x3x3 to 7x7x7 by int…

2023

MGMAE: Motion Guided Masking for Video Masked Autoencoding

ICCV 2023poster

Masked autoencoding has shown excellent performance on self-supervised video representation learning. Temporal redundancy has led to a high masking ratio and customized masking strategy in VideoMAE. In this paper, we aim to further improve the performance of video masked autoencoding by introducing…

Cited by 39PDFcodeScholar
2023

Memory-and-Anticipation Transformer for Online Action Understanding

ICCV 2023poster

Most existing forecasting systems are memory-based methods, which attempt to mimic human forecasting ability by employing various memory mechanisms and have progressed in temporal modeling for memory dependency. Nevertheless, an obvious weakness of this paradigm is that it can only model limited his…

Cited by 43PDFcodeScholar
2023

PDPP:Projected Diffusion for Procedure Planning in Instructional Videos

CVPR 2023highlight

In this paper, we study the problem of procedure planning in instructional videos, which aims to make goal-directed plans given the current visual observations in unstructured real-life videos. Previous works cast this problem as a sequence planning problem and leverage either heavy intermediate vis…

2023

SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera Videos

ICCV 2023poster

Camera-based 3D object detection in BEV (Bird's Eye View) space has drawn great attention over the past few years. Dense detectors typically follow a two-stage pipeline by first constructing a dense BEV feature and then performing object detection in BEV space, which suffers from complex view transf…

Cited by 127PDFcodeScholar
2023

SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes

ICCV 2023poster

Multi-object tracking (MOT) in sports scenes plays a critical role in gathering players statistics, supporting further applications, such as automatic tactical analysis. Yet existing MOT benchmarks cast little attention on this domain. In this work, we present a new large-scale multi-object tracking…

Cited by 103PDFcodeScholar
2023

StageInteractor: Query-based Object Detector with Cross-stage Interaction

ICCV 2023poster

Previous object detectors make predictions based on dense grid points or numerous preset anchors. Most of these detectors are trained with one-to-many label assignment strategies. On the contrary, recent query-based object detectors are based a sparse set of learnable queries refined by a series of…

Cited by 12PDFcodeScholar
2023

UniFormerV2: Unlocking the Potential of Image ViTs for Video Understanding

ICCV 2023poster

The prolific performances of Vision Transformers (ViTs) in image tasks have prompted research into adapting the image ViTs for video tasks. However, the substantial gap between image and video impedes the spatiotemporal learning of these image-pretrained models. Though video-specialized models like…

Cited by 58PDFcodeScholar
2023

Unmasked Teacher: Towards Training-Efficient Video Foundation Models

ICCV 2023oral

Video Foundation Models (VFMs) have received limited exploration due to high computational costs and data scarcity. Previous VFMs rely on Image Foundation Models (IFMs), which face challenges in transferring to the video domain. Although VideoMAE has trained a robust ViT from limited data, its low-l…

Cited by 189PDFcodeScholar
2023

VideoMAE V2: Scaling Video Masked Autoencoders With Dual Masking

CVPR 2023poster

Scale is the primary factor for building a powerful foundation model that could well generalize to a variety of downstream tasks. However, it is still challenging to train video foundation models with billions of parameters. This paper shows that video masked autoencoder (VideoMAE) is a scalable and…

2022

Cross-Architecture Self-Supervised Video Representation Learning

CVPR 2022poster

In this paper, we present a new cross-architecture contrastive learning (CACL) framework for self-supervised video representation learning. CACL consists of a 3D CNN and a video transformer which are used in parallel to generate diverse positive pairs for contrastive learning. This allows the model…

Cited by 31PDFcodeScholar
2022

DCAN: Improving Temporal Action Detection via Dual Context Aggregation

AAAI 2022technical

Temporal action detection aims to locate the boundaries of action in the video. The current method based on boundary matching enumerates and calculates all possible boundary matchings to generate proposals. However, these methods neglect the long-range context aggregation in boundary prediction. At…

2022

Joint-Modal Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

ECCV 2022poster

"This paper focuses on the weakly-supervised audio-visual video parsing task, which aims to recognize all events belonging to each modality and localize their temporal boundaries. This task is challenging because only overall labels indicating the video events are provided for training. However, an…

2022

Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding

AAAI 2022technical

Temporal grounding aims to localize a video moment which is semantically aligned with a given natural language query. Existing methods typically apply a detection or regression pipeline on the fused representation with the research focus on designing complicated prediction heads or fusion strategies…

2022

OCSampler: Compressing Videos to One Clip With Single-Step Sampling

CVPR 2022poster

Videos incorporate rich semantics as well as redundant information. Seeking a compact yet effective video representation, e.g., sample informative frames from the entire video, is critical to efficient video recognition. There have been works that formulate frame sampling as a sequential decision ta…

Cited by 33PDFcodeScholar
2022

PointTAD: Multi-Label Temporal Action Detection with Learnable Query Points

NeurIPS 2022accept

Traditional temporal action detection (TAD) usually handles untrimmed videos with small number of action instances from a single label (e.g., ActivityNet, THUMOS). However, this setting might be unrealistic as different classes of actions often co-occur in practice. In this paper, we focus on the ta…

2022

Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary Detection

CVPR 2022poster

Generic event boundary detection is an important yet challenging task in video understanding, which aims at detecting the moments where humans naturally perceive event boundaries. The main challenge of this task is perceiving various temporal variations of diverse event boundaries. To this end, this…

Cited by 20PDFcodeScholar
2022

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

NeurIPS 2022accept

Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets. In this paper, we show that video masked autoencoders (VideoMAE) are data-efficient learners for self-supervised video pre-training (SSVP). We are inspired…

2021

MGSampler: An Explainable Sampling Strategy for Video Action Recognition

ICCV 2021poster

Frame sampling is a fundamental problem in video action recognition due to the essential redundancy in time and limited computation resources. The existing sampling strategy often employs a fixed frame selection and lacks the flexibility to deal with complex variations in videos. In this paper, we p…

Cited by 91PDFcodeScholar
2021

MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions

ICCV 2021poster

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions. This paper aims to present a new multi-person dataset of spat…

Cited by 124PDFcodeScholar
2021

PyMAF: 3D Human Pose and Shape Regression With Pyramidal Mesh Alignment Feedback Loop

ICCV 2021poster

Regression-based methods have recently shown promising results in reconstructing human meshes from monocular images. By directly mapping raw pixels to model parameters, these methods can produce parametric models in a feed-forward manner via neural networks. However, minor deviation in parameters ma…

Cited by 386PDFcodeScholar
2021

Target Adaptive Context Aggregation for Video Scene Graph Generation

ICCV 2021poster

This paper deals with a challenging task of video scene graph generation (VidSGG), which could serve as a structured video representation for high-level understanding tasks. We present a new detect-to-track paradigm for this task by decoupling the context modeling for relation prediction from the co…

Cited by 80PDFcodeScholar
2020

Boundary-Aware Cascade Networks for Temporal Action Segmentation

ECCV 2020poster

Identifying human action segments in an untrimmed video is still challenging due to boundary ambiguity and over-segmentation issues. To address these problems, we present a new boundary-aware cascade network by introducing two novel components. First, we devise a new cascading paradigm, called Stage…

2020

Context-Aware RCNN: A Baseline for Action Detection in Videos

ECCV 2020poster

Video action detection approaches usually conduct actor-centric action recognition over RoI-pooled features following the standard pipeline of Faster-RCNN. In this work, we first empirically find the recognition accuracy is highly correlated with the bounding box size of an actor, and thus higher re…

2020

SketchyCOCO: Image Generation From Freehand Scene Sketches

CVPR 2020oral

We introduce the first method for automatic image generation from scene-level freehand sketches. Our model allows for controllable image generation by specifying the synthesis goal via freehand sketches. The key contribution is an attribute vector bridged Generative Adversarial Network called EdgeGA…

Cited by 151PDFScholar
2020

TEA: Temporal Excitation and Aggregation for Action Recognition

CVPR 2020poster

Temporal modeling is key for action recognition in videos. It normally considers both short-range motions and long-range aggregations. In this paper, we propose a Temporal Excitation and Aggregation (TEA) block, including a motion excitation (ME) module and a multiple temporal aggregation (MTA) modu…

Cited by 638PDFScholar
2020

V4D: 4D Convolutional Neural Networks for Video-level Representation Learning

ICLR 2020poster

Most existing 3D CNN structures for video representation learning are clip-based methods, and do not consider video-level temporal evolution of spatio-temporal features. In this paper, we propose Video-level 4D Convolutional Neural Networks, namely V4D, to model the evolution of long-range spatio-te…

Cited by 123SourceScholar
2019

Learning Actor Relation Graphs for Group Activity Recognition

CVPR 2019poster

Modeling relation between actors is important for recognizing group activity in a multi-person scene. This paper aims at learning discriminative relation between actors efficiently using deep models. To this end, we propose to build a flexible and efficient \rm Actor Relation Graph (ARG) to simult…

Cited by 333PDFcodeScholar
2019

Translate-to-Recognize Networks for RGB-D Scene Recognition

CVPR 2019poster

Cross-modal transfer is helpful to enhance modality-specific discriminative power for scene recognition. To this end, this paper presents a unified framework to integrate the tasks of cross-modal translation and modality-specific recognition, termed as Translate-to-Recognize Network TRecgNet. Specif…

Cited by 64PDFcodeScholar
2017

Temporal Action Detection With Structured Segment Networks

ICCV 2017poster

Detecting actions in untrimmed videos is an important yet challenging task. In this paper, we present the structured segment network (SSN), a novel framework which models the temporal structure of each action instance via a structured temporal pyramid. On top of the pyramid, we further introduce a d…

Cited by 1154PDFcodeScholar
2017

Thin-Slicing Network: A Deep Structured Model for Pose Estimation in Videos

CVPR 2017oral

Deep ConvNets have been shown to be effective for the task of human pose estimation from single images. However, several challenging issues arise in the video-based case such as self-occlusion, motion blur, and uncommon poses with few or no examples in the training data. Temporal information can pro…

Cited by 159PDFScholar
2017

UntrimmedNets for Weakly Supervised Action Recognition and Detection

CVPR 2017poster

Current action recognition methods heavily rely on trimmed videos for model training. However, it is expensive and time-consuming to acquire a large-scale trimmed video dataset. This paper presents a new weakly supervised architecture, called UntrimmedNet, which is able to directly learn action reco…

Cited by 651PDFcodeScholar
2016

Real-Time Action Recognition With Enhanced Motion Vector CNNs

CVPR 2016poster

The deep two-stream architecture exhibited excellent performance on video based action recognition. The most computationally expensive step in this approach comes from the calculation of optical flow which prevents it to be real-time. This paper accelerates this architecture by replacing optical flo…

Cited by 546PDFcodeScholar