← Search

Quan Kong

18 accepted papers

2026

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

CVPR 2026

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely provide the fine-grained, grounded evidence needed to rigorously

Cited by 0SourceScholar
2026

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

CVPR 2026

Current vision-language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only supervision. We introduce InstAP, an Instance-Aware Pre-training framework that jointly optimizes global vision-text alignment and fine-grained, inst

Cited by 0SourceScholar
2026

MoVie: Broaden Your Views with Human Motion for Action Detection

CVPR 2026

Human action detection in videos requires both semantic recognition and accurate modeling of motion. While recent video foundation models have advanced visual semantics, they still struggle to capture complex and compositional actions due to the limited representation ability of motion. Human skelet

Cited by 0SourceScholar
2026

ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding

CVPR 2026

Although current Video-LLMs achieve impressive performance in video understanding tasks, their autoregressive decoding efficiency remains constrained by the massive number of video tokens. Visual token pruning can partially ease this bottleneck, yet existing approaches still suffer from information

Cited by 0SourcecodeScholar
2026

SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism

ICLR 2026poster

Speculative decoding (SD) has emerged as a promising technique to accelerate LLM inference by employing a small, efficient draft model to propose draft tokens in advance, and subsequently validating them in parallel with the large target model. However, the existing SD methods still remain fundament…

Cited by 0SourcecodeScholar
2026

TrajTok: Learning Trajectory Tokens Enhances Video Understanding

CVPR 2026

Tokenization in video models, typically through patchification, generates an excessive and redundant number of tokens. This severely limits video efficiency and scalability. While the recent trajectory-based tokenizers offer a promising solution by decoupling video duration from token count, they re

Cited by 0SourcecodeScholar
2025

GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding

CVPR 2025poster

We propose a novel 3D gaze estimation approach that learns spatial relationships between the subject and objects in the scene, and outputs 3D gaze direction. Our method targets unconstrained settings, including cases where close-up views of the subject's eyes are unavailable, such as when the subjec…

Cited by 0SourcePDFScholar
2025

Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection

CVPR 2025highlight

Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such…

2025

Mixture of Experts Guided by Gaussian Splatters Matters: A new Approach to Weakly-Supervised Video Anomaly Detection

ICCV 2025poster

Video Anomaly Detection (VAD) is a challenging task due to the variability of anomalous events and the limited availability of labeled data. Under the Weakly-Supervised VAD (WSVAD) paradigm, only video-level labels are provided during training, while predictions are made at the frame level. Although…

2025

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory

ICCV 2025poster

Effective video tokenization is critical for scaling transformer models for long videos. Current approaches tokenize videos using space-time patches, leading to excessive tokens and computational inefficiencies. The best token reduction strategies degrade performance and barely reduce the number of…

Cited by 0SourcePDFScholar
2024

Reprojection Errors as Prompts for Efficient Scene Coordinate Regression

ECCV 2024poster

"Scene coordinate regression (SCR) methods have emerged as a promising area of research due to their potential for accurate visual localization. However, many existing SCR approaches train on samples from all image regions, including dynamic objects and texture-less areas. Utilizing these areas for…

Cited by 1SourcePDFScholar
2023

DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-To-Fine Contrastive Ranking

CVPR 2023poster

Understanding dense action in videos is a fundamental challenge towards the generalization of vision models. Several works show that compositionality is key to achieving generalization by combining known primitive elements, especially for handling novel composited structures. Compositional temporal…

Cited by 15SourcePDFScholar
2023

LAC - Latent Action Composition for Skeleton-based Action Segmentation

ICCV 2023poster

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their perfo…

Cited by 14PDFScholar
2023

Self-Supervised Video Representation Learning via Latent Time Navigation

AAAI 2023technical

Self-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent information related to temporal relationships, rendering actions such as `enter' and `leav…

Cited by 11SourcePDFScholar
2020

Anticipating the Start of User Interaction for Service Robot in the Wild

ICRA 2020poster

A service robot is expected to provide proactive service for visitors who require its help. In contrast to passive service, e.g., providing service only after being spoken to, proactive service initiates an interaction at an early stage, e.g., talking to potential visitors who need the robot’s help…

Cited by 9SourceScholar
2020

Cycle-Contrast for Self-Supervised Video Representation Learning

NeurIPS 2020poster

We present Cycle-Contrastive Learning (CCL), a novel self-supervised method for learning video representation. Following a nature that there is a belong and inclusion relation of video and its frames, CCL is designed to find correspondences across frames and videos considering the contrastive repres…

Cited by 54SourcePDFScholar
2019

MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding

ICCV 2019poster

Unlike vision modalities, body-worn sensors or passive sensing can avoid the failure of action understanding in vision related challenges, e.g. occlusion and appearance variation. However, a standard large-scale dataset does not exist, in which different types of modalities across vision and sensors…

Cited by 124PDFScholar