← Search

YinChao Ma

8 accepted papers

2026

Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning

AAAI 2026technical

Video large language models have demonstrated remarkable capabilities in video understanding tasks. However, the redundancy of video tokens introduces significant computational overhead during inference, limiting their practical deployment. Many compression algorithms are proposed to prioritize reta

Cited by 0SourcePDFScholar
2026

Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking

CVPR 2026

3D single object tracking (SOT) in point clouds is essential for real-world 3D perception, yet it remains challenging due to data sparsity and large variations in scale and structure across diverse object categories. Most existing methods rely on a category-specific paradigm that trains separate mod

Cited by 0SourceScholar
2026

Learning Generalized Trackers with Elastic Token Budgets

ICML 2026poster

Visual tracking aims to estimate target states in video sequences, with applications spanning diverse computational requirements. Recent methods optimize trackers using manually pruned image tokens with a fixed budget to reduce computational costs. However, these trackers, once trained, are constrai…

Cited by 0SourceScholar
2026

ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis

ICLR 2026poster

While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasoning remains underdeveloped. This gap stems primarily from a critical data bottleneck: existing datasets lack the challengi…

Cited by 0SourcecodeScholar
2024

Unifying Visual and Vision-Language Tracking via Contrastive Learning

AAAI 2024technical

Single object tracking aims to locate the target object in a video sequence according to the state specified by different modal references, including the initial bounding box (BBOX), natural language (NL), or both (NL+BBOX). Due to the gap between different modalities, most existing trackers are des…

2023

Foreground-Background Distribution Modeling Transformer for Visual Object Tracking

ICCV 2023poster

Visual object tracking is a fundamental research topic with a broad range of applications. Benefiting from the rapid development of Transformer, pure Transformer trackers have achieved great progress. However, the feature learning of these Transformer-based trackers is easily disturbed by complex ba…

Cited by 38PDFScholar
2022

A Keypoint-Based Global Association Network for Lane Detection

CVPR 2022poster

Lane detection is a challenging task that requires predicting complex topology shapes of lane lines and distinguishing different types of lanes simultaneously. Earlier works follow a top-down roadmap to regress predefined anchors into various shapes of lane lines, which lacks enough flexibility to f…

Cited by 154PDFcodeScholar
2021

Interpreting and Boosting Dropout from a Game-Theoretic View

ICLR 2021poster

This paper aims to understand and improve the utility of the dropout operation from the perspective of game-theoretical interactions. We prove that dropout can suppress the strength of interactions between input variables of deep neural networks (DNNs). The theoretical proof is also verified by vari…

Cited by 55SourcePDFScholar