← Search

Shiyu Hu

13 accepted papers

2026

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking

IJCAI 2026

Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic supervision. This mismatch is particularly acute in highly homogeneous scenarios that require fine-grained discrimination over complex compositional se

Cited by 0Scholar
2026

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos

AAAI 2026technical

Recent advances in large language models (LLMs) have improved reasoning in text and image domains, yet achieving robust video reasoning remains a significant challenge. Existing video benchmarks mainly assess shallow understanding and reasoning and allow models to exploit global context, failing to

Cited by 0SourcePDFScholar
2026

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation

ICLR 2026poster

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video generation tasks is not only to extend video duration but also…

Cited by 0SourceScholar
2026

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains

AAAI 2026technical

Large language models (LLMs) increasingly rely on reinforcement learning (RL) to enhance their reasoning capabilities through feedback. A critical challenge is verifying the consistency of model-generated responses and reference answers, since these responses are often lengthy, diverse, and nuanced.

Cited by 0SourcePDFScholar
2025

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

ICCV 2025poster

Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, i…

2025

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

ICML 2025poster

Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the mod…

2025

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

ICASSP 2025accepted

Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalit…

Cited by 0SourceScholar
2024

Beyond Accuracy: Tracking more like Human via Visual Search

NeurIPS 2024poster

Human visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and…

2024

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts

NeurIPS 2024poster

Vision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed…

Cited by 5SourcePDFScholar
2023

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal Relationship

NeurIPS 2023poster

Tracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented b…

Cited by 13SourcePDFScholar