← Search

Yang Du

7 accepted papers

2025

Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?

ICLR 2025poster

Egocentric video-language pretraining is a crucial step in advancing the understanding of hand-object interactions in first-person scenarios. Despite successes on existing testbeds, we find that current EgoVLMs can be easily misled by simple modifications, such as changing the verbs or nouns in inte…

2025

MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning

COLING 2025main

Extensive research has been conducted to explore the capability of Large Language Models (LLMs) for table reasoning and has significantly improved the performance on existing benchmarks. However, tables and user questions in real-world applications are more complex and diverse, presenting an unignor…

2025

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

NeurIPS 2025poster

Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability…

Cited by 0SourcecodeScholar
2025

VC4VG: Optimizing Video Captions for Text-to-Video Generation

EMNLP 2025

Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video captions specifically for T2V training remain underexplored. In thi

2018

Interaction-aware Spatio-temporal Pyramid Attention Networks for Action Classification

ECCV 2018poster

Local features at neighboring spatial positions in feature maps have high correlation since their receptive fields are often overlapped. Self-attention usually uses the weighted sum (or other functions) with internal elements of each local feature to obtain its weight score, which ignores interactio…

Cited by 118SourcePDFScholar
2017

Spatio-Temporal Self-Organizing Map Deep Network for Dynamic Object Detection From Videos

CVPR 2017poster

In dynamic object detection, it is challenging to construct an effective model to sufficiently characterize the spatial-temporal properties of the background. This paper proposes a new Spatio-Temporal Self-Organizing Map (STSOM) deep network to detect dynamic objects in complex scenarios. The propos…

Cited by 18PDFScholar