← Search

Dingxin Cheng

2 accepted papers

2025

VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

AAAI 2025technical

Video Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing. Unlike traditional task-specific models, Video Large Language Models (video LLMs) can handle multiple…

2024

Long Term Memory-Enhanced Via Causal Reasoning for Text-To-Video Retrieval

ICASSP 2024accepted

The T2VR task aims to retrieve videos that are semantically relevant to the given query text in a large number of unlabeled videos. Most of the existing methods adopt a representation encoding strategy that can only focus on limited contextual information, and lack the ability to focus on the long m…

Cited by 0SourceScholar