← Search

Dezhao Luo

5 accepted papers

2026

ViMo: A Generative Visual GUI World Model for App Agents

ICLR 2026poster

App agents, which autonomously operate mobile Apps through GUIs, have gained significant interest in real-world applications. Yet, they often struggle with long-horizon planning, failing to find the optimal actions for complex tasks with longer steps. To address this, world models are used to predic…

Cited by 0SourceScholar
2025

Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval

AAAI 2025technical

Video moment retrieval (VMR) aims to locate the most likely video moment(s) corresponding to a text query in untrimmed videos. Training of existing methods is limited by the lack of diverse and generalisable VMR datasets, hindering their ability to generalise moment-text associations to queries cont…

Cited by 0SourcePDFScholar
2025

The Role of Video Generation in Enhancing Data-Limited Action Understanding

IJCAI 2025

Video action understanding tasks in real-world scenarios often suffer from data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that leverages a text-to-video diffusion transformer to generate annotated data fo

Cited by 0SourcePDFScholar
2023

Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training

CVPR 2023poster

The correlation between the vision and text is essential for video moment retrieval (VMR), however, existing methods heavily rely on separate pre-training feature extractors for visual and textual understanding. Without sufficient temporal boundary annotations, it is non-trivial to learn universal v…

Cited by 40SourcePDFScholar
2020

Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation Learning

CVPR 2020poster

In self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capabilities of learned models. In this paper, we propose a novel self-supervised method, referred to as video Playback Rate…

Cited by 210PDFcodeScholar