← Search

Ruyang Liu

10 accepted papers

2026

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

ICLR 2026poster

The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most existing video LLMs cannot handle hour-long videos, while methods custom for long vi…

Cited by 0SourcecodeScholar
2026

Video Spatial Reasoning with Object-Centric 3D Rollout

AAAI 2026technical

Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning—the ability to comprehend object locations, orientations, and inter-object relationships in dynamic 3D scenes—remains

Cited by 0SourcePDFScholar
2025

Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow

ICCV 2025poster

Long-form video understanding has always been a challenging problem due to the significant redundancy in both temporal and spatial contents. This challenge is further exacerbated by the limited context length of Multimodal Large Language Models (MLLMs). To address this issue, many previous works hav…

Cited by 0SourcePDFScholar
2025

MUSE: Mamba Is Efficient Multi-scale Learner for Text-video Retrieval

AAAI 2025technical

Text-Video Retrieval (TVR) aims to align and associate relevant video content with corresponding natural language queries. Most existing TVR methods are based on large-scale pre-trained vision-language models (e.g., CLIP). However, due to CLIP's inherent plain structure, few TVR methods explore the…

2024

BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

CVPR 2024poster

The recent progress in Large Language Models (LLM) has spurred various advancements in image-language conversation agents while how to build a proficient video-based dialogue system is still under exploration. Considering the extensive scale of LLM and visual backbone minimal GPU memory is left for…

2024

RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter

ACL 2024findings

Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most of the state-of-the-art TVR methods learn image-to-video transfer learning based on the large-scale pre-trained vision-language models (e.g., CLIP). However, fully fine-tuning these pre-train…

Cited by 17SourcePDFScholar
2024

ST-LLM: Large Language Models Are Effective Temporal Learners

ECCV 2024poster

"Large Language Models (LLMs) have showcased impressive capabilities in text comprehension and generation, prompting research efforts towards video LLMs to facilitate human-AI interaction at the video level. However, how to effectively encode and understand videos in video-based dialogue systems rem…

2023

Causality Compensated Attention for Contextual Biased Visual Recognition

ICLR 2023poster

Visual attention does not always capture the essential object representation desired for robust predictions. Attention modules tend to underline not only the target object but also the common co-occurring context that the module thinks helpful in the training. The problem is rooted in the confoundin…

Cited by 23SourcePDFScholar
2023

Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge Transferring

CVPR 2023poster

Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual representation learning in the video domain. In this paper, based on the CLIP model…

2022

Contextual Debiasing for Visual Recognition With Causal Mechanisms

CVPR 2022poster

As a common problem in the visual world, contextual bias means the recognition may depend on the co-occurrence context rather than the objects themselves, which is even more severe in multi-label tasks due to multiple targets and the absence of location. Although some studies have focused on tacklin…

Cited by 46PDFcodeScholar