← Search

Rongjie Li

7 accepted papers

2025

GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

NeurIPS 2025poster

While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction,…

Cited by 0SourceScholar
2025

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

AAAI 2025technical

Open-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representations with open-vocabulary textual representations. This enables the identification of novel visual relationships, making it applicable to real-world scena…

Cited by 1SourcePDFScholar
2024

From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models

CVPR 2024poster

Scene graph generation (SGG) aims to parse a visual scene into an intermediate graph representation for downstream reasoning tasks. Despite recent advancements existing methods struggle to generate scene graphs with novel visual relation concepts. To address this challenge we introduce a new open-vo…

2024

Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning

CVPR 2024poster

Generative vision-language models (VLMs) have shown impressive performance in zero-shot vision-language tasks like image captioning and visual question answering.However improving their zero-shot reasoning typically requires second-stage instruction tuning which relies heavily on human-labeled or la…

2021

Bipartite Graph Network With Adaptive Message Passing for Unbiased Scene Graph Generation

CVPR 2021poster

Scene graph generation is an important visual understanding task with a broad range of vision applications. Despite recent tremendous progress, it remains challenging due to the intrinsic long-tailed class distribution and large intra-class variation. To address these issues, we introduce a novel co…

Cited by 281PDFcodeScholar
2019

Pose-Aware Multi-Level Feature Network for Human Object Interaction Detection

ICCV 2019oral

Reasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large variations in human-object configurations, multiple co-occurring relation instances and subtle visual difference between rel…

Cited by 276PDFcodeScholar