← Search

Ruiyang Zhang

4 accepted papers

2026

SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models

ICLR 2026poster

Despite the empirical success of extensive, step-by-step reasoning in large multimodal models, long reasoning processes inevitably incur substantial computational overhead, i.e., in terms of higher token costs and increased response time, which undermines inference efficiency. In contrast, humans of…

Cited by 0SourcecodeScholar
2025

Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object Detection

ICCV 2025poster

Unsupervised 3D object detection aims to identify objects of interest from unlabeled raw data, such as LiDAR points. Recent approaches usually adopt pseudo 3D bounding boxes (3D bboxes) from clustering algorithm to initialize the model training. However, pseudo bboxes inevitably contain noise, and s…

2025

MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval

EMNLP 2025

Document Understanding is a foundational AI capability with broad applications, and Document Question Answering (DocQA) is a key evaluation task. Traditional methods convert the document into text for processing by Large Language Models (LLMs), but this process strips away critical multi-modal infor

2023

A Step Towards Conditional Autonomy - Robotic Appendectomy

RA-L 2023

In recent years, Robot-Assisted Minimally Invasive Surgery (RAMIS) has been widely adopted worldwide due to its high precision, improved ergonomics and intuitive control. With the advances in artificial intelligence and surgical robot technologies, it is anticipated that the cognitive load on the su

Cited by 13SourceScholar