← Search

Qingchuan Ma

4 accepted papers

2026

A²RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation

ICML 2026poster

Abstract reasoning ability reflects the intelligence and generalization capacity of LLMs to extract and apply abstract rules. However, accurately measuring this ability remains challenging: existing benchmarks either rely on expensive manual annotation, limiting their scale, or risk measuring memori…

Cited by 1SourceScholar
2026

Surgeon Supervised Autonomous Surgical System for Oral and Maxillofacial Surgery (I)

ICRA 2026poster

Oral and maxillofacial surgery (OMS) imposes an increasing workload on even the most experienced surgeons due to long operation time, high skill requirements, limited observation field, constrained workspace, and fast-growing patient population. Robot-assisted OMS is particularly challenging, requir…

Cited by 0Scholar
2025

Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective

ICML 2025poster

In this paper, we aim to establish a simple, effective, and theoretically grounded benchmark for rigorously probing abstract reasoning in Large Language Models (LLMs). To achieve this, we first develop a mathematic framework that defines abstract reasoning as the ability to: (i) extract essential pa…

2025

CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset

CVPR 2025poster

X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence that can significantly reduce diagnostic burdens and patient wait times. Despite significant progress, we believe that the task has reached a bottleneck due to the limited benchmark datasets and the existi…