← Search

Shengyuan Wang

8 accepted papers

2025

CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluation

EMNLP 2025

Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptimal evaluation. We introduce a **Cl**inically grounded tabular framework with **E**xpert-curated labels and **A**ttribute

2025

Control and Localization of Magnetic Nanorobot Swarms in Human-Sized Vascular Phantom

IROS 2025

Magnetically controlled micro-nano robots hold revolutionary significance in the clinical targeted treatment of brain tumors. Imaging and tracking miniature robots can provide feedback for precise magnetic field control. The cooperation among micro-nano robots, magnetic field control system, and ima

Cited by 0SourceScholar
2025

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning

EMNLP 2025

Large language models (LLMs) possess extensive world knowledge, including geospatial knowledge, which has been successfully applied to various geospatial tasks such as mobility prediction and social indicator prediction. However, LLMs often generate inaccurate geospatial knowledge, leading to geospa

2025

Multimodal Upstream Motion of Magnetically Controlled Micro/Nano Robots in High-Viscosity Fluids

IROS 2025

The efficacy of targeted cancer drug therapy is significantly compromised by imprecise drug delivery mechanisms. Micro/nano robots (MNRs), characterized by their controllable motion, present a promising solution to this challenge. However, the non-Newtonian nature of blood, with its high viscosity a

Cited by 0SourceScholar
2025

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence

ICCV 2025poster

Urban research involves a wide range of scenarios and tasks that require the understanding of multi-modal data, such as structured geospatial data, trajectory data, satellite image data, and street view image data. Current methods often focus on specific data types and lack a unified framework in ur…

2024

AlignBench: Benchmarking Chinese Alignment of Large Language Models

ACL 2024long

Alignment has become a critical step for instruction-tuned Large Language Models (LLMs) to become helpful assistants. However, effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluat…

2024

CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

ACL 2024long

Since the natural language processing (NLP) community started to make large language models (LLMs) act as a critic to evaluate the quality of generated texts, most of the existing works train a critique generation model on the evaluation data labeled by GPT-4’s direct prompting. We observe that thes…

2024

Dung Beetle Optimizer-based High-precision Localization for Magnetic-Controlled Capsule Robot

IROS 2024poster

As a medical microrobot, magnetic-controlled capsule robots (MCRs) are pivotal in internal diagnostics and therapeutic interventions. Achieving high-precision localization of MCRs is essential for the successful execution of medical procedures. This paper introduces a novel Dung Beetle Optimizer (DB…

Cited by 0SourceScholar