← Search

Zhichao Yang

17 accepted papers

2026

Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks

CVPR 2026

Image aesthetic assessment (IAA) has extensive applications in content creation, album management, and recommendation systems, etc. In such applications, it is commonly needed to pick out the most aesthetically pleasing image from a series of images with subtle aesthetic variations, a topic we refer

Cited by 0SourcecodeScholar
2026

Fine-grained Image Quality Assessment for Perceptual Image Restoration

AAAI 2026technical

Recent years have witnessed remarkable achievements in perceptual image restoration (IR), creating an urgent demand for accurate image quality assessment (IQA), which is essential for both performance comparison and algorithm optimization. Unfortunately, the existing IQA metrics exhibit inherent wea

Cited by 0SourcePDFScholar
2026

LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations

AAAI 2026technical

The increasing popularity of long Text-to-Image (T2I) generation has created an urgent need for automatic and interpretable models that can evaluate the image-text alignment in long prompt scenarios. However, the existing T2I alignment benchmarks predominantly focus on short prompt scenarios and onl

Cited by 0SourcePDFScholar
2026

Medical thinking with multiple images

ICLR 2026poster

Large language models and vision-language models score high on many medical QA benchmarks; however, real-world clinical reasoning remains challenging because cases often involve multiple images and require cross-view fusion. We present MedThinkVQA, a benchmark that asks models to think with multiple…

Cited by 6SourcecodeScholar
2026

PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning

AAAI 2026technical

Inspired by the dual-process theory of human cognition from Thinking, Fast and Slow, we introduce PRIME (Planning and Retrieval-Integrated Memory for Enhanced Reasoning), a multi-agent reasoning framework that dynamically integrates System 1 (fast, intuitive thinking) and System 2 (slow, deliberate

Cited by 0SourcePDFScholar
2026

TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning

AAAI 2026technical

Livestreaming has become increasingly prevalent in modern visual communication, where automatic camera quality tuning is essential for delivering superior user Quality of Experience (QoE). Such tuning requires accurate blind image quality assessment (BIQA) to guide parameter optimization decisions.

Cited by 0SourcePDFScholar
2025

MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback

NAACL 2025long

Automatic question generation (QG) is essential for AI and NLP, particularly in intelligent tutoring, dialogue systems, and fact verification. Generating multiple-choice questions (MCQG) for professional exams, like the United States Medical Licensing Examination (USMLE), is particularly challenging…

2025

Online Anti-Swing Trajectory Refinement for Variable-Length Cable-Suspended Aerial Transportation Robot

IROS 2025

Aerial robots have demonstrated significant potential in suspended cargo transportation, especially in industries such as logistics and food delivery. Due to the underactuated and nonlinear dynamics of the cable-suspended system, directly tracking a given trajectory with a multicopter without modify

Cited by 0SourceScholar
2025

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models

ACL 2025long

This work introduces RARE (Retrieval-Augmented Reasoning Enhancement), a versatile extension to the mutual reasoning framework (rStar), aimed at enhancing reasoning accuracy and factual integrity across large language models (LLMs) for complex, knowledge-intensive tasks such as medical and commonsen…

2025

Synth-SBDH: A Synthetic Dataset of Social and Behavioral Determinants of Health for Clinical Text

EMNLP 2025

Social and behavioral determinants of health (SBDH) play a crucial role in health outcomes and are frequently documented in clinical text. Automatically extracting SBDH information from clinical text relies on publicly available good-quality datasets. However, existing SBDH datasets exhibit substant

2024

Large Language Models are In-context Teachers for Knowledge Reasoning

EMNLP 2024finding

In this work, we study in-context teaching(ICT), where a teacher provides in-context example rationales to teach a student to reasonover unseen cases. Human teachers are usually required to craft in-context demonstrations, which are costly and have high variance. We ask whether a large language mode…

Cited by 2SourcePDFScholar
2024

NoteChat: A Dataset of Synthetic Patient-Physician Conversations Conditioned on Clinical Notes

ACL 2024findings

We introduce NoteChat, a novel cooperative multi-agent framework leveraging Large Language Models (LLMs) to generate patient-physician dialogues. NoteChat embodies the principle that an ensemble of role-specific LLMs, through structured role-play and strategic prompting, can perform their assigned r…

2024

README: Bridging Medical Jargon and Lay Understanding for Patient Education through Data-Centric NLP

EMNLP 2024finding

The advancement in healthcare has shifted focus toward patient-centric approaches, particularly in self-care and patient education, facilitated by access to Electronic Health Records (EHR). However, medical jargon in EHRs poses significant challenges in patient comprehension. To address this, we int…

2023

Interpretable Math Word Problem Solution Generation via Step-by-step Planning

ACL 2023long

Solutions to math word problems (MWPs) with step-by-step explanations are valuable, especially in education, to help students better comprehend problem-solving strategies. Most existing approaches only focus on obtaining the final correct answer. A few recent approaches leverage intermediate solutio…

Cited by 12SourcePDFScholar
2023

Multi-Label Few-Shot ICD Coding as Autoregressive Generation with Prompt

AAAI 2023technical

Automatic International Classification of Diseases (ICD) coding aims to assign multiple ICD codes to a medical note with an average of 3,000+ tokens. This task is challenging due to the high-dimensional space of multi-label assignment (155,000+ ICD code candidates) and the long-tail challenge - Many…

2023

Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss Information

ACL 2023long

Visual Word Sense Disambiguation (VWSD) is a task to find the image that most accurately depicts the correct sense of the target word for the given context. Previously, image-text matching models often suffered from recognizing polysemous words. This paper introduces an unsupervised VWSD approach th…

2022

Knowledge Injected Prompt Based Fine-tuning for Multi-label Few-shot ICD Coding

EMNLP 2022finding

Automatic International Classification of Diseases (ICD) coding aims to assign multiple ICD codes to a medical note with average length of 3,000+ tokens. This task is challenging due to a high-dimensional space of multi-label assignment (tens of thousands of ICD codes) and the long-tail challenge: o…