← Search

Jaeyoung Do

15 accepted papers

2026

Medic-AD: Towards Medical Vision-Language Model's Clinical Intelligence

CVPR 2026

Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language Models (VLMs) still lack mechanisms that translate their broad knowledge into clinically actionable outputs. To bridge this gap, we present Medic-AD, a

Cited by 0SourcecodeScholar
2026

RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models

ICLR 2026poster

Large Reasoning Models (LRMs) exhibit strong performance, yet often produce rationales that sound plausible but fail to reflect their true decision process, undermining reliability and trust. We introduce a formal framework for *reasoning faithfulness*, defined by two testable conditions: *stance co…

Cited by 0SourcecodeScholar
2026

VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

ICML 2026oral

Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path, yet three gaps persist– extraction often ignores hierarchic…

Cited by 0SourceScholar
2025

Don’t Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation

NeurIPS 2025poster

Classifier guidance is a widely adopted technique in diffusion language models, used to steer generation toward desired attributes. However, such guidance often introduces instability during the generation process, where token-level updates fluctuate across timesteps. We identify and formally charac…

Cited by 0SourceScholar
2025

MathReader : Text-to-Speech for Mathematical Documents

ICASSP 2025accepted

TTS (Text-to-Speech) document reader from Microsoft, Adobe, Apple, and OpenAI have been serviced worldwide. They provide relatively good TTS results for general plain text, but sometimes skip contents or provide unsatisfactory results for mathematical expressions. This is because most modern academi…

Cited by 0SourceScholar
2025

MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula

AAAI 2025technical

In various academic and professional settings, such as mathematics lectures or research presentations, it is often necessary to convey mathematical expressions orally. However, reading mathematical expressions aloud without accompanying visuals can significantly hinder comprehension, especially for…

2025

SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding

ICML 2025poster

Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a critical challenge to achieving accurate visual understanding. To address this issue, we propose SECOND: Selective and Contrastive Decoding, a novel approac…

2024

Aligning Large Language Models via Fine-grained Supervision

ACL 2024short

Pre-trained large-scale language models (LLMs) excel at producing coherent articles, yet their outputs may be untruthful, toxic, or fail to align with user expectations. Current approaches focus on using reinforcement learning with human feedback (RLHF) to improve model alignment, which works by tra…

Cited by 2SourcePDFScholar
2023

Grounding Counterfactual Explanation of Image Classifiers to Textual Concept Space

CVPR 2023poster

Concept-based explanation aims to provide concise and human-understandable explanations of an image classifier. However, existing concept-based explanation methods typically require a significant amount of manually collected concept-annotated images. This is costly and runs the risk of human biases…

Cited by 11SourcePDFScholar
2023

Large-scale Lifelong Learning of In-context Instructions and How to Tackle It

ACL 2023long

Jointly fine-tuning a Pre-trained Language Model (PLM) on a pre-defined set of tasks with in-context instructions has been proven to improve its generalization performance, allowing us to build a universal language model that can be deployed across task boundaries. In this work, we explore for the f…

Cited by 15SourcePDFScholar
2023

Scalable and Safe Remediation of Defective Actions in Self-Learning Conversational Systems

ACL 2023industry

Off-Policy reinforcement learning has been the driving force for the state-of-the-art conversational AIs leading to more natural human-agent interactions and improving the user satisfaction for goal-oriented agents. However, in large-scale commercial settings, it is often challenging to balance betw…

Cited by 0SourcePDFScholar
2023

Weakly Supervised Referring Image Segmentation with Intra-Chunk and Inter-Chunk Consistency

ICCV 2023poster

Referring image segmentation (RIS) aims to localize the object in an image referred by a natural language expression. Most previous studies learn RIS with a large-scale dataset containing segmentation labels, but they are costly. We present a weakly supervised learning method for RIS that only uses…

Cited by 29PDFScholar