← Search

Kaitao Chen

5 accepted papers

2026

Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification

ICML 2026poster

Extended Chain-of-Thought (CoT) reasoning has significantly bolstered the capabilities of medical large language models (LLMs). However, current models exhibit static computational expenditure, applying lengthy reasoning processes indiscriminately to both simple queries and complex diagnostic cases.…

Cited by 0SourceScholar
2026

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

ICML 2026poster

Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle with medical images, which typically exhibit extremely sparse visual evidence to inform clinical decision-making. We recognize that pruning visual tokens outside…

Cited by 0SourceScholar
2025

Counterfactual Debiasing for Physical Audiovisual Commonsense Reasoning

AAAI 2025technical

Physical commonsense is an essential aspect of human cognition, involving an intuitive understanding of the physical properties and interactions of everyday objects and materials. Though physical commonsense reasoning should inherently be a multisensory task, integrating both video and audio signals…

Cited by 0SourcePDFScholar
2025

RPMIL: Rethinking Uncertainty-Aware Probabilistic Multiple Instance Learning for Whole Slide Pathology Diagnosis

IJCAI 2025

Whole slide images (WSIs) are gigapixel digital scans of traditional pathology slides, offering substantial support for cancer diagnosis. Current multiple instance learning (MIL) methods for WSIs typically extract instance features and aggregate these into a single bag feature for prediction. We obs

Cited by 0SourcePDFScholar
2024

CaMIL: Causal Multiple Instance Learning for Whole Slide Image Classification

AAAI 2024technical

Whole slide image (WSI) classification is a crucial component in automated pathology analysis. Due to the inherent challenges of high-resolution WSIs and the absence of patch-level labels, most of the proposed methods follow the multiple instance learning (MIL) formulation. While MIL has been equipp…

Cited by 13SourcePDFScholar