← Search

Kaiwen Shi

4 accepted papers

2026

Can LLMs Move Beyond Short Exchanges to Realistic Therapy Conversations?

ICLR 2026poster

Recent incidents have revealed that large language models (LLMs) deployed in mental health contexts can generate unsafe guidance, including reports of chatbots encouraging self-harm. Such risks highlight the urgent need for rigorous, clinically valid evaluation before integration into care. However,…

Cited by 0SourceScholar
2026

DRIFT-BENCH: Diagnosing CoopeRative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction

ICML 2026poster

As Large Language Models transition to autonomous agents, user inputs frequently violate cooperative assumptions (e.g., implicit intent, missing parameters, false presuppositions, or ambiguous expressions), creating execution risks that text-only evaluations do not capture. Existing benchmarks typic…

Cited by 0SourceScholar
2026

FlowGen: Synthesizing Diverse Flowcharts to Enhance and Benchmark MLLM Reasoning

ICLR 2026poster

Flowcharts are widely used to represent processes and relationships through intuitive visual representations. However, accurately interpreting these diagrams remains challenging due to their structural complexity and high visual diversity. Existing flowchart datasets often lack fine-grained control…

Cited by 0SourcecodeScholar
2025

Linear Spherical Sliced Optimal Transport: A Fast Metric for Comparing Spherical Data

ICLR 2025spotlight

Efficient comparison of spherical probability distributions becomes important in fields such as computer vision, geosciences, and medicine. Sliced optimal transport distances, such as spherical and stereographic spherical sliced Wasserstein distances, have recently been developed to address this nee…

Cited by 0SourcePDFScholar