← Search

Kai Shu

21 accepted papers

2026

Benchmarking LLMs for Political Science: A United Nations Perspective

AAAI 2026technical

Large Language Models (LLMs) have achieved significant advances in natural language processing, yet their potential for high-stake political decision-making remains largely unexplored. This paper addresses the gap by focusing on the application of LLMs to the United Nations (UN) decision-making proc

Cited by 0SourcePDFScholar
2026

CryoLVM: Self-supervised Learning from Cryo-EM Density Maps with Large Vision Models

ICLR 2026poster

Cryo-electron microscopy (cryo-EM) has revolutionized structural biology by enabling near-atomic-level visualization of biomolecular assemblies. However, the exponential growth in cryo-EM data throughput and complexity, coupled with diverse downstream analytical tasks, necessitates unified computati…

Cited by 0SourceScholar
2026

Model Editing as a Double-Edged Sword: Steering Agent Behavior Toward Beneficence or Harm

AAAI 2026technical

Agents based on Large Language Models (LLMs) have demonstrated strong capabilities across a wide range of tasks. However, deploying LLM-based agents in high-stakes domains comes with significant safety and ethical risks. Unethical behavior by these agents can directly result in serious real-world co

Cited by 0SourcePDFScholar
2026

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

ICLR 2026poster

Generative foundation models (GenFMs), such as large language models and text-to-image systems, have demonstrated remarkable capabilities in various downstream applications. As they are increasingly deployed in high-stakes applications, assessing their trustworthiness has become both a critical nece…

Cited by 0SourceScholar
2025

Can Knowledge Editing Really Correct Hallucinations?

ICLR 2025poster

Large Language Models (LLMs) suffer from hallucinations, referring to the non-factual information in generated content, despite their superior capacities across tasks. Meanwhile, knowledge editing has been developed as a new popular paradigm to correct erroneous factual knowledge encoded in LLMs wi…

2025

ConQRet: A New Benchmark for Fine-Grained Automatic Evaluation of Retrieval Augmented Computational Argumentation

NAACL 2025long

Computational argumentation, which involves generating answers or summaries for controversial topics like abortion bans and vaccination, has become increasingly important in today’s polarized environment. Sophisticated LLM capabilities offer the potential to provide nuanced, evidence-based answers t…

2025

From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge

EMNLP 2025

Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios. Recent advancements in Large Language Models (LL

2025

Measuring Sycophancy of Language Models in Multi-turn Dialogues

EMNLP 2025

Large Language Models (LLMs) are expected to provide helpful and harmless responses, yet they often exhibit sycophancy —conforming to user beliefs regardless of factual accuracy or ethical soundness. Prior research on sycophancy has primarily focused on single-turn factual correctness, overlooking t

2025

Piecing It All Together: Verifying Multi-Hop Multimodal Claims

COLING 2025main

Existing claim verification datasets often do not require systems to perform complex reasoning or effectively interpret multimodal evidence. To address this, we introduce a new task: multi-hop multimodal claim verification. This task challenges models to reason over multiple pieces of evidence from…

2025

Taxonomy-Guided Zero-Shot Recommendations with LLMs

COLING 2025main

With the emergence of large language models (LLMs) and their ability to perform a variety of tasks, their application in recommender systems (RecSys) has shown promise. However, we are facing significant challenges when deploying LLMs into RecSys, such as limited prompt length, unstructured item inf…

2024

Can Large Language Model Agents Simulate Human Trust Behavior?

NeurIPS 2024poster

Large Language Model (LLM) agents have been increasingly adopted as simulation tools to model humans in social science and role-playing applications. However, one fundamental question remains: can LLM agents really simulate human behavior? In this paper, we focus on one critical and elemental behavi…

2024

Fine-Grained Discrepancy Contrastive Learning for Robust Fake News Detection

ICASSP 2024accepted

In recent years, fake news on social media has become a significant threat to societal security, elevating fake news detection to a research priority. Among various strategies, fact-checking detection methods stand out for their accuracy, leveraging evidence from dedicated fact databases. However, t…

Cited by 0SourceScholar
2024

Position: TrustLLM: Trustworthiness in Large Language Models

ICML 2024poster

Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLM…

Cited by 95SourcePDFScholar
2023

Explainable Claim Verification via Knowledge-Grounded Reasoning with Large Language Models

EMNLP 2023long findings

Claim verification plays a crucial role in combating misinformation. While existing works on claim verification have shown promising results, a crucial piece of the puzzle that remains unsolved is to understand how to verify claims without relying on human-annotated data, which is expensive to creat…

Cited by 0SourcecodeScholar
2022

BOND: Benchmarking Unsupervised Outlier Node Detection on Static Attributed Graphs

NeurIPS 2022accept

Detecting which nodes in graphs are outliers is a relatively new machine learning task with numerous applications. Despite the proliferation of algorithms developed in recent years for this task, there has been no standard comprehensive setting for performance evaluation. Consequently, it has been d…

2022

WALNUT: A Benchmark on Semi-weakly Supervised Learning for Natural Language Understanding

NAACL 2022long

Building machine learning models for natural language understanding (NLU) tasks relies heavily on labeled data. Weak supervision has been proven valuable when large amount of labeled data is unavailable or expensive to obtain. Existing works studying weak supervision for NLU either mostly focus on a…