← Search

Junchen Ding

2 accepted papers

2026

Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test Oracles

AAAI 2026technical

Large Language Models (LLMs) have achieved significant progress in language understanding and reasoning. Evaluating and analyzing their logical reasoning abilities has therefore become essential. However, existing datasets and benchmarks are often limited to overly simplistic, unnatural, or contextu

Cited by 0SourcePDFScholar
2025

TombRaider: Entering the Vault of History to Jailbreak Large Language Models

EMNLP 2025

**Warning: This paper contains content that may involve potentially harmful behaviours, discussed strictly for research purposes.**Jailbreak attacks can hinder the safety of Large Language Model (LLM) applications, especially chatbots. Studying jailbreak techniques is an important AI red teaming tas

Cited by 0SourcePDFScholar