← Search

Longxiang Wang

3 accepted papers

2026

EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer

AAAI 2026technical

Most existing spatial reasoning benchmarks focus on static or globally observable environments, failing to capture the challenges of long-horizon reasoning and memory utilization under partial observability and dynamic changes. We introduce two dynamic spatial benchmarks—locally observable maze navi

Cited by 0SourcePDFScholar
2026

PsyPARSE: Retrieval-Augmented Slow Thinking for Personalized Empathetic Counseling

AAAI 2026technical

The escalating global demand for mental health services highlights the potential of Large Language Models (LLMs) in psychological counseling. However, current LLM-based approaches, particularly fine-tuned models, are constrained by data distribution biases, leading to limited therapeutic diversity a

Cited by 0SourcePDFScholar
2025

CALM: Curiosity-Driven Auditing for Large Language Models

AAAI 2025technical

Auditing Large Language Models (LLMs) is a crucial and challenging task. In this study, we focus on auditing black-box LLMs without access to their parameters, only to the provided service. We treat this type of auditing as a black-box optimization problem where the goal is to automatically uncover…