← Search

Xuzhao Li

4 accepted papers

2026

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos

AAAI 2026technical

Recent advances in large language models (LLMs) have improved reasoning in text and image domains, yet achieving robust video reasoning remains a significant challenge. Existing video benchmarks mainly assess shallow understanding and reasoning and allow models to exploit global context, failing to

Cited by 0SourcePDFScholar
2026

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains

AAAI 2026technical

Large language models (LLMs) increasingly rely on reinforcement learning (RL) to enhance their reasoning capabilities through feedback. A critical challenge is verifying the consistency of model-generated responses and reference answers, since these responses are often lengthy, diverse, and nuanced.

Cited by 0SourcePDFScholar