← Search

Sida Li

4 accepted papers

2026

LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena

ICLR 2026poster

With the rapid progress of large language models (LLMs) trained on every available piece of data, it becomes increasingly challenging to reliably evaluate their intelligence due to potential data contamination and benchmark overfitting. To overcome these challenges, we investigate a new angle of ben…

Cited by 0SourceScholar
2026

Mapping Overlaps in Benchmarks through Perplexity in the Wild

ICLR 2026poster

We construct benchmark signatures that capture the capacity required for strong performance to characterize large language model (LLM) benchmarks and their meaningful overlaps. Formally, we define them as sets of salient tokens drawn from **in-the-wild** corpora whose LLM token perplexity, reflectin…

Cited by 0SourcecodeScholar
2025

ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning

NeurIPS 2025poster

Recent models such as OpenAI o1 and DeepSeek-R1 have demonstrated strong performance on reasoning-intensive tasks by generating extended Chain-of-Thought (CoT) traces. While longer reasoning helps with thorough exploration of solution paths for complex problems, it also often leads to inefficient an…

Cited by 0SourceScholar