← Search

Wenye Lin

3 accepted papers

2026

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

ICML 2026poster

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by surging model progress. To address …

Cited by 0SourceScholar
2025

GAMEBoT: Transparent Assessment of LLM Reasoning in Games

ACL 2025long

Large Language Models (LLMs) are increasingly deployed in real-world applications that demand complex reasoning. To track progress, robust benchmarks are required to evaluate their capabilities beyond superficial pattern recognition. However, current LLM reasoning benchmarks often face challenges su…

Cited by 0SourcePDFScholar
2023

A Simple Yet Effective Approach to Structured Knowledge Distillation

ICASSP 2023accepted

Structured prediction models aim at solving tasks where the output is a complex structure, rather than a single variable. Performing knowledge distillation for such problems is non- trivial due to their exponentially large output space. Previous works address this problem by developing particular di…

Cited by 0SourceScholar