← Search

Weining Shen

8 accepted papers

2026

SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports

ICLR 2026poster

Artificial Intelligence brings powerful new tools to sports, from automated officiating to tactical analysis, but these applications all depend on a core reasoning capability. Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning—a challe…

Cited by 0SourcecodeScholar
2026

VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding

ICML 2026poster

Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints and the need to capture information distributed across thousands of frames. Existing approaches either sample frames uniformly (risking information loss) …

Cited by 5SourceScholar
2025

FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models

NAACL 2025long

Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored, and their performance on complex tasks like financial agent remains unknown. This paper presents FinEval, a be…

2025

SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

ICLR 2025poster

Multimodal Large Language Models (MLLMs) are advancing the ability to reason about complex sports scenarios by integrating textual and visual information. To comprehensively evaluate their capabilities, we introduce SPORTU, a benchmark designed to assess MLLMs across multi-level sports reasoning tas…

2025

VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding

EMNLP 2025

Multimodal large language models (MLLMs) hold great promise for automating complex financial analysis. To comprehensively evaluate their capabilities, we introduce VisFinEval, the first large-scale Chinese benchmark that spans the full front-middle-back office lifecycle of financial tasks. VisFinEva

2024

SportQA: A Benchmark for Sports Understanding in Large Language Models

NAACL 2024long

A deep understanding of sports, a field rich in strategic and dynamic content, is crucial for advancing Natural Language Processing (NLP). This holds particular significance in the context of evaluating and advancing Large Language Models (LLMs), given the existing gap in specialized benchmarks. To…