← Search

Bracha Shapira

2 accepted papers

2025

Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

EMNLP 2025

Large language models (LLMs) often appear to excel on public benchmarks, but these high scores may mask an overreliance on dataset-specific surface cues rather than true language understanding. We introduce the **Chameleon Benchmark Overfit Detector (C-BOD)**, a meta-evaluation framework designed to

2022

Q-Ball: Modeling Basketball Games Using Deep Reinforcement Learning

AAAI 2022technical

Basketball is one of the most popular types of sports in the world. Recent technological developments have made it possible to collect large amounts of data on the game, analyze it, and discover new insights. We propose a novel approach for modeling basketball games using deep reinforcement learning…

Cited by 17SourcePDFScholar