Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
The rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, and commonsense, leading to the inception of certain widely-used benchmark suites such as the H6 benchmark. However, thes…