← Search

Attapol Rutherford

5 accepted papers

2025

Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models

COLING 2025main

The rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, and commonsense, leading to the inception of certain widely-used benchmark suites such as the H6 benchmark. However, thes…

2024

Learning Job Title Representation from Job Description Aggregation Network

ACL 2024findings

Learning job title representation is a vital process for developing automatic human resource tools. To do so, existing methods primarily rely on learning the title representation through skills extracted from the job description, neglecting the rich and diverse content within. Thus, we propose an al…

Cited by 1SourcePDFScholar
2022

More Than Words: Collocation Retokenization for Latent Dirichlet Allocation Models

ACL 2022findings

Traditionally, Latent Dirichlet Allocation (LDA) ingests words in a collection of documents to discover their latent topics using word-document co-occurrences. Previous studies show that representing bigrams collocations in the input can improve topic coherence in English. However, it is unclear how…

2022

Thai Nested Named Entity Recognition Corpus

ACL 2022findings

This paper presents the first Thai Nested Named Entity Recognition (N-NER) dataset. Thai N-NER consists of 264,798 mentions, 104 classes, and a maximum depth of 8 layers obtained from 4,894 documents in the domains of news articles and restaurant reviews. Our work, to the best of our knowledge, pres…

2020

Syllable-based Neural Thai Word Segmentation

COLING 2020main

Word segmentation is a challenging pre-processing step for Thai Natural Language Processing due to the lack of explicit word boundaries. The previous systems rely on powerful neural network architecture alone and ignore linguistic substructures of Thai words. We utilize the linguistic observation th…