← Search

Takuya Ohko

3 accepted papers

2026

GneissWeb: Preparing High Quality Data for LLMs at Scale

ICLR 2026poster

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. In this paper, we introduce **GneissWeb**, a large dataset of aro…

Cited by 0SourceScholar
2024

Incorporating Syntax and Lexical Knowledge to Multilingual Sentiment Classification on Large Language Models

ACL 2024findings

This paper exploits a sentiment extractor supported by syntactic and lexical resources to enhance multilingual sentiment classification solved through the generative approach, without retraining LLMs. By adding external information of words and phrases that have positive/negative polarities, the mul…

Cited by 4SourcePDFScholar
2023

Incorporating Syntactic Knowledge into Pre-trained Language Model using Optimization for Overcoming Catastrophic Forgetting

EMNLP 2023long findings

Syntactic knowledge is invaluable information for many tasks which handle complex or long sentences, but typical pre-trained language models do not contain sufficient syntactic knowledge. Thus it results in failures in downstream tasks that require syntactic knowledge. In this paper, we explore addi…

Cited by 0SourceScholar