← Search

Eiji Aramaki

9 accepted papers

2025

Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy

EMNLP 2025

Counterfactual data augmentation (CDA) is a promising strategy for improving hate speech classification, but automating counterfactual text generation remains a challenge. Strong attribute control can distort meaning, while prioritizing semantic preservation may weaken attribute alignment. We propos

Cited by 0SourcePDFScholar
2025

Exploring LLM Annotation for Adaptation of Clinical Information Extraction Models under Data-sharing Restrictions

ACL 2025finding

In-hospital text data contains valuable clinical information, yet deploying fine-tuned small language models (SLMs) for information extraction remains challenging due to differences in formatting and vocabulary across institutions. Since access to the original in-hospital data (source domain) is oft…

Cited by 0SourcePDFScholar
2025

Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors

EMNLP 2025

This paper investigates how LLMs encode inputs with typos. We hypothesize that specific neurons and attention heads recognize typos and fix them internally using local and global contexts. We introduce a method to identify typo neurons and typo heads that work actively when inputs contain typos. Our

Cited by 0SourcePDFScholar
2025

MultiMSD: A Corpus for Multilingual Medical Text Simplification from Online Medical References

ACL 2025finding

We release a parallel corpus for medical text simplification, which paraphrases medical terms into expressions easily understood by patients. Medical texts written by medical practitioners contain a lot of technical terms, and patients who are non-experts are often unable to use the information effe…

2025

RecordTwin: Towards Creating Safe Synthetic Clinical Corpora

ACL 2025finding

The scarcity of publicly available clinical corpora hinders developing and applying NLP tools in clinical research. While existing work tackles this issue by utilizing generative models to create high-quality synthetic corpora, their methods require learning from the original in-hospital clinical do…

Cited by 0SourcePDFScholar
2024

A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages

COLING 2024main

User-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilin…

2024

QA-based Event Start-Points Ordering for Clinical Temporal Relation Annotation

COLING 2024main

Temporal relation annotation in the clinical domain is crucial yet challenging due to its workload and the medical expertise required. In this paper, we propose a novel annotation method that integrates event start-points ordering and question-answering (QA) as the annotation format. By focusing onl…

2023

Comparative evaluation of boundary-relaxed annotation for Entity Linking performance

ACL 2023long

Entity Linking performance has a strong reliance on having a large quantity of high-quality annotated training data available. Yet, manual annotation of named entities, especially their boundaries, is ambiguous, error-prone, and raises many inconsistencies between annotators. While imprecise boundar…

2020

Offensive Language Detection on Video Live Streaming Chat

COLING 2020main

This paper presents a prototype of a chat room that detects offensive expressions in a video live streaming chat in real time. Focusing on Twitch, one of the most popular live streaming platforms, we created a dataset for the task of detecting offensive expressions. We collected 2,000 chat posts acr…

Cited by 23SourcePDFScholar