← Search

Shoko Wakamiya

5 accepted papers

2025

Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy

EMNLP 2025

Counterfactual data augmentation (CDA) is a promising strategy for improving hate speech classification, but automating counterfactual text generation remains a challenge. Strong attribute control can distort meaning, while prioritizing semantic preservation may weaken attribute alignment. We propos

Cited by 0SourcePDFScholar
2025

Exploring LLM Annotation for Adaptation of Clinical Information Extraction Models under Data-sharing Restrictions

ACL 2025finding

In-hospital text data contains valuable clinical information, yet deploying fine-tuned small language models (SLMs) for information extraction remains challenging due to differences in formatting and vocabulary across institutions. Since access to the original in-hospital data (source domain) is oft…

Cited by 0SourcePDFScholar
2025

MultiMSD: A Corpus for Multilingual Medical Text Simplification from Online Medical References

ACL 2025finding

We release a parallel corpus for medical text simplification, which paraphrases medical terms into expressions easily understood by patients. Medical texts written by medical practitioners contain a lot of technical terms, and patients who are non-experts are often unable to use the information effe…

2025

RecordTwin: Towards Creating Safe Synthetic Clinical Corpora

ACL 2025finding

The scarcity of publicly available clinical corpora hinders developing and applying NLP tools in clinical research. While existing work tackles this issue by utilizing generative models to create high-quality synthetic corpora, their methods require learning from the original in-hospital clinical do…

Cited by 0SourcePDFScholar
2020

Offensive Language Detection on Video Live Streaming Chat

COLING 2020main

This paper presents a prototype of a chat room that detects offensive expressions in a video live streaming chat in real time. Focusing on Twitch, one of the most popular live streaming platforms, we created a dataset for the task of detecting offensive expressions. We collected 2,000 chat posts acr…

Cited by 23SourcePDFScholar