← Search

Youngwook KIM

10 accepted papers

2025

Analyzing Offensive Language Dataset Insights from Training Dynamics and Human Agreement Level

COLING 2025main

Implicit hate speech detection is challenging due to its subjectivity and context dependence, with existing models often struggling in outof-domain scenarios. We propose CONELA, a novel data refinement strategy that enhances model performance and generalization by integrating human annotation agreem…

Cited by 0SourcePDFScholar
2024

SharedCon: Implicit Hate Speech Detection using Shared Semantics

ACL 2024findings

The ever-growing presence of hate speech on social network services and other online platforms not only fuels online harassment but also presents a growing challenge for hate speech detection. As this task is akin to binary classification, one of the promising approaches for hate speech detection is…

Cited by 6SourcePDFScholar
2023

Bridging the Gap Between Model Explanations in Partially Annotated Multi-Label Classification

CVPR 2023poster

Due to the expensive costs of collecting labels in multi-label classification datasets, partially annotated multi-label classification has become an emerging field in computer vision. One baseline approach to this task is to assume unobserved labels as negative labels, but this assumption induces la…

2023

ConPrompt: Pre-training a Language Model with Machine-Generated Data for Implicit Hate Speech Detection

EMNLP 2023long findings

Implicit hate speech detection is a challenging task in text classification since no explicit cues (e.g., swear words) exist in the text. While some pre-trained language models have been developed for hate speech detection, they are not specialized in implicit hate speech. Recently, an implicit hate…

Cited by 0SourceScholar
2023

SurProGenes: Survival Risk-Ordered Representation of Cancer Patients and Genes for the Identification of Prognostic Genes

ICML 2023poster

Identifying prognostic genes associated with patient survival is an important goal in cancer genomics, as this information could inform treatment approaches and improve patient outcomes. However, the identification of prognostic genes is complicated by the high dimensionality of genetic data, which…

Cited by 0SourcePDFScholar
2022

Dual Task Framework for Improving Persona-Grounded Dialogue Dataset

AAAI 2022technical

This paper introduces a simple yet effective data-centric approach for the task of improving persona-conditioned dialogue agents. Prior model-centric approaches unquestioningly depend on the raw crowdsourced benchmark datasets such as Persona-Chat. In contrast, we aim to fix annotation artifacts in…

Cited by 9SourcePDFScholar
2022

Generalizable Implicit Hate Speech Detection Using Contrastive Learning

COLING 2022main

Hate speech detection has gained increasing attention with the growing prevalence of hateful contents. When a text contains an obvious hate word or expression, it is fairly easy to detect it. However, it is challenging to identify implicit hate speech in nuance or context when there are insufficient…

2022

Large Loss Matters in Weakly Supervised Multi-Label Classification

CVPR 2022poster

Weakly supervised multi-label classification (WSML) task, which is to learn a multi-label classification using partially observed labels per image, is becoming increasingly important due to its huge annotation cost. In this work, we first regard unobserved labels as negative labels, casting the WSML…

Cited by 79PDFcodeScholar
2022

Modularized Transfer Learning with Multiple Knowledge Graphs for Zero-shot Commonsense Reasoning

NAACL 2022long

Commonsense reasoning systems should be able to generalize to diverse reasoning cases. However, most state-of-the-art approaches depend on expensive data annotations and overfit to a specific benchmark without learning how to perform general semantic reasoning. To overcome these drawbacks, zero-shot…

Cited by 16SourcePDFScholar
2022

TrustAL: Trustworthy Active Learning Using Knowledge Distillation

AAAI 2022technical

Active learning can be defined as iterations of data labeling, model training, and data acquisition, until sufficient labels are acquired. A traditional view of data acquisition is that, through iterations, knowledge from human labels and models is implicitly distilled to monotonically increase the…

Cited by 10SourcePDFScholar