← Search

Sang Yun Kwon

3 accepted papers

2025

JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

NAACL 2025long

Recent advancements in instruction fine-tuning, alignment methods such as reinforcement learning from human feedback (RLHF), and optimization techniques like direct preference optimization (DPO), have significantly enhanced the adaptability of large language models (LLMs) to user preferences. Howeve…

2025

Voice of a Continent: Mapping Africa’s Speech Technology Frontier

EMNLP 2025

Africa’s rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent’s speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench

2024

Gazelle: An Instruction Dataset for Arabic Writing Assistance

EMNLP 2024finding

Writing has long been considered a hallmark of human intelligence and remains a pinnacle task for artificial intelligence (AI) due to the intricate cognitive processes involved. Recently, rapid advancements in generative AI, particularly through the development of Large Language Models (LLMs), have…

Cited by 0SourcePDFScholar