← Search

Derry Tanti Wijaya

15 accepted papers

2026

mR3: Multilingual Rubric-Agnostic Reward Reasoning Models

ICLR 2026poster

Evaluation using Large Language Model (LLM) judges has been widely adopted in English and shown to be effective for automatic evaluation. However, their performance does not generalize well to non-English settings, and it remains unclear what constitutes effective multilingual training for such judg…

Cited by 0SourcecodeScholar
2025

A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information

ACL 2025finding

Online discourse is increasingly trapped in a vicious cycle where polarizing language fuelstoxicity and vice versa. Identity, one of the most divisive issues in modern politics, oftenincreases polarization. Yet, prior NLP research has mostly treated toxicity and polarization asseparate problems. In…

Cited by 0SourcePDFScholar
2025

Do Language Models Understand Honorific Systems in Javanese?

ACL 2025long

The Javanese language features a complex system of honorifics that vary according to the social status of the speaker, listener, and referent. Despite its cultural and linguistic significance, there has been limited progress in developing a comprehensive corpus to capture these variations for natura…

Cited by 0SourcePDFScholar
2025

MetaMetrics: Calibrating Metrics for Generation Tasks Using Human Preferences

ICLR 2025poster

Understanding the quality of a performance evaluation metric is crucial for ensuring that model outputs align with human preferences. However, it remains unclear how well each metric captures the diverse aspects of these preferences, as metrics often excel in one particular area but not across all d…

2025

NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts

ACL 2025long

Indonesia is rich in languages and scripts. However, most NLP progress has been made using romanized text. In this paper, we present NusaAksara, a novel public benchmark for Indonesian languages that includes their original scripts. Our benchmark covers both text and image modalities and encompasses…

Cited by 0SourcePDFScholar
2025

What Do Indonesians Really Need from Language Technology? A Nationwide Survey

EMNLP 2025

Despite emerging efforts to develop NLP for Indonesia’s 700+ local languages, progress remains costly due to the need for direct engagement with native speakers. However, it is unclear what these language communities truly need from language technology. To address this, we conduct a nationwide surve

2025

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

NAACL 2025long

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we introduce WorldCuisines, a massive-scale benchmark for multilingual and multicul…

2024

Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text Generation

COLING 2024main

Predicting emotions elicited by news headlines can be challenging as the task is largely influenced by the varying nature of people’s interpretations and backgrounds. Previous works have explored classifying discrete emotions directly from news headlines. We provide a different approach to tackling…

Cited by 1SourcePDFScholar
2024

Mitigating Translationese in Low-resource Languages: The Storyboard Approach

COLING 2024main

Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which can introduce the translationese effect. This phenomenon results in translated sentences that lack fluency and naturalness in the target language. In this pape…

Cited by 2SourcePDFScholar
2024

Monitoring Hate Speech in Indonesia: An NLP-based Classification of Social Media Texts

EMNLP 2024system demonstrations

Online hate speech propagation is a complex issue, deeply influenced by both the perpetrator and the target’s cultural, historical, and societal contexts. Consequently, developing a universally robust hate speech classifier for diverse social media texts remains a challenging and unsolved task. The…

Cited by 1SourcePDFScholar
2023

RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs

ACL 2023long

Despite their unprecedented success, even the largest language models make mistakes. Similar to how humans learn and improve using feedback, previous work proposed providing language models with natural language feedback to guide them in repairing their outputs. Because human-generated critiques are…

2021

Cultural and Geographical Influences on Image Translatability of Words across Languages

NAACL 2021long

Neural Machine Translation (NMT) models have been observed to produce poor translations when there are few/no parallel sentences to train the models. In the absence of parallel data, several approaches have turned to the use of images to learn translations. Since images of words, e.g., horse may be…

2021

Detecting Frames in News Headlines and Lead Images in U.S. Gun Violence Coverage

EMNLP 2021finding

News media structure their reporting of events or issues using certain perspectives. When describing an incident involving gun violence, for example, some journalists may focus on mental health or gun regulation, while others may emphasize the discussion of gun rights. Such perspectives are called “…

Cited by 21SourcePDFScholar
2021

OpenFraming: Open-sourced Tool for Computational Framing Analysis of Multilingual Data

EMNLP 2021system demonstrations

When journalists cover a news story, they can cover the story from multiple angles or perspectives. These perspectives are called “frames,” and usage of one frame or another may influence public perception and opinion of the issue at hand. We develop a web-based system for analyzing frames in multil…

2021

“Wikily” Supervised Neural Translation Tailored to Cross-Lingual Tasks

EMNLP 2021main

We present a simple but effective approach for leveraging Wikipedia for neural machine translation as well as cross-lingual tasks of image captioning and dependency parsing without using any direct supervision from external parallel data or supervised models in the target language. We show that firs…