← Search

Yueqi Song

11 accepted papers

2026

VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

ICML 2026poster

Current multimodal benchmarks often conflate reasoning with domain-specific knowledge, making it difficult to isolate and evaluate general reasoning abilities in non-expert settings. To address this, we introduce VisualPuzzles, a benchmark that targets visual reasoning while deliberately minimizing …

Cited by 0SourceScholar
2025

Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia

ACL 2025long

Despite Southeast Asia’s (SEA) extraordinary linguistic and cultural diversity, the region remains significantly underrepresented in vision-language (VL) research, resulting in AI models that inadequately capture SEA cultural nuances. To fill this gap, we present SEA-VL, an open-source initiative de…

2025

Grounding Multilingual Multimodal LLMs With Cultural Knowledge

EMNLP 2025

Multimodal Large Language Models excel in high-resource settings, but often misinterpret long-tail cultural entities and underperform in low-resource languages. To address this gap, we propose a data-centric approach that directly grounds MLLMs in cultural knowledge. Leveraging a large scale knowled

Cited by 0SourcePDFScholar
2025

OpenHands: An Open Platform for AI Software Developers as Generalist Agents

ICLR 2025poster

Software is one of the most powerful tools that we humans have at our disposal; it allows a skilled programmer to interact with the world in complex and profound ways. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that…

Cited by 32SourcePDFScholar
2025

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

ICLR 2025poster

Despite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultural contexts underrepresented. This paper introduces PANGEA, a multilingual mu…

Cited by 14SourcePDFScholar
2025

Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics

EMNLP 2025

As large language models (LLMs) are increasingly used in morally sensitive domains, it is crucial to understand how persona traits affect their moral reasoning and persuasive behavior. We present the first large-scale study of multi-dimensional persona effects in AI-AI debates over real-world moral

Cited by 0SourcePDFScholar
2025

What Is Missing in Multilingual Visual Reasoning and How to Fix It

NAACL 2025findings

NLP models today strive for supporting multiple languages and modalities, improving accessibility for diverse users. In this paper, we evaluate their multilingual, multimodal capabilities by testing on a visual reasoning task. We observe that proprietary systems like GPT-4V obtain the best performan…

2024

An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance

EMNLP 2024main

Given the rise of multimedia content, human translators increasingly focus on culturally adapting not only words but also other modalities such as images to convey the same meaning. While several applications stand to benefit from this, machine translation systems remain confined to dealing with lan…

2023

GlobalBench: A Benchmark for Global Progress in Natural Language Processing

EMNLP 2023long main

Despite the major advances in NLP, significant disparities in NLP system performance across languages still exist. Arguably, these are due to uneven resource allocation and sub-optimal incentives to work on less resourced languages. To track and further incentivize the global development of equitabl…

Cited by 0SourceScholar