← Search

Hiromi Wakaki

12 accepted papers

2026

Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion

ICML 2026poster

Masked diffusion models have shown promising performance in generating high-quality samples in a wide range of domains, but accelerating their sampling process remains relatively underexplored. To investigate efficient samplers for masked diffusion, this paper theoretically analyzes the MaskGIT samp…

Cited by 0SourceScholar
2026

Learning to Route Languages for Multilingual Preference Optimization

ICML 2026poster

Large language models (LLMs) are trained on heterogeneous multilingual corpora, yet existing preference optimization methods often implicitly restrict each training question to a single response language or rely on a fixed dominant language for supervision. We propose language-routed preference opti…

Cited by 0SourceScholar
2026

Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

ICML 2026poster

Multimodal LLMs lack a systematic understanding of visual dynamics in complex human world activities, which requires the model to predict or simulate multiple levels of dynamic constituents, such as the general progression of actions and the associated changes of low-level details in the world. To a…

Cited by 0SourceScholar
2025

CARE: Multilingual Human Preference Learning for Cultural Awareness

EMNLP 2025

Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied. In this paper, we systematically analyze how native human cultural preferences can be incorpora

2025

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

EMNLP 2025

Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model’s ability to analyze and interpret various musical elements. These improvements primarily focused on integrating both music and text inputs. However, the potential

2025

Distillation of Discrete Diffusion through Dimensional Correlations

ICML 2025poster

Diffusion models have demonstrated exceptional performances in various fields of generative modeling, but suffer from slow sampling speed due to their iterative nature. While this issue is being addressed in continuous domains, discrete diffusion models face unique challenges, particularly in captur…

2025

VinaBench: Benchmark for Faithful and Consistent Visual Narratives

CVPR 2025poster

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge const…

Cited by 1SourcePDFScholar
2024

DiffuCOMET: Contextual Commonsense Knowledge Diffusion

ACL 2024long

Inferring contextually-relevant and diverse commonsense to understand narratives remains challenging for knowledge models. In this work, we develop a series of knowledge models, DiffuCOMET, that leverage diffusion to learn to reconstruct the implicit semantic connections between narrative contexts a…

2024

Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning

ACL 2024findings

We consider the task of building a dialogue system that can motivate users to adopt positive lifestyle changes, Motivational Interviewing (MI). Addressing such a task requires a system that could infer how to motivate the user effectively. We propose DIIR, a framework that is capable of learning and…

2024

On the Language Encoder of Contrastive Cross-modal Models

ACL 2024findings

Contrastive cross-modal models such as CLIP and CLAP aid various vision-language (VL) and audio-language (AL) tasks. However, there has been limited investigation of and improvement in their language encoder – the central component of encoding natural language descriptions of image/audio into vector…

Cited by 0SourcePDFScholar
2023

PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives

ACL 2023long

Sustaining coherent and engaging narratives requires dialogue or storytelling agents to understandhow the personas of speakers or listeners ground the narrative. Specifically, these agents must infer personas of their listeners to produce statements that cater to their interests. They must also lear…

2022

ComFact: A Benchmark for Linking Contextual Commonsense Knowledge

EMNLP 2022finding

Understanding rich narratives, such as dialogues and stories, often requires natural language processing systems to access relevant knowledge from commonsense knowledge graphs. However, these systems typically retrieve facts from KGs using simple heuristics that disregard the complex challenges of i…