← Search

Beomseok Lee

6 accepted papers

2026

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

ICML 2026poster

Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, error accumulation, and limited self-correction. Diffusion Multimodal Large Language Models (dMLLMs) unmask tokens in an order-agnostic process, improving efficienc…

Cited by 0SourceScholar
2025

Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection

COLING 2025main

While crowdsourcing is an established solution for facilitating and scaling the collection of speech data, the involvement of non-experts necessitates protocols to ensure final data quality. To reduce the costs of these essential controls, this paper investigates the use of Speech Foundation Models…

2024

XDetox: Text Detoxification with Token-Level Toxicity Explanations

EMNLP 2024main

Methods for mitigating toxic content through masking and infilling often overlook the decision-making process, leading to either insufficient or excessive modifications of toxic tokens. To address this challenge, we propose XDetox, a novel method that integrates token-level toxicity explanations wit…

Cited by 0SourcePDFScholar
2022

Language Model Augmented Monotonic Attention for Simultaneous Translation

NAACL 2022long

The state-of-the-art adaptive policies for Simultaneous Neural Machine Translation (SNMT) use monotonic attention to perform read/write decisions based on the partial source and target sequences. The lack of sufficient information might cause the monotonic attention to take poor read/write decisions…

Cited by 9SourcePDFScholar
2021

Task Aware Multi-Task Learning for Speech to Text Tasks

ICASSP 2021accepted

In general, the direct Speech-to-text translation (ST) is jointly trained with Automatic Speech Recognition (ASR), and Machine Translation (MT) tasks. However, the issues with the current joint learning strategies inhibit the knowledge transfer across these tasks. We propose a task modulation networ…

Cited by 0SourceScholar
2020

End-end Speech-to-Text Translation with Modality Agnostic Meta-Learning

ICASSP 2020accepted

Collecting large amounts of data to train end-to-end Speech Translation (ST) models is more difficult compared to the ASR and MT tasks. Previous studies have proposed the use of transfer learning approaches to overcome the above difficulty. These approaches benefit from weakly supervised training da…

Cited by 0SourceScholar