← Search

Jaeyoon Jung

6 accepted papers

2026

D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

ICLR 2026poster

Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments---particularly gaming---offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining…

Cited by 0SourcecodeScholar
2026

EXPLORING FINE-TUNING OF LARGE AUDIO LANGUAGE MODELS FOR SPOKEN LANGUAGE UNDERSTANDING UNDER LIMITED SPEECH DATA

ICASSP 2026oral

Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different fine-tuning schemes including text-only, direct mixing, and curriculu…

Cited by 0SourcePDFScholar
2026

PepTri: Tri-Guided All-Atom Diffusion for Peptide Design via Physics, Evolution, and Mutual Information

ICLR 2026poster

Peptides, short chains of amino acids capable of high-specificity protein binding, represent a powerful class of therapeutics. While deep generative models have shown promise for peptide design, existing approaches are often structure-centric and therefore generate sequences and structures in a deco…

Cited by 0SourcecodeScholar
2025

CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction

ICRA 2025

Real-life robot navigation involves more than just reaching a destination; it requires optimizing movements while addressing scenario-specific goals. An intuitive way for humans to express these goals is through abstract cues like verbal commands or rough sketches. Such human guidance may lack detai

Cited by 4SourceScholar
2025

Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion

ACL 2025finding

Query expansion methods powered by large language models (LLMs) have demonstrated effectiveness in zero-shot retrieval tasks. These methods assume that LLMs can generate hypothetical documents that, when incorporated into a query vector, enhance the retrieval of real evidence. However, we challenge…

Cited by 0SourcePDFScholar
2024

EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning

ICASSP 2024accepted

We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked codec modeling that improves acoustic awareness of the pret…

Cited by 0SourceScholar