← Search

Kim Sung-Bin

7 accepted papers

2026

A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search

ICML 2026poster

Fine-tuning Large Language Models (LLMs) with Low-Rank Adaptation (LoRA) offers a resource-efficient way to personalize or specialize. However, LoRA is highly sensitive to hyperparameter choices, and performing an exhaustive hyperparameter search remains computationally intensive. To address these c…

Cited by 0SourceScholar
2025

AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

ICLR 2025poster

Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understandi…

2025

Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics

CVPR 2025highlight

Recent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the perceptual alignment between varying speech characteristics and corresponding lip movements. In this work, we claim that three…

2025

SoundBrush: Sound as a Brush for Visual Scene Editing

AAAI 2025technical

We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorporate audio information for editing visual scenes. Inspired by existing image-editing works, we frame this task as a supe…

Cited by 0SourcePDFScholar
2025

VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models

ICCV 2025poster

We present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting voice-impaired individuals. Building on the success of Neural Codec Language Mode…

Cited by 0SourcePDFScholar
2024

SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models

NAACL 2024findings

Despite the recent advances in artificial intelligence, building social intelligence remains a challenge.Among social signals, laughter is one of the distinctive expressions that occurs during social interactions between humans.In this work, we tackle a new challenge for machines to understand the r…

2023

Sound to Visual Scene Generation by Audio-to-Visual Latent Alignment

CVPR 2023poster

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We design a model that works by scheduling the learning procedur…