← Search

Seonho Lee

7 accepted papers

2026

Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

ICML 2026poster

Recent multimodal large language models have demonstrated strong reasoning ability, yet their reliability as automated evaluators remains limited by a critical weakness: when visual evidence conflicts with textual cues, MLLM judges tend to reward plausible narratives over perceptually correct answer…

Cited by 0SourceScholar
2026

What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging

ICLR 2026poster

State-of-the-art vision-language models (VLMs) suffer from a critical failure in understanding negation, often referred to as affirmative bias. This limitation is particularly severe in described object detection (DOD) tasks. To address this, we propose two primary contributions: (1) a new dataset p…

Cited by 0SourceScholar
2025

3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation

EMNLP 2025

Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in their understanding of 3D spatial structures.We propose Geometric Distillation, a lightweight, annotation-free fine-tuning framework that injects human-ins

2025

Doodle to Detect: A Goofy but Powerful Approach to Skeleton-based Hand Gesture Recognition

NeurIPS 2025poster

Skeleton-based hand gesture recognition plays a crucial role in enabling intuitive human–computer interaction. Traditional methods have primarily relied on hand-crafted features—such as distances between joints or positional changes across frames—to alleviate issues from viewpoint variation or body…

Cited by 0SourcecodeScholar
2025

DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation

ICLR 2025poster

Score distillation sampling (SDS) has emerged as an effective framework in text-driven 3D editing tasks, leveraging diffusion models for 3D-consistent editing. However, existing SDS-based 3D editing methods suffer from long training times and produce low-quality results. We identify that the root ca…

Cited by 13SourcePDFScholar
2025

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

CVPR 2025poster

Open-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories. We identify two primary challenges in OVPS: (1) the difficulty in aligning part-level image-text correspondence, and (2) the lack of structural understanding in segmenting object pa…

2024

Understanding Multi-Granularity for Open-Vocabulary Part Segmentation

NeurIPS 2024poster

Open-vocabulary part segmentation (OVPS) is an emerging research area focused on segmenting fine-grained entities using diverse and previously unseen vocabularies. Our study highlights the inherent complexities of part segmentation due to intricate boundaries and diverse granularity, reflecting the…

Cited by 2SourcePDFScholar