← Search

Hyeongkeun Lee

5 accepted papers

2026

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

CVPR 2026

Dense Video Captioning (DVC) is a challenging multimodal task that involves temporally localizing multiple events within a video and describing them with natural language. While query-based frameworks enable the simultaneous, end-to-end processing of localization and captioning, their reliance on sh

Cited by 0SourceScholar
2025

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

ICASSP 2025accepted

Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse audio and visual data, missing important semantic cues from eac…

Cited by 0SourceScholar
2024

EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning

ICML 2024poster

Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified in many learning methods, audio-visual learning has struggled to fully harness…