← Search

Janghoon Cho

4 accepted papers

2026

FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding

ICLR 2026poster

Recent studies in long video understanding have harnessed the advanced visual-language reasoning capabilities of Large Multimodal Models (LMMs), driving the evolution of video-LMMs specialized for processing extended video sequences. However, the scalability of these models is severely limited by th…

Cited by 0SourceScholar
2025

Generalized Contrastive Learning for Universal Multimodal Retrieval

NeurIPS 2025poster

Despite their consistent performance improvements, cross-modal retrieval models (e.g., CLIP) show degraded performances with retrieving keys composed of fused image-text modality (e.g., Wikipedia pages with both images and text). To address this critical challenge, multimodal retrieval has been rece…

Cited by 5SourceScholar
2024

Balanced Learning for Multi-Domain Long-Tailed Speaker Recognition

ICASSP 2024accepted

This paper considers two types of imbalance problems commonly inherent in large-scale datasets: multiple domain and class imbalance. Class imbalance causes the algorithm to be biased toward the majority classes, and multiple-domain data results in significant performance disparities for different do…

Cited by 0SourceScholar
2021

Subspectral Normalization for Neural Audio Data Processing

ICASSP 2021accepted

Convolutional Neural Networks are widely used in various machine learning domains. In image processing, the features can be obtained by applying 2D convolution to all spatial dimensions of the input. However, in the audio case, frequency domain input like Mel-Spectrogram has different and unique cha…

Cited by 0SourceScholar