← Search

SouYoung Jin

11 accepted papers

2026

MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs). However, existing approaches often struggle with fine-grained visual grounding due to semantic entanglement in visual …

Cited by 0SourceScholar
2024

LangNav: Language as a Perceptual Representation for Navigation

NAACL 2024findings

We explore the use of language as a perceptual representation for vision-and-language navigation (VLN), with a focus on low-data settings. Our approach uses off-the-shelf vision systems for image captioning and object detection to convert an agent’s egocentric panoramic view at each time step into n…

2023

Learning Human Action Recognition Representations Without Real Humans

NeurIPS 2023poster

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets contain images of people and hence are accompanied with issues related to privacy, ethics, and data protection, often pr…

2023

Leveraging Temporal Context in Low Representational Power Regimes

CVPR 2023poster

Computer vision models are excellent at identifying and exploiting regularities in the world. However, it is computationally costly to learn these regularities from scratch. This presents a challenge for low-parameter models, like those running on edge devices (e.g. smartphones). Can the performance…

Cited by 2SourcePDFScholar
2022

Cross-Modal Discrete Representation Learning

ACL 2022long

In contrast to recent advances focusing on high-level representation learning across modalities, in this work we present a self-supervised learning framework that is able to learn a representation that captures finer levels of granularity across different modalities such as concepts or events repres…

2022

How Transferable are Video Representations Based on Synthetic Data?

NeurIPS 2022accept

Action recognition has improved dramatically with massive-scale video datasets. Yet, these datasets are accompanied with issues related to curation cost, privacy, ethics, bias, and copyright. Compared to that, only minor efforts have been devoted toward exploring the potential of synthetic video dat…

2021

Spoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions

CVPR 2021poster

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where, who and how) of the observed event and exclude background info…

Cited by 86PDFScholar
2019

Automatic Adaptation of Object Detectors to New Domains Using Self-Training

CVPR 2019poster

This work addresses the unsupervised adaptation of an existing object detector to a new target domain. We assume that a large number of unlabeled videos from this domain are readily available. We automatically obtain labels on the target data by using high-confidence detections from the existing det…

Cited by 184PDFScholar
2018

Unsupervised Hard Example Mining from Videos for Improved Object Detection

ECCV 2018poster

Important gains have recently been obtained in object detection by using training objectives that focus on {em hard negative} examples, i.e., negative examples that are currently rated as positive or ambiguous by the detector. These examples can strongly influence parameters when the network is trai…

Cited by 90SourcePDFScholar
2017

End-To-End Face Detection and Cast Grouping in Movies Using Erdos-Renyi Clustering

ICCV 2017spotlight

We present an end-to-end system for detecting and clustering faces by identity in full-length movies. Unlike works that start with a predefined set of detected faces, we consider the end-to-end problem of detection and clustering together. We make three separate contributions. First, we combine a st…

Cited by 50PDFScholar