← Search

Jiyoung Lee

30 accepted papers

2026

LEARNING WHAT TO HEAR: BOOSTING SOUND-SOURCE ASSOCIATION FOR ROBUST AUDIOVISUAL INSTANCE SEGMENTATION

ICASSP 2026poster

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform additive fusion prevents queries from specializing to different sound sources, whil…

Cited by 0SourcePDFScholar
2025

Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations

CVPR 2025poster

View-invariant representation learning from egocentric (first-person, ego) and exocentric (third-person, exo) videos is a promising approach toward generalizing video understanding systems across multiple viewpoints. However, this area has been underexplored due to the substantial differences in per…

2025

Read, Watch and Scream! Sound Generation from Text and Video

AAAI 2025technical

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely, text-to-audio generation methods generate high-quality audio b…

2025

Single Ground Truth Is Not Enough: Adding Flexibility to Aspect-Based Sentiment Analysis Evaluation

NAACL 2025long

Aspect-based sentiment analysis (ABSA) is a challenging task of extracting sentiments along with their corresponding aspects and opinion terms from the text.The inherent subjectivity of span annotation makes variability in the surface forms of extracted terms, complicating the evaluation process.Tra…

2025

Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties

NeurIPS 2025poster

Large Language Models (LLMs) are predominantly evaluated on Standard American English (SAE), often overlooking the diversity of global English varieties. This narrow focus may raise fairness concerns as degraded performance on non-standard varieties can lead to unequal benefits for users worldwide.…

Cited by 0SourceScholar
2024

KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge

ACL 2024findings

To reliably deploy Large Language Models (LLMs) in a specific country, they must possess an understanding of the nation’s culture and basic knowledge. To this end, we introduce National Alignment, which measures the alignment between an LLM and a targeted country from two aspects: social value align…

2024

Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation

ICLR 2024poster

Text-to-3D generation has shown rapid progress in recent days with the advent of score distillation sampling (SDS), a methodology of using pretrained text-to-2D diffusion models to optimize a neural radiance field (NeRF) in a zero-shot setting. However, the lack of 3D awareness in the 2D diffusion m…

2023

Dense Text-to-Image Generation with Attention Modulation

ICCV 2023poster

Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image region. To address this, we propose DenseDiffusion, a training-free method that adapts a pre-trained text-to-image model t…

Cited by 125PDFcodeScholar
2023

Exploration Into Translation-Equivariant Image Quantization

ICASSP 2023accepted

This is an exploratory study that discovers the current image quantization (vector quantization) do not satisfy translation equivariance in the quantized space due to aliasing. Instead of focusing on anti-aliasing, we propose a simple yet effective way to achieve translation-equivariant image quanti…

Cited by 0SourceScholar
2023

Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning

ICCV 2023poster

Compositional zero-shot learning (CZSL) aims to recognize unseen compositions with prior knowledge of known primitives (attribute and object). Previous works for CZSL often suffer from grasping the contextuality between attribute and object, as well as the discriminability of visual features, and th…

Cited by 22PDFcodeScholar
2023

MIDMs: Matching Interleaved Diffusion Models for Exemplar-Based Image Translation

AAAI 2023technical

We present a novel method for exemplar-based image translation, called matching interleaved diffusion models (MIDMs). Most existing methods for this task were formulated as GAN-based matching-then-generation framework. However, in this framework, matching errors induced by the difficulty of semantic…

2023

Robust Camera Pose Refinement for Multi-Resolution Hash Encoding

ICML 2023poster

Multi-resolution hash encoding has recently been proposed to reduce the computational cost of neural renderings, such as NeRF. This method requires accurate camera poses for the neural renderings of given scenes. However, contrary to previous methods jointly optimizing camera poses and 3D scenes, th…

Cited by 28SourcePDFScholar
2023

VisAlign: Dataset for Measuring the Alignment between AI and Humans in Visual Perception

NeurIPS 2023poster

AI alignment refers to models acting towards human-intended goals, preferences, or ethical principles. Analyzing the similarity between models and humans can be a proxy measure for ensuring AI safety. In this paper, we focus on the models' visual perception alignment with humans, further referred to…

2022

Multi-Domain Unsupervised Image-to-Image Translation with Appearance Adaptive Convolution

ICASSP 2022accepted

Over the past few years, image-to-image (I2I) translation methods have been proposed to translate a given image into diverse outputs. Despite the impressive results, they mainly focus on the I2I translation between two domains, so the multi-domain I2I translation still remains a challenge. To addres…

Cited by 0SourceScholar
2022

Mutual Information Divergence: A Unified Metric for Multimodal Generative Models

NeurIPS 2022accept

Text-to-image generation and image captioning are recently emerged as a new experimental paradigm to assess machine intelligence. They predict continuous quantity accompanied by their sampling techniques in the generation, making evaluation complicated and intractable to get marginal distributions.…

2022

Pin the Memory: Learning To Generalize Semantic Segmentation

CVPR 2022poster

The rise of deep neural networks has led to several breakthroughs for semantic segmentation. In spite of this, a model trained on source domain often fails to work properly in new challenging domains, that is directly concerned with the generalization capability of the model. In this paper, we prese…

Cited by 68PDFcodeScholar
2022

PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation

ECCV 2022poster

"Online stereo adaptation tackles the domain shift problem, caused by different environments between synthetic (training) and real (test) datasets, to promptly adapt stereo models in dynamic real-world applications such as autonomous driving. However, previous methods often fail to counteract partic…

2022

Specializing Multi-domain NMT via Penalizing Low Mutual Information

EMNLP 2022main

Multi-domain Neural Machine Translation (NMT) trains a single model with multiple domains. It is appealing because of its efficacy in handling multiple domains within one model. An ideal multi-domain NMT learns distinctive domain characteristics simultaneously, however, grasping the domain peculiari…

2021

Bridge To Answer: Structure-Aware Graph Interaction Network for Video Question Answering

CVPR 2021poster

This paper presents a novel method, termed Bridge to Answer, to infer correct answers for questions about a given video by leveraging adequate graph interactions of heterogeneous crossmodal graphs. To realize this, we learn question conditioned visual graphs by exploiting the relation between video…

Cited by 118PDFScholar
2021

Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech Separation

CVPR 2021poster

In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between co-occurring audio and video. Thus, their performance heavily depend…

Cited by 56PDFScholar
2018

Spatiotemporal Attention Based Deep Neural Networks for Emotion Recognition

ICASSP 2018accepted

We propose a spatiotemporal attention based deep neural networks for dimensional emotion recognition in facial videos. To learn the spatiotemporal attention that selectively focuses on emotional sailient parts within facial videos, we formulate the spatiotemporal encoder-decoder network using Convol…

Cited by 0SourceScholar