← Search

Joonseok Lee

31 accepted papers

2026

A More Word-like Image Tokenization for MLLMs

CVPR 2026

Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optim

Cited by 0SourcecodeScholar
2026

A Scalable Inter-edge Correlation Modeling in CopulaGNN for Link Sign Prediction

ICLR 2026poster

Link sign prediction on a signed graph is a task to determine whether the relationship represented by an edge is positive or negative. Since the presence of negative edges violates the graph homophily assumption that adjacent nodes are similar, regular graph methods have not been applicable without…

Cited by 0SourceScholar
2026

Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching

ICML 2026poster

Diffusion models often yield highly curved trajectories and noisy score targets due to an uninformative, memoryless forward process that induces independent data-noise coupling. We propose Adjoint Schrödinger Bridge Matching (ASBM), a generative modeling framework that recovers optimal trajectories …

Cited by 0SourceScholar
2026

Equivariant Latent Alignment via Flow Matching under Group Symmetries

ICML 2026poster

Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency. In parallel, equivariant representation learning has emerged as a powerful framework for constructing latent spaces where analytically known group transformations coul…

Cited by 0SourceScholar
2026

Fine-Grained Multi Image Object Hallucination Benchmark

CVPR 2026

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination--generating plausible yet factually inconsistent descriptions about objects. Exi

Cited by 0SourceScholar
2026

QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning

ICML 2026poster

GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate sample…

Cited by 0SourceScholar
2026

Sparsity-promoting Fine-tuning for Equivariant Materials Foundation Model

ICLR 2026poster

Pre-trained materials foundation models, or machine learning interatomic potentials, leverage general physicochemical knowledge to effectively approximate potential energy surfaces. However, they often require domain-specific calibration due to physicochemical diversity and mismatches between practi…

Cited by 0SourceScholar
2026

TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization

ICLR 2026poster

The exponential growth of video content highlights the importance of video summarization, a task that efficiently extracts key information from long videos. However, existing video summarization studies face inherent limitations in understanding complex, multimodal videos. This limitation stems from…

Cited by 0SourcecodeScholar
2025

Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance

ICASSP 2025accepted

Modern music retrieval systems often rely on fixed representations of user preferences, limiting their ability to capture users’ diverse and uncertain retrieval needs. To address this limitation, we introduce Diff4Steer, a novel generative retrieval framework that employs lightweight diffusion model…

Cited by 0SourceScholar
2025

Latent Expression Generation for Referring Image Segmentation and Grounding

ICCV 2025poster

Visual grounding tasks, such as referring image segmentation (RIS) and referring expression comprehension (REC), aim to localize a target object based on a given textual description. The target object in an image can be described in multiple ways, reflecting diverse attributes such as color, positio…

Cited by 0SourcePDFScholar
2025

SummDiff: Generative Modeling of Video Summarization with Diffusion

ICCV 2025poster

Video summarization is a task of shortening a video by choosing a subset of frames while preserving its essential moments. Despite the innate subjectivity of the task, previous works have deterministically regressed to an averaged frame score over multiple raters, ignoring the inherent subjectivity…

Cited by 0SourcePDFScholar
2025

Towards Scalable Human-aligned Benchmark for Text-guided Image Editing

CVPR 2025highlight

A variety of text-guided image editing models have been proposed recently. However, there is no widely-accepted standard evaluation method mainly due to the subjective nature of the task, letting researchers rely on manual user study. To address this, we introduce a novel Human-Aligned benchmark for…

2024

Dictionary Contrastive Learning for Efficient Local Supervision without Auxiliary Networks

ICLR 2024spotlight

While backpropagation (BP) has achieved widespread success in deep learning, it faces two prominent challenges: computational inefficiency and biological implausibility. In response to these challenges, local supervision, encompassing Local Learning (LL) and Forward Learning (FL), has emerged as a p…

Cited by 0SourcePDFScholar
2024

Isometric Representation Learning for Disentangled Latent Space of Diffusion Models

ICML 2024poster

The latent space of diffusion model mostly still remains unexplored, despite its great success and potential in the field of generative modeling. In fact, the latent space of existing diffusion models are entangled, with a distorted mapping from its latent space to image space. To tackle this proble…

2024

Towards a Complete Benchmark on Video Moment Localization

AISTATS 2024poster

In this paper, we propose and conduct a comprehensive benchmark on moment localization task, which aims to retrieve a segment that corresponds to a text query from a single untrimmed video. Our study starts from an observation that most moment localization papers report experimental results only on…

2024

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

AAAI 2024technical

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the exploration of video-acoustic signatures has been confined to sp…

Cited by 12SourcePDFScholar
2023

Activity Grammars for Temporal Action Segmentation

NeurIPS 2023poster

Sequence prediction on temporal data requires the ability to understand compositional structures of multi-level semantics beyond individual and contextual properties of parts. The task of temporal action segmentation remains challenging for the reason, aiming at translating an untrimmed activity vid…

2023

Exploration Into Translation-Equivariant Image Quantization

ICASSP 2023accepted

This is an exploratory study that discovers the current image quantization (vector quantization) do not satisfy translation equivariance in the quantized space due to aliasing. Instead of focusing on anti-aliasing, we propose a simple yet effective way to achieve translation-equivariant image quanti…

Cited by 0SourceScholar
2023

Perspective Projection-Based 3d CT Reconstruction from Biplanar X-Rays

ICASSP 2023accepted

X-ray computed tomography (CT) is one of the most common imaging techniques used to diagnose various diseases in the medical field. Its high contrast sensitivity and spatial resolution allow the physician to observe details of body parts such as bones, soft tissue, blood vessels, etc. As it involves…

Cited by 0SourceScholar
2023

Towards Physically Reliable Molecular Representation Learning

UAI 2023poster

Estimating the energetic properties of molecular systems is a critical task in material design. Machine learning has shown remarkable promise on this task over classical force fields, but a fully data-driven approach suffers from limited labeled data; not just the amount of available data lacks, but…

Cited by 2SourcePDFScholar
2023

Towards Robust and Smooth 3D Multi-Person Pose Estimation from Monocular Videos in the Wild

ICCV 2023poster

3D pose estimation is an invaluable task in computer vision with various practical applications. Especially, 3D pose estimation for multi-person from a monocular video (3DMPPE) is particularly challenging and is still largely uncharted, far from applying to in-the-wild scenarios yet. We pose three u…

Cited by 18PDFScholar
2023

VisAlign: Dataset for Measuring the Alignment between AI and Humans in Visual Perception

NeurIPS 2023poster

AI alignment refers to models acting towards human-intended goals, preferences, or ethical principles. Analyzing the similarity between models and humans can be a proxy measure for ensuring AI safety. In this paper, we focus on the models' visual perception alignment with humans, further referred to…

2022

A Conservative Approach for Unbiased Learning on Unknown Biases

CVPR 2022poster

Although convolutional neural networks (CNNs) achieve state-of-the-art in image classification, recent works address their unreliable predictions due to their excessive dependence on biased training data. Existing unbiased modeling postulates that the bias in the dataset is obvious to know, but it i…

Cited by 20PDFcodeScholar
2021

Vid-ODE: Continuous-Time Video Generation with Neural Ordinary Differential Equation

AAAI 2021technical

Video generation models often operate under the assumption of fixed frame rates, which leads to suboptimal performance when it comes to handling flexible frame rates (e.g., increasing the frame rate of the more dynamic portion of the video as well as handling missing video frames). To resolve the re…

2020

Large Scale Video Representation Learning via Relational Graph Clustering

CVPR 2020poster

Representation learning is widely applied for various tasks on multimedia data, e.g., retrieval and search. One approach for learning useful representation is by utilizing the relationships or similarities between examples. In this work, we explore two promising scalable representation learning appr…

Cited by 27PDFScholar
2019

N-GCN: Multi-scale Graph Convolution for Semi-supervised Node Classification

UAI 2019poster

Graph Convolutional Networks (GCNs) have shown significant improvements in semi-supervised learning on graph-structured data. Concurrently, unsupervised learning of graph embeddings has benefited from the information contained in random walks. In this paper, we propose a model: Network of GCNs (N-GC…