← Search

Jungbeom Lee

17 accepted papers

2026

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation

RSS 2026poster

Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level actions under distribution shifts. For example, even in a simple pick-up task with identical scene layouts, camera viewpoi…

Cited by 0SourceScholar
2026

Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers

ICML 2026poster

Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omission, where specified objects or attributes fail to emerge in the generated image. By performing linear probing on text tokens, we demonstrate that t…

Cited by 0SourceScholar
2026

FlowFixer: Towards Detail-Preserving Subject-Driven Generation

CVPR 2026

We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation from visual references, avoiding ambiguities in language pr

Cited by 0SourcecodeScholar
2026

The Truth Stays in the Family: Enhancing Contextual Truthfulness via Inherited Heads in Model Lineages

ICML 2026poster

Recent advances in large language models (LLMs) have led to the emergence of specialized multimodal LLMs (MLLMs), forming distinct model families that share a common foundation language models. Despite this evolutionary trend, it remains unexplored whether a fundamental behavioral link exists betwee…

Cited by 0SourceScholar
2025

Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

ICCV 2025poster

While CLIP has significantly advanced multimodal understanding by bridging vision and language, the inability to grasp negation -- such as failing to differentiate concepts like "parking" from "no parking" -- poses substantial challenges.By analyzing the data used in the public CLIP model's pre-trai…

Cited by 0SourcePDFScholar
2024

Toward Interactive Regional Understanding in Vision-Large Language Models

NAACL 2024long

Recent Vision-Language Pre-training (VLP) models have demonstrated significant advancements. Nevertheless, these models heavily rely on image-text pairs that capture only coarse and global information of an image, leading to a limitation in their regional understanding ability. In this work, we intr…

Cited by 1SourcePDFScholar
2023

Improving Visual Prompt Tuning for Self-supervised Vision Transformers

ICML 2023poster

Visual Prompt Tuning (VPT) is an effective tuning method for adapting pretrained Vision Transformers (ViTs) to downstream tasks. It leverages extra learnable tokens, known as prompts, which steer the frozen pretrained ViTs. Although VPT has demonstrated its applicability with supervised vision trans…

2023

Weakly Supervised Referring Image Segmentation with Intra-Chunk and Inter-Chunk Consistency

ICCV 2023poster

Referring image segmentation (RIS) aims to localize the object in an image referred by a natural language expression. Most previous studies learn RIS with a large-scale dataset containing segmentation labels, but they are costly. We present a weakly supervised learning method for RIS that only uses…

Cited by 29PDFScholar
2022

Bridging the Gap Between Classification and Localization for Weakly Supervised Object Localization

CVPR 2022poster

Weakly supervised object localization aims to find a target object region in a given image with only weak supervision, such as image-level labels. Most existing methods use a class activation map (CAM) to generate a localization map; however, a CAM identifies only the most discriminative parts of a…

Cited by 56PDFcodeScholar
2022

Perception Prioritized Training of Diffusion Models

CVPR 2022poster

Diffusion models learn to restore noisy data, which is corrupted with different levels of noise, by optimizing the weighted sum of the corresponding loss terms, i.e., denoising score matching loss. In this paper, we show that restoring data corrupted with certain noise levels offers a proper pretext…

Cited by 254PDFcodeScholar
2022

Weakly Supervised Semantic Segmentation Using Out-of-Distribution Data

CVPR 2022poster

Weakly supervised semantic segmentation (WSSS) methods are often built on pixel-level localization maps obtained from a classifier. However, training on class labels only, classifiers suffer from the spurious correlation between foreground and background cues (e.g. train and rail), fundamentally bou…

Cited by 127PDFcodeScholar
2021

Anti-Adversarially Manipulated Attributions for Weakly and Semi-Supervised Semantic Segmentation

CVPR 2021poster

Weakly supervised semantic segmentation produces a pixel-level localization from class labels; but a classifier trained on such labels is likely to restrict its focus to a small discriminative region of the target object. AdvCAM is an attribution map of an image that is manipulated to increase the c…

Cited by 307PDFcodeScholar
2021

BBAM: Bounding Box Attribution Map for Weakly Supervised Semantic and Instance Segmentation

CVPR 2021poster

Weakly supervised segmentation methods using bounding box annotations focus on obtaining a pixel-level mask from each box containing an object. Existing methods typically depend on a class-agnostic mask generator, which operates on the low-level information intrinsic to an image. In this work, we ut…

Cited by 232PDFcodeScholar
2021

Reducing Information Bottleneck for Weakly Supervised Semantic Segmentation

NeurIPS 2021poster

Weakly supervised semantic segmentation produces pixel-level localization from class labels; however, a classifier trained on such labels is likely to focus on a small discriminative region of the target object. We interpret this phenomenon using the information bottleneck principle: the final layer…

2019

FickleNet: Weakly and Semi-Supervised Semantic Image Segmentation Using Stochastic Inference

CVPR 2019poster

The main obstacle to weakly supervised semantic image segmentation is the difficulty of obtaining pixel-level information from coarse image-level annotations. Most methods based on image-level annotations use localization maps obtained from the classifier, but these only focus on the small discrimin…

Cited by 557PDFScholar
2019

Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic Segmentation

ICCV 2019poster

When a deep neural network is trained on data with only image-level labeling, the regions activated in each image tend to identify only a small region of the target object. We propose a method of using videos automatically harvested from the web to identify a larger region of the target object by us…

Cited by 45PDFScholar