← Search

Jaeseok Jeong

12 accepted papers

2026

Improving Black-Box Generative Attacks via Generator Semantic Consistency

ICLR 2026poster

Transfer attacks optimize on a surrogate and deploy to a black-box target. While iterative optimization attacks in this paradigm are limited by their per-input cost limits efficiency and scalability due to multistep gradient updates for each input, generative attacks alleviate these by producing adv…

Cited by 0SourcecodeScholar
2026

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

ICLR 2026poster

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a promising condition for temporally controlled video generatio…

Cited by 0SourceScholar
2025

StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance

ICCV 2025poster

In the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled more precise control over style and content. However, existing methods often suffer from content leakage, where undesired…

Cited by 0SourcePDFScholar
2025

TCFG: Tangential Damping Classifier-free Guidance

CVPR 2025poster

Diffusion models have achieved remarkable success in text-to-image synthesis, largely attributed to the use of classifier-free guidance (CFG), which enables high-quality, condition-aligned image generation. CFG combines the conditional score (e.g., text-conditioned) with the unconditional score to c…

Cited by 0SourcePDFScholar
2024

Diffusion-Guided Weakly Supervised Semantic Segmentation

ECCV 2024poster

"Weakly Supervised Semantic Segmentation (WSSS) with classification labels typically uses Class Activation Maps to localize the object based on Convolutional Neural Networks (CNN). With limited receptive fields, CNN-based CAMs often fail to localize the whole object. The emergence of a Vision Transf…

2024

Phase Concentration and Shortcut Suppression for Weakly Supervised Semantic Segmentation

ECCV 2024poster

"Weakly Supervised Semantic Segmentation (WSSS) with image-level supervision typically acquires object localization information from Class Activation Maps (CAMs). While Vision Transformers (ViTs) in WSSS have been increasingly explored for their superior performance in understanding global context,…

2024

T4P: Test-Time Training of Trajectory Prediction via Masked Autoencoder and Actor-specific Token Memory

CVPR 2024poster

Trajectory prediction is a challenging problem that requires considering interactions among multiple actors and the surrounding environment. While data-driven approaches have been used to address this complex problem they suffer from unreliable predictions under distribution shifts during test time.…

2024

Towards Real-world Event-guided Low-light Video Enhancement and Deblurring

ECCV 2024poster

"In low-light conditions, capturing videos with frame-based cameras often requires long exposure times, resulting in motion blur and reduced visibility. While frame-based motion deblurring and low-light enhancement have been studied, they still pose significant challenges. Event cameras have emerged…

2019

SpherePHD: Applying CNNs on a Spherical PolyHeDron Representation of 360deg Images

CVPR 2019poster

Omni-directional cameras have many advantages overconventional cameras in that they have a much wider field-of-view (FOV). Accordingly, several approaches have beenproposed recently to apply convolutional neural networks(CNNs) to omni-directional images for various visual tasks.However, most of…

Cited by 130PDFScholar