← Search

Jae-Pil Heo

43 accepted papers

2026

Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer Normalization

ICLR 2026poster

This work analyzes the training dynamics of Image Restoration (IR) Transformers and uncovers a critical yet overlooked issue: conventional LayerNorm (LN) drives feature magnitudes to diverge to a million scale and collapses channel-wise entropy. We analyze this in the perspective of networks attempt…

Cited by 0SourcecodeScholar
2026

From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning

ICLR 2026poster

Unsupervised object-centric learning models, particularly slot-based architectures, have shown great promise in decomposing complex scenes. However, their reliance on reconstruction-based training creates a fundamental conflict between the sharp, high-frequency attention maps of the encoder and the…

Cited by 0SourcecodeScholar
2026

Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation

CVPR 2026

A sliding-window inference strategy is commonly adopted in recent training-free open-vocabulary semantic segmentation methods to overcome limitation of the CLIP in processing high-resolution images. However, this approach introduces a new challenge: each window is processed independently, leading to

Cited by 0SourcecodeScholar
2026

Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding

CVPR 2026

Recent advances in 3D scene-language understanding have leveraged Large Language Models (LLMs) for 3D reasoning by transferring their general reasoning ability to 3D multi-modal contexts. However, existing methods typically adopt standard decoders from language modeling, which rely on a causal atten

Cited by 0SourcecodeScholar
2026

Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning

CVPR 2026

Video Object-Centric Learning seeks to decompose raw videos into a small set of object slots, but existing slot-attention models often suffer from severe over-fragmentation. This is because the model is implicitly encouraged to occupy all slots to minimize the reconstruction objective, thereby repre

Cited by 0SourcecodeScholar
2026

SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Models

CVPR 2026

Diffusion models are a strong backbone for visual generation, but their inherently sequential denoising process leads to slow inference. Previous methods accelerate sampling by caching and reusing intermediate outputs based on feature distances between adjacent timesteps. However, existing caching s

Cited by 0SourcecodeScholar
2025

Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval

AAAI 2025technical

Partially Relevant Video Retrieval~(PRVR) aims to retrieve a video where a specific segment is relevant to a given text query. Typical training processes of PRVR assume a one-to-one relationship where each text query is relevant to only one video. However, we point out the inherent ambiguity between…

Cited by 0SourcePDFScholar
2025

Auto-Encoded Supervision for Perceptual Image Super-Resolution

CVPR 2025poster

This work tackles the fidelity objective in the perceptual super-resolution (SR) task. Specifically, we address the shortcomings of pixel-level \mathcal L _\text p loss (\mathcal L _\text pix ) in the GAN-based SR framework. Since \mathcal L _\text pix is known to have a trade-off relationship aga…

2025

Bridging the Semantic Granularity Gap Between Text and Frame Representations for Partially Relevant Video Retrieval

AAAI 2025technical

Partially Relevant Video Retrieval (PRVR) addresses the challenges of text-to-video retrieval in real-world scenarios where untrimmed videos are prevalent. Traditional PRVR methods encode videos at two feature scales: (1) frame-level to capture fine details, and (2) clip-level to recognize broader c…

Cited by 0SourcePDFScholar
2025

Diffusion Feature Field for Text-based 3D Editing with Gaussian Splatting

NeurIPS 2025poster

Recent advances in text-based image editing have motivated the extension of these techniques into the 3D domain. However, existing methods typically apply 2D diffusion models independently to multiple viewpoints, resulting in significant artifacts, most notably the Janus problem, due to inconsisten…

Cited by 0SourceScholar
2025

Fine-Tuning Visual Autogressive Models for Subject-Driven Generation

ICCV 2025poster

Recent advances in text-to-image generative models have enabled numerous practical applications, including subject-driven generation, which fine-tunes pre-trained models to capture subject semantics from only a few examples. While diffusion-based models produce high-quality images, their extensive d…

2025

Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot Segmentation

AAAI 2025technical

We propose Foreground-Covering Prototype Generation and Matching to resolve Few-Shot Segmentation (FSS), which aims to segment target regions in unlabeled query images based on labeled support images. Unlike previous research, which typically estimates target regions in the query using support proto…

2025

Mitigating Semantic Collapse in Partially Relevant Video Retrieval

NeurIPS 2025poster

Partially Relevant Video Retrieval (PRVR) seeks videos where only part of the content matches a text query. Existing methods treat every annotated text–video pair as a positive and all others as negatives, ignoring the rich semantic variation both within a single video and across different videos.…

Cited by 4SourceScholar
2025

Prediction-Feedback DETR for Temporal Action Detection

AAAI 2025technical

Temporal Action Detection (TAD) is fundamental yet challenging for real-world video applications. Leveraging the unique benefits of transformers, various DETR-based approaches have been adopted in TAD. However, it has recently been identified that the attention collapse in self-attention causes the…

Cited by 1SourcePDFScholar
2025

Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval

ICCV 2025poster

In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhance…

Cited by 0SourcePDFScholar
2025

Selective Contrastive Learning for Weakly Supervised Affordance Grounding

ICCV 2025poster

Facilitating an entity's interaction with objects requires accurately identifying parts that afford specific actions. Weakly supervised affordance grounding (WSAG) seeks to imitate human learning from third-person demonstrations, where humans intuitively grasp functional parts without needing pixel-…

Cited by 0SourcePDFScholar
2025

Temporal Alignment-Free Video Matching for Few-shot Action Recognition

CVPR 2025poster

Few-Shot Action Recognition (FSAR) aims to train a model with only a few labeled video instances. A key challenge in FSAR is handling divergent narrative trajectories for precise video matching. While the frame- and tuple-level alignment approaches have been promising, their methods heavily rely on…

2025

Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion Models

ICCV 2025poster

Text-to-Image (T2I) diffusion models have made significant progress in generating diverse high-quality images from textual prompts. However, these models still face challenges in suppressing content that is strongly entangled with specific words. For example, when generating an image of "Charlie Cha…

2024

Diversity-aware Channel Pruning for StyleGAN Compression

CVPR 2024poster

StyleGAN has shown remarkable performance in unconditional image generation. However its high computational cost poses a significant challenge for practical applications. Although recent efforts have been made to compress StyleGAN while preserving its performance existing compressed models still lag…

2024

GSGAN: Adversarial Learning for Hierarchical Generation of 3D Gaussian Splats

NeurIPS 2024poster

Most advances in 3D Generative Adversarial Networks (3D GANs) largely depend on ray casting-based volume rendering, which incurs demanding rendering costs. One promising alternative is rasterization-based 3D Gaussian Splatting (3D-GS), providing a much faster rendering speed and explicit 3D represen…

2024

Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models for Style Transfer

CVPR 2024highlight

Despite the impressive generative capabilities of diffusion models existing diffusion model-based style transfer methods require inference-stage optimization (e.g. fine-tuning or textual inversion of style) which is time-consuming or fails to leverage the generative ability of large-scale diffusion…

2024

Task-Disruptive Background Suppression for Few-Shot Segmentation

AAAI 2024technical

Few-shot segmentation aims to accurately segment novel target objects within query images using only a limited number of annotated support images. The recent works exploit support background as well as its foreground to precisely compute the dense correlations between query and support. However, the…

2024

Towards Squeezing-Averse Virtual Try-On via Sequential Deformation

AAAI 2024technical

In this paper, we first investigate a visual quality degradation problem observed in recent high-resolution virtual try-on approach. The tendency is empirically found that the textures of clothes are squeezed at the sleeve, as visualized in the upper row of Fig.1(a). A main reason for the issue aris…

2024

VLCounter: Text-Aware Visual Representation for Zero-Shot Object Counting

AAAI 2024technical

Zero-Shot Object Counting~(ZSOC) aims to count referred instances of arbitrary classes in a query image without human-annotated exemplars. To deal with ZSOC, preceding studies proposed a two-stage pipeline: discovering exemplars and counting. However, there remains a challenge of vulnerability to er…

2023

Disentangled Representation Learning for Unsupervised Neural Quantization

CVPR 2023poster

The inverted index is a widely used data structure to avoid the infeasible exhaustive search. It accelerates retrieval significantly by splitting the database into multiple disjoint sets and restricts distance computation to a small fraction of the database. Moreover, it even improves search quality…

Cited by 3SourcePDFScholar
2023

Leveraging Hidden Positives for Unsupervised Semantic Segmentation

CVPR 2023poster

Dramatic demand for manpower to label pixel-level annotations triggered the advent of unsupervised semantic segmentation. Although the recent work employing the vision transformer (ViT) backbone shows exceptional performance, there is still a lack of consideration for task-specific training guidance…

2023

Minority-Oriented Vicinity Expansion with Attentive Aggregation for Video Long-Tailed Recognition

AAAI 2023technical

A dramatic increase in real-world video volume with extremely diverse and emerging topics naturally forms a long-tailed video distribution in terms of their categories, and it spotlights the need for Video Long-Tailed Recognition (VLTR). In this work, we summarize the challenges in VLTR and explore…

2023

Progressive Few-Shot Adaptation of Generative Model with Align-Free Spatial Correlation

AAAI 2023technical

In few-shot generative model adaptation, the model for target domain is prone to the mode-collapse. Recent studies attempted to mitigate the problem by matching the relationship among samples generated from the same latent codes in source and target domains. The objective is further extended to imag…

Cited by 3SourcePDFScholar
2023

Query-Dependent Video Representation for Moment Retrieval and Highlight Detection

CVPR 2023poster

Recently, video moment retrieval and highlight detection (MR/HD) are being spotlighted as the demand for video understanding is drastically increased. The key objective of MR/HD is to localize the moment and estimate clip-wise accordance level, i.e., saliency score, to the given text query. Although…

2023

Robust Image Denoising of No-Flash Images Guided by Consistent Flash Images

AAAI 2023technical

Images taken in low light conditions typically contain distracting noise, and eliminating such noise is a crucial computer vision problem. Additional photos captured with a camera flash can guide an image denoiser to preserve edges since the flash images often contain fine details with reduced noise…

2022

Difficulty-Aware Simulator for Open Set Recognition

ECCV 2022poster

"Open set recognition (OSR) assumes unknown instances appear out of the blue at the inference time. The main challenge of OSR is that the response of models for unknowns is totally unpredictable. Furthermore, the diversity of open set makes it harder since instances have different difficulty levels.…

2021

Self-Supervised Video GANs: Learning for Appearance Consistency and Motion Coherency

CVPR 2021poster

A video can be represented by the composition of appearance and motion. Appearance (or content) expresses the information invariant throughout time, and motion describes the time-variant movement. Here, we propose self-supervised approaches for video Generative Adversarial Networks (GANs) to achieve…

Cited by 25PDFScholar
2016

Shortlist Selection With Residual-Aware Distance Estimator for K-Nearest Neighbor Search

CVPR 2016poster

In this paper, we introduce a novel shortlist computation algorithm for approximate, high-dimensional nearest neighbor search. Our method relies on a novel distance estimator: the residual-aware distance estimator, that accounts for the residual distances of data points to their respective quantized…

Cited by 13PDFScholar