← Search

Kyeongbo Kong

10 accepted papers

2026

An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models

ICLR 2026poster

Large Vision-Language Models (LVLMs) have adopted visual token pruning strategies to mitigate substantial computational overhead incurred by extensive visual token sequences. While prior works primarily focus on either attention-based or diversity-based pruning methods, in-depth analysis of these a…

Cited by 0SourcecodeScholar
2026

MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting

ICLR 2026poster

Recent advances in dynamic scene reconstruction have significantly benefited from 3D Gaussian Splatting, yet existing methods show inconsistent performance across diverse scenes, indicating no single approach effectively handles all dynamic challenges. To overcome these limitations, we propose Mixtu…

Cited by 0SourceScholar
2026

PADS-TAL: Padding-Annealed Diffusion Sampling in Text-Aware Latent Space for Robust and Diverse Text-to-Music Generation

ICML 2026poster

Text-to-Music diffusion models are increasingly used in real-world applications, yet deployment remains challenging: generations can collapse to limited patterns even with diverse initial noise and prompts, and inference-time diversity control often harms text alignment and fidelity by distorting ke…

Cited by 0SourcecodeScholar
2025

Optimizing 4D Gaussians for Dynamic Scene Video from Single Landscape Images

ICLR 2025poster

To achieve realistic immersion in landscape images, fluids such as water and clouds need to move within the image while revealing new scenes from various camera perspectives. Recently, a field called dynamic scene video has emerged, which combines single image animation with 3D photography. These me…

2024

AttentionHand: Text-driven Controllable Hand Image Generation for 3D Hand Reconstruction in the Wild

ECCV 2024oral

"Recently, there has been a significant amount of research conducted on 3D hand reconstruction to use various forms of human-computer interaction. However, 3D hand reconstruction in the wild is challenging due to extreme lack of in-the-wild 3D hand datasets. Especially, when hands are in complex pos…

2024

Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation

ECCV 2024poster

"We present an Encoder-Decoder Attention Transformer, ED-AFormer, which consists of the Embedding-Free Transformer (EFT) encoder and the all-attention decoder leveraging our Embedding-Free Attention (EFA) structure. The proposed EFA is a novel global context modeling mechanism that focuses on functi…

2024

Person in Place: Generating Associative Skeleton-Guidance Maps for Human-Object Interaction Image Editing

CVPR 2024poster

Recently there were remarkable advances in image editing tasks in various ways. Nevertheless existing image editing models are not designed for Human-Object Interaction (HOI) image editing. One of these approaches (e.g. ControlNet) employs the skeleton guidance to offer precise representations of hu…

2023

FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation

AAAI 2023technical

With the success of Vision Transformer (ViT) in image classification, its variants have yielded great success in many downstream vision tasks. Among those, the semantic segmentation task has also benefited greatly from the advance of ViT variants. However, most studies of the transformer for semanti…

2023

SEFD: Learning to Distill Complex Pose and Occlusion

ICCV 2023poster

This paper addresses the problem of three-dimensional (3D) human mesh estimation in complex poses and occluded situations. Although many improvements have been made in 3D human mesh estimation using the two-dimensional (2D) pose with occlusion between humans, occlusion from complex poses and other o…

Cited by 13PDFcodeScholar
2022

Selective TransHDR: Transformer-Based Selective HDR Imaging Using Ghost Region Mask

ECCV 2022poster

"The primary issue in high dynamic range (HDR) imaging is the removal of ghost artifacts afforded when merging multi-exposure low dynamic range images. In the weakly misaligned region, ghost artifacts can be suppressed using convolutional neural network (CNN)-based methods. However, in highly misali…

Cited by 29SourcePDFScholar