← Search

Gyeong-Moon Park

24 accepted papers

2026

AudioAvatar: Personalized Audio-driven Whole-body Talking Avatars

CVPR 2026

Prior expressive whole-body conversational avatar systems map audio to parametric poses and then render, creating a lossy bottleneck where quantization, retargeting, and tracking errors accumulate. This degrades audio-motion synchronization and suppresses micro-articulations critical for realism--su

Cited by 0SourceScholar
2026

Detecting Unknown Objects via Energy-based Separation for Open World Object Detection

CVPR 2026

In this work, we tackle the problem of Open World Object Detection (OWOD). This challenging scenario requires the detector to incrementally learn to classify known objects without forgetting while identifying unknown objects without supervision. Previous OWOD methods have enhanced the unknown discov

Cited by 0SourceScholar
2026

Dynamic Texture Modeling of 3D Clothed Gaussian Avatars from a Single Video

ICLR 2026poster

Recent advances in neural rendering, particularly 3D Gaussian Splatting (3DGS), have enabled animatable 3D human avatars from single videos with efficient rendering and high fidelity. However, current methods struggle with dynamic appearances, especially in loose garments (e.g., skirts), causing unr…

Cited by 0SourceScholar
2026

Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editing

CVPR 2026

Editing a 3D indoor scene from natural language is conceptually straightforward but technically challenging. Existing open-vocabulary systems often regenerate large portions of a scene or rely on image-space edits that disrupt spatial structure, resulting in unintended global changes or physically i

Cited by 0SourceScholar
2025

ConceptSplit: Decoupled Multi-Concept Personalization of Diffusion Models via Token-wise Adaptation and Attention Disentanglement

ICCV 2025poster

In recent years, multi-concept personalization for text-to-image (T2I) diffusion models to represent several subjects at an image has gained much more attention. The main challenge of this task is "concept mixing", where multiple learned concepts interfere or blend undesirably in the output image. T…

2025

Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition

NeurIPS 2025spotlight

Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods—based on saliency—produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Langu…

Cited by 0SourceScholar
2025

Do Not Mimic My Voice : Speaker Identity Unlearning for Zero-Shot Text-to-Speech

ICML 2025poster

The rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to selectively remove the knowledge to replicate unwanted individu…

2025

ESC: Erasing Space Concept for Knowledge Deletion

CVPR 2025highlight

As concerns regarding privacy in deep learning continue to grow, individuals are increasingly apprehensive about the potential exploitation of their personal knowledge in trained models. Despite several research efforts to address this, they often fail to consider the real-world demand from users fo…

2025

ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning

ICCV 2025poster

In this work, we tackle the problem of video class-incremental learning (VCIL). Many existing VCIL methods mitigate catastrophic forgetting by rehearsal training with a few temporally dense samples stored in episodic memory, which is memory-inefficient. Alternatively, some methods store temporally s…

Cited by 0SourcePDFScholar
2025

GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head Avatar

ICCV 2025poster

Despite recent progress in 3D head avatar generation, balancing identity preservation, i.e., reconstruction, with novel poses and expressions, i.e., animation, remains a challenge. Existing methods struggle to adapt Gaussians to varying geometrical deviations across facial regions, resulting in subo…

Cited by 4SourcePDFScholar
2025

Multispectral Pedestrian Detection with Sparsely Annotated Label

AAAI 2025technical

Although existing Sparsely Annotated Object Detection (SAOD) approches have made progress in handling sparsely annotated environments in multispectral domain, where only some pedestrians are annotated, they still have the following limitations: (i) they lack considerations for improving the quality…

2025

Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion Models

NeurIPS 2025poster

Recent advances in diffusion models have enabled high-quality synthesis of specific subjects, such as identities or objects. This capability, while unlocking new possibilities in content creation, also introduces significant privacy risks, as personalization techniques can be misused by malicious us…

Cited by 0SourcecodeScholar
2025

Test-Time Fine-Tuning of Image Compression Models for Multi-Task Adaptability

CVPR 2025poster

The field of computer vision was initially inspired by the human visual system and has progressively expanded to include a broader range of machine vision applications. Consequently, image compressors should be designed to effectively accommodate not only human visual perception but also machine vis…

Cited by 0SourcePDFScholar
2025

Universal Domain Adaptation for Semantic Segmentation

CVPR 2025poster

Unsupervised domain adaptation for semantic segmentation (UDA-SS) aims to transfer knowledge from labeled synthetic data (source) to unlabeled real-world data (target). Traditional UDA-SS methods work on the assumption that the category settings between the source and target domains are known in adv…

2025

When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series

ICML 2025poster

Recently, forecasting future abnormal events has emerged as an important scenario to tackle realworld necessities. However, the solution of predicting specific future time points when anomalies will occur, known as Anomaly Prediction (AP), remains under-explored. Existing methods dealing with time s…

Cited by 10SourcePDFScholar
2024

CED: Comparing Embedding Differences for Detecting Out-of-Distribution and Hallucinated Text

EMNLP 2024finding

Detecting out-of-distribution (OOD) samples is crucial for ensuring the safety and robustness of models deployed in real-world scenarios. While most studies on OOD detection focus on fine-tuned models trained on in-distribution (ID) data, detecting OOD in pre-trained models is also important due to…

Cited by 0SourcePDFScholar
2024

Generative Unlearning for Any Identity

CVPR 2024poster

Recent advances in generative models trained on large-scale datasets have made it possible to synthesize high-quality samples across various domains. Moreover the emergence of strong inversion networks enables not only a reconstruction of real-world images but also the modification of attributes thr…

2024

Open-Set Domain Adaptation for Semantic Segmentation

CVPR 2024poster

Unsupervised domain adaptation (UDA) for semantic segmentation aims to transfer the pixel-wise knowledge from the labeled source domain to the unlabeled target domain. However current UDA methods typically assume a shared label space between source and target limiting their applicability in real-wor…

2024

Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners

CVPR 2024poster

Few-Shot Class Incremental Learning (FSCIL) is a task that requires a model to learn new classes incrementally without forgetting when only a few samples for each class are given. FSCIL encounters two significant challenges: catastrophic forgetting and overfitting and these challenges have driven pr…

2023

LINe: Out-of-Distribution Detection by Leveraging Important Neurons

CVPR 2023poster

It is important to quantify the uncertainty of input samples, especially in mission-critical domains such as autonomous driving and healthcare, where failure predictions on out-of-distribution (OOD) data are likely to cause big problems. OOD detection problem fundamentally begins in that the model c…

2023

Online Class Incremental Learning on Stochastic Blurry Task Boundary via Mask and Visual Prompt Tuning

ICCV 2023poster

Continual learning aims to learn a model from a continuous stream of data, but it mainly assumes a fixed number of data and tasks with clear task boundaries. However, in real-world scenarios, the number of input data and tasks is constantly changing in a statistical way, not a static way. Although r…

Cited by 27PDFcodeScholar