← Search

Junha Hyung

10 accepted papers

2026

ACG: Action Coherence Guidance for Flow-Based Vision-Language-Action Models

ICRA 2026poster

Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstration…

2026

EgoX: Egocentric Video Generation from a Single Exocentric Video

CVPR 2026

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive understanding but remains highly challenging due to extreme c

Cited by 0SourcecodeScholar
2025

Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs

EMNLP 2025

Masked diffusion models (MDMs) offer a promising non-autoregressive alternative for large language modeling. Standard decoding methods for MDMs, such as confidence-based sampling, select tokens independently based on individual token confidences at each diffusion step. However, we observe that this

Cited by 0SourcePDFScholar
2025

Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling

CVPR 2025poster

Diffusion models have emerged as a powerful tool for generating high-quality images, videos, and 3D content. While sampling guidance techniques like CFG improve quality, they reduce diversity and motion. Autoguidance mitigates these issues but demands extra weak model training, limiting its practica…

2025

SurFhead: Affine Rig Blending for Geometrically Accurate 2D Gaussian Surfel Head Avatars

ICLR 2025poster

Recent advancements in head avatar rendering using Gaussian primitives have achieved significantly high-fidelity results. Although precise head geometry is crucial for applications like mesh reconstruction and relighting, current methods struggle to capture intricate geometric details and render uns…

Cited by 1SourcePDFScholar
2025

Temporal In‑Context Fine‑Tuning for Versatile Control of Video Diffusion Models

NeurIPS 2025poster

Recent advances in text-to-video diffusion models have enabled high-quality video synthesis, but controllable generation remains challenging—particularly under limited data and compute. Existing fine-tuning methods often rely on external encoders or architectural modifications, which demand large da…

Cited by 0SourceScholar
2024

Effective Rank Analysis and Regularization for Enhanced 3D Gaussian Splatting

NeurIPS 2024poster

3D reconstruction from multi-view images is one of the fundamental challenges in computer vision and graphics. Recently, 3D Gaussian Splatting (3DGS) has emerged as a promising technique capable of real-time rendering with high-quality 3D reconstruction. This method utilizes 3D Gaussian representati…

2023

FaceCLIPNeRF: Text-driven 3D Face Manipulation using Deformable Neural Radiance Fields

ICCV 2023poster

As recent advances in Neural Radiance Fields (NeRF) have enabled high-fidelity 3D face reconstruction and novel view synthesis, its manipulation also became an essential task in 3D vision. However, existing manipulation methods require extensive human labor, such as a user-provided semantic mask and…

Cited by 16PDFcodeScholar
2023

Local 3D Editing via 3D Distillation of CLIP Knowledge

CVPR 2023poster

3D content manipulation is an important computer vision task with many real-world applications (e.g., product design, cartoon generation, and 3D Avatar editing). Recently proposed 3D GANs can generate diverse photo-realistic 3D-aware contents using Neural Radiance fields (NeRF). However, manipulatio…

Cited by 29SourcePDFScholar