← Search

Jooyeol Yun

7 accepted papers

2026

ACG: Action Coherence Guidance for Flow-Based Vision-Language-Action Models

ICRA 2026poster

Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstration…

2026

Selectively Extracting and Injecting Visual Attributes into Text-to-Image Models

CVPR 2026

Text-to-image models are increasingly utilized in design workflows, but articulating nuanced design intentions solely through text remains a challenge. This work proposes a method that extracts visual attributes from a reference image and injects them directly into the generation pipeline. Specifica

Cited by 0SourceScholar
2026

SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation

AAAI 2026technical

The increasing demand for AR/VR applications has highlighted the need for high-quality content, such as 360° live wallpapers. However, generating high-quality 360° panoramic contents remains a challenging task due to the severe distortions introduced by equirectangular projection (ERP). Existing ap

Cited by 0SourcePDFScholar
2025

Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention

CVPR 2025poster

While large-scale text-to-image diffusion models enable the generation of high-quality, diverse images from text prompts, these prompts struggle to capture intricate details, such as textures, preventing the user intent from being reflected. This limitation has led to efforts to generate images cond…

Cited by 0SourcePDFScholar
2025

Enabling Region-Specific Control via Lassos in Point-Based Colorization

AAAI 2025technical

Point-based interactive colorization techniques allow users to effortlessly colorize grayscale images using user-provided color hints. However, point-based methods often face challenges when different colors are given to semantically similar areas, leading to color intermingling and unsatisfactory r…

Cited by 0SourcePDFScholar
2023

Learning to Generate Semantic Layouts for Higher Text-Image Correspondence in Text-to-Image Synthesis

ICCV 2023poster

Existing text-to-image generation approaches have set high standards for photorealism and text-image correspondence, largely benefiting from web-scale text-image datasets, which can include up to 5 billion pairs. However, text-to-image generation models trained on domain-specific datasets, such as u…

Cited by 12PDFcodeScholar