← Search

Jun-Kun Chen

10 accepted papers

2025

Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image

NeurIPS 2025poster

This paper proposes Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an auto-regressive, segment-by-segment generation process, eliminating the need for resource-intensive generation and len…

Cited by 0SourcecodeScholar
2024

ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing

CVPR 2024poster

This paper proposes ConsistDreamer - a novel framework that lifts 2D diffusion models with 3D awareness and 3D consistency thus enabling high-fidelity instruction-guided scene editing. To overcome the fundamental limitation of missing 3D consistency in 2D diffusion models our key insight is to intro…

Cited by 9SourcePDFScholar
2024

Diarist: Streaming Speech Translation with Speaker Diarization

ICASSP 2024accepted

End-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and handling of overlapping speech in a streaming fashion. In this work, we propose DiariST, the first streaming ST and SD solu…

Cited by 0SourceScholar
2024

Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion

CVPR 2024poster

This paper proposes Instruct 4D-to-4D that achieves 4D awareness and spatial-temporal consistency for 2D diffusion models to generate high-quality instruction-guided dynamic scene editing results. Traditional applications of 2D diffusion models in dynamic scene editing often result in inconsistency…

Cited by 9SourcePDFScholar
2024

Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation

ICASSP 2024accepted

The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential for user applications. Traditional approaches to automatic speech recognition (AS…

Cited by 0SourceScholar
2023

Contrastive Learning Relies More on Spatial Inductive Bias Than Supervised Learning: An Empirical Study

ICCV 2023poster

Though self-supervised contrastive learning (CL) has shown its potential to achieve state-of-the-art accuracy without any supervision, its behavior still remains under investigated by academia. Different from most previous work that understands CL from learning objectives, we focus on an unexplored…

Cited by 2PDFScholar
2023

NeuralEditor: Editing Neural Radiance Fields via Manipulating Point Clouds

CVPR 2023poster

This paper proposes NeuralEditor that enables neural radiance fields (NeRFs) natively editable for general shape editing tasks. Despite their impressive results on novel-view synthesis, it remains a fundamental challenge for NeRFs to edit the shape of the scene. Our key insight is to exploit the exp…