← Search

Jen-Hao Rick Chang

17 accepted papers

2026

Velox: Learning Representations of 4D Geometry and Appearance

CVPR 2026

We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud, to construct. Speci

Cited by 0SourceScholar
2025

Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization

NAACL 2025findings

In this work, we propose Mutual Reinforcing Data Synthesis (MRDS) within LLMs to improve few-shot dialogue summarization task. Unlike prior methods that require external knowledge, we mutually reinforce the LLM’s dialogue synthesis and summarization capabilities, allowing them to complement each oth…

Cited by 0SourcePDFScholar
2024

Corpus Synthesis for Zero-Shot ASR Domain Adaptation Using Large Language Models

ICASSP 2024accepted

While Automatic Speech Recognition (ASR) systems are widely used in many real-world applications, they often do not generalize well to new domains and need to be fine-tuned on data from these domains. However, target-domain data usually are not readily available in many scenarios. In this paper, we…

Cited by 0SourceScholar
2024

Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum

NeurIPS 2024poster

Large language models (LLMs) are commonly trained on datasets consisting of fixed-length token sequences. These datasets are created by randomly concatenating documents of various lengths and then chunking them into sequences of a predetermined target length (concat-and-chunk). Recent attention impl…

2024

Efficient-3Dim: Learning a Generalizable Single-image Novel-view Synthesizer in One Day

ICLR 2024poster

The task of novel view synthesis aims to generate unseen perspectives of an object or scene from a limited set of input images. Nevertheless, synthesizing novel views from a single image remains a significant challenge. Previous approaches tackle this problem by adopting mesh prediction, multi-plane…

Cited by 0SourcePDFScholar
2024

HUGS: Human Gaussian Splats

CVPR 2024poster

Recent advances in neural rendering have improved both training and rendering times by orders of magnitude. While these methods demonstrate state-of-the-art quality and speed they are designed for photogrammetry of static scenes and do not generalize well to freely moving humans in the environment.…

2024

Probabilistic Speech-Driven 3D Facial Motion Synthesis: New Benchmarks Methods and Applications

CVPR 2024poster

We consider the task of animating 3D facial geometry from speech signal. Existing works are primarily deterministic focusing on learning a one-to-one mapping from speech signal to 3D face meshes on small datasets with limited speakers. While these models can achieve high-quality lip articulation for…

Cited by 14SourcePDFScholar
2023

Pointersect: Neural Rendering With Cloud-Ray Intersection

CVPR 2023poster

We propose a novel method that renders point clouds as if they are surfaces. The proposed method is differentiable and requires no scene-specific optimization. This unique capability enables, out-of-the-box, surface normal estimation, rendering room-scale point clouds, inverse rendering, and ray tra…

Cited by 20SourcePDFScholar
2023

Text is all You Need: Personalizing ASR Models Using Controllable Speech Synthesis

ICASSP 2023accepted

Adapting generic speech recognition models to specific individuals is a challenging problem due to the scarcity of personalized data. Recent works have proposed boosting the amount of training data using personalized text-to-speech synthesis. Here, we ask two fundamental questions about this strateg…

Cited by 0SourceScholar
2022

Data Incubation - Synthesizing Missing Data for Handwriting Recognition

ICASSP 2022accepted

In this paper, we demonstrate how a generative model can be used to build a better recognizer through the control of content and style. We are building an online handwriting recognizer from a modest amount of training samples. By training our controllable handwriting synthesizer on the same data, we…

Cited by 0SourceScholar
2022

SYNT++: Utilizing Imperfect Synthetic Data to Improve Speech Recognition

ICASSP 2022accepted

With recent advances in speech synthesis, synthetic data is becoming a viable alternative to real data for training speech recognition models. However, machine learning with synthetic data is not trivial due to the gap between the synthetic and the real data distributions. Synthetic datasets may con…

Cited by 0SourceScholar
2022

Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models

ICML 2022spotlight

Controllable generative sequence models with the capability to extract and replicate the style of specific examples enable many applications, including narrating audiobooks in different voices, auto-completing and auto-correcting written handwriting, and generating missing training samples for downs…

Cited by 26SourcePDFScholar
2021

SapAugment: Learning A Sample Adaptive Policy for Data Augmentation

ICASSP 2021accepted

Data augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a Normal distribution with a fixed standard deviation, for all samples. We hypothesize that a hard sample with high trainin…

Cited by 0SourceScholar