← Search

Roger Hsiao

9 accepted papers

2026

3DGS$^2$-TR: A Scalable Second-Order Trust-Region Method for 3D Gaussian Splatting

ICML 2026poster

We propose 3DGS$^2$-TR, a second-order optimizer for accelerating the scene training problem in 3D Gaussian Splatting (3DGS). Unlike existing second-order approaches that rely on explicit or dense curvature representations, such as 3DGS-LM (Höllein et al., 2025) or 3DGS2 (Lan et al., 2025), our meth…

Cited by 0SourceScholar
2023

Neural Transducer Training: Reduced Memory Consumption with Sample-Wise Computation

ICASSP 2023accepted

The neural transducer is an end-to-end model for automatic speech recognition (ASR). While the model is well-suited for streaming ASR, the training process remains challenging. During training, the memory requirements may quickly exceed the capacity of state-of-the-art GPUs, limiting batch size and…

Cited by 0SourceScholar
2023

Variable Attention Masking for Configurable Transformer Transducer Speech Recognition

ICASSP 2023accepted

This work studies the use of attention masking in transformer transducer based speech recognition for building a single configurable model for different deployment scenarios. We present a comprehensive set of experiments comparing fixed masking, where the same attention mask is applied at every fram…

Cited by 0SourceScholar
2020

Improving Language Identification for Multilingual Speakers

ICASSP 2020accepted

Spoken language identification (LID) technologies have improved in recent years from discriminating largely distinct languages to discriminating highly similar languages or even dialects of the same language. One aspect that has been mostly neglected, however, is discrimination of languages for mult…

Cited by 0SourceScholar
2017

Analysis of keyword spotting performance across IARPA babel languages

ICASSP 2017accepted

With the completion of the IARPA Babel program, it is possible to systematically analyze the performance of speech recognition systems across a wide variety of languages. We select 16 languages from the dataset and compare performance using a deep neural network-based acoustic model. The focus is on…

Cited by 0SourceScholar
2017

The 2016 BBN Georgian telephone speech keyword spotting system

ICASSP 2017accepted

In this paper we describe the 2016 BBN conversational telephone speech keyword spotting system; the culmination of four years of research and development under the IARPA Babel program. The system was constructed in response to the NIST Open Keyword Search (OpenKWS) evaluation of 2016. We present our…

Cited by 0SourceScholar
2017

Unsupervised adaptation for deep neural networks using Alternating Direction Method of Multipliers

ICASSP 2017accepted

In this paper, we continue our work on linear least squares based adaptation (LLS) for deep neural networks. We show that our previously proposed algorithm is a special case of an optimization algorithm called Alternating Direction Method of Multipliers (ADMM). We demonstrate that the adaptation alg…

Cited by 0SourceScholar