← Search

Ruizhi Li

6 accepted papers

2026

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

ICLR 2026poster

Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguis…

Cited by 0SourcecodeScholar
2022

An Exact Algorithm with New Upper Bounds for the Maximum k-Defective Clique Problem in Massive Sparse Graphs

AAAI 2022technical

The Maximum k-Defective Clique Problem (MDCP), as a clique relaxation model, has been used to solve various problems. Because it is a hard computational task, previous works can hardly solve the MDCP for massive sparse graphs derived from real-world applications. In this work, we propose a novel bra…

2020

A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech Recognition

ICASSP 2020accepted

The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study offered a promising direction within end-to-end automatic speech recognition, where parallel encoders aim to capture dive…

Cited by 0SourceScholar
2019

Deriving Spectro-temporal Properties of Hearing from Speech Data

ICASSP 2019accepted

Human hearing and human speech are intrinsically tied together, as the properties of speech almost certainly developed in order to be heard by human ears. As a result of this connection, it has been shown that certain properties of human hearing are mimicked within data-driven systems that are train…

Cited by 0SourceScholar
2019

M-vectors: Sub-band Based Energy Modulation Features for Multi-stream Automatic Speech Recognition

ICASSP 2019accepted

In this paper, we propose a novel method to capture energy modulations from different frequency bands in speech into frame-level feature vectors, called Modulation-vectors or M-vectors, for use in Automatic Speech Recognition (ASR) systems. We show that in different multi-stream setups, with paralle…

Cited by 0SourceScholar
2019

Stream Attention-based Multi-array End-to-end Speech Recognition

ICASSP 2019accepted

Automatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array shares and contributes is crucial in this task. Motivated by the advances of joint Connectionist Temporal Classification…

Cited by 21SourceScholar