← Search

Shiyu Zhou

6 accepted papers

2026

Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition

AAAI 2026technical

Automatic speech recognition (ASR) systems have achieved remarkable performance in common conditions but often struggle to leverage long-context information in contextualized scenarios that require domain-specific knowledge, such as conference presentations. This challenge arises primarily due to co

Cited by 0SourcePDFScholar
2025

Learning Robust and Flexible Locomotion of Wheel-Legged Quadruped Robots in Complex Terrains

IROS 2025

The wheel-legged quadruped robot, equipped with leg and end-wheel structures, possesses the capability to traverse continuous surfaces at relatively high speeds while also being able to navigate unstructured terrains. However, designing its controller using traditional methods presents significant c

Cited by 0SourceScholar
2023

Load Awareness: Sensorless Body Payload Sensing and Localization for Heavy Quadruped Robot

IROS 2023poster

Heavy quadrupedal drives have great potential for overcoming obstacles, showing great possibilities for transportation industries in complex environments. Ground reaction force (GRF) is a crucial state variable for quadrupedal control. Most GRF observations are implemented in lightweight quadrupeds,…

Cited by 2SourceScholar
2022

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

ICASSP 2022accepted

Nowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between simil…

Cited by 59SourceScholar
2021

Cif-Based Collaborative Decoding for End-to-End Contextual Speech Recognition

ICASSP 2021accepted

End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure and the E2E training hamper injecting context information into them for contextual biasing. Though contextual LAS (CLAS)…

Cited by 0SourceScholar