← Search

Tianyu Zhao

19 accepted papers

2026

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

ICLR 2026poster

Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing E2E approaches primarily fall into two categories: (1) Methods that genera…

Cited by 0SourceScholar
2026

Knee-Inspired Hinge Absorbs Longitudinal Impacts to Enhance Robot-Environment Interaction Safety

ICRA 2026poster

As robots integrate into human society, safe robot-environment interaction has emerged as a growing priority. A promising solution is introducing compliance to existing robots, akin to musculoskeletal systems, to absorb impacts. However, mimicking longitudinal compliance in biological joints remains…

Cited by 0SourceScholar
2026

TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration

CVPR 2026

Diffusion models achieve strong generative performance but remain slow at inference due to the need for repeated full-model denoising passes. We present Token-Adaptive Predictor (TAP), a training-free, probe-driven framework that adaptively selects a predictor for each token at every sampling step.

Cited by 0SourceScholar
2025

Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice

ICASSP 2025accepted

LLM-based text-to-speech(TTS) system has becoming the new trend and SOTA due to its high naturalness and zero-shot capability. However, it relies heavily on training data, usually requires at least thousands hours of labeled audio. In this report, we describe how to use pretrained CosyVoice model, t…

Cited by 0SourceScholar
2025

Fast Adaptation of Pretrained Speaker Verification System for Source Speaker Tracking

ICASSP 2025accepted

Traditional speaker verification system aims at distinguish speaker identity in real world audio, and has achieved satisfying performance in many scenarios. However, it is also very vulnerable, and can be easily attacked by voice anonymization system. In this report, we describe how to fast adapt a…

Cited by 0SourceScholar
2025

Universal Online Temporal Calibration for Optimization-Based Visual-Inertial Navigation Systems

ICRA 2025

6-Degree of Freedom (6DoF) motion estimation with a combination of visual and inertial sensors is a growing area with numerous real-world applications. However, precise calibration of the time offset between these two sensor types is a prerequisite for accurate and robust tracking. To address this,

Cited by 0SourcecodeScholar
2024

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

ACL 2024findings

Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, mo…

2024

Release of Pre-Trained Models for the Japanese Language

COLING 2024main

AI democratization aims to create a world in which the average person can utilize AI techniques. To achieve this goal, numerous research institutes have attempted to make their results accessible to the public. In particular, large pre-trained models trained on large-scale data have shown unpreceden…

Cited by 18SourcePDFScholar
2024

Rethinking the Power of Graph Canonization in Graph Representation Learning with Stability

ICLR 2024poster

The expressivity of Graph Neural Networks (GNNs) has been studied broadly in recent years to reveal the design principles for more powerful GNNs. Graph canonization is known as a typical approach to distinguish non-isomorphic graphs, yet rarely adopted when developing expressive GNNs. This paper pro…

Cited by 8SourcePDFScholar
2024

SchurVINS: Schur Complement-Based Lightweight Visual Inertial Navigation System

CVPR 2024poster

Accuracy and computational efficiency are the most important metrics to Visual Inertial Navigation System (VINS). The existing VINS algorithms with either high accuracy or low computational complexity are difficult to provide the high precision localization in resource-constrained devices. To this e…

2023

Ensuring DNN Solution Feasibility for Optimization Problems with Linear Constraints

ICLR 2023top-25%

We propose preventive learning as the first framework to guarantee Deep Neural Network (DNN) solution feasibility for optimization problems with linear constraints without post-processing, upon satisfying a mild condition on constraint calibration. Without loss of generality, we focus on problems wi…

Cited by 13SourcePDFScholar
2023

Focused Prefix Tuning for Controllable Text Generation

ACL 2023short

In a controllable text generation dataset, there exist unannotated attributes that could provide irrelevant learning signals to models that use it for training and thus degrade their performance. We propose focused prefix tuning (FPT) to mitigate the problem and to enable the control to focus on the…

Cited by 10SourcePDFScholar
2021

CORSAIR: Convolutional Object Retrieval and Symmetry-AIded Registration

IROS 2021poster

This paper considers online object-level mapping using partial point-cloud observations obtained online in an unknown environment. We develop an approach for fully Convolutional Object Retrieval and Symmetry-AIded Registration (CORSAIR). Our model extends the Fully Convolutional Geo-metric Features…

Cited by 10SourceScholar
2020

Topic-relevant Response Generation using Optimal Transport for an Open-domain Dialog System

COLING 2020main

Conventional neural generative models tend to generate safe and generic responses which have little connection with previous utterances semantically and would disengage users in a dialog system. To generate relevant responses, we propose a method that employs two types of constraints - topical const…

Cited by 7SourcePDFScholar