← Search

Yuzi Yan

5 accepted papers

2025

3D-Properties: Identifying Challenges in DPO and Charting a Path Forward

ICLR 2025poster

Aligning large language models (LLMs) with human preferences has gained significant attention, with Proximal Policy Optimization (PPO) as a standard yet computationally expensive method and Direct Preference Optimization (DPO) as a more efficient alternative. While DPO offers simplicity, it remains…

Cited by 10SourcePDFScholar
2024

Exploring the LLM Journey from Cognition to Expression with Linear Representations

ICML 2024poster

This paper presents an in-depth examination of the evolution and interplay of cognitive and expressive capabilities in large language models (LLMs), with a specific focus on Baichuan-7B and Baichuan-33B, an advanced bilingual (Chinese and English) LLM series. We define and explore the model's cognit…

Cited by 3SourcePDFScholar
2023

Approximation Error Back-Propagation for Q-Function in Scalable Reinforcement Learning with Tree Dependence Structure

ICASSP 2023accepted

This paper applies the exponential decay property of scalable RL theory to a specific scenario where the network structure is a tree, and use KL (Kullback-Leibler) divergence to analyze the propagation of approximation error along the structure over time, in order to quantify its backtracking result…

Cited by 0SourceScholar
2022

Relative Distributed Formation and Obstacle Avoidance with Multi-agent Reinforcement Learning

ICRA 2022poster

Multi-agent formation as well as obstacle avoid-ance is one of the most actively studied topics in the field of multi-agent systems. Although some classic controllers like model predictive control (MPC) and fuzzy control achieve a certain measure of success, most of them require precise global infor…

Cited by 24SourceScholar
2021

Adaspeech 2: Adaptive Text to Speech with Untranscribed Data

ICASSP 2021accepted

Text to speech (TTS) is widely used to synthesize personal voice for a target speaker, where a well-trained source TTS model is fine-tuned with few paired adaptation data (speech and its transcripts) on this target speaker. However, in many scenarios, only untranscribed speech data is available for…

Cited by 0SourceScholar