← Search

Tianhao Zhang

16 accepted papers

2025

BBScoreV2: Learning Time-Evolution and Latent Alignment from Stochastic Representation

EMNLP 2025

Autoregressive generative models play a key role in various language tasks, especially for modeling and evaluating long text sequences. While recent methods leverage stochastic representations to better capture sequence dynamics, encoding both temporal and structural dependencies and utilizing such

2025

In-context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-separation

NeurIPS 2025poster

This paper investigates approximation-theoretic aspects of the in-context learning capability of the transformers in representing a family of noisy linear dynamical systems. Our first theoretical result establishes an upper bound on the approximation error of multi-layer transformers with respect to…

Cited by 0SourceScholar
2024

BBScore: A Brownian Bridge Based Metric for Assessing Text Coherence

AAAI 2024technical

Measuring the coherence of text is a vital aspect of evaluating the quality of written content. Recent advancements in neural coherence modeling have demonstrated their efficacy in capturing entity coreference and discourse relations, thereby enhancing coherence evaluation. However, many existing me…

2023

Self-Convolution for Automatic Speech Recognition

ICASSP 2023accepted

Self-attention plays a significant role in recent automatic speech recognition (ASR) models with promising results. However, it suffers from high computational complexity and weak capability in modeling local information. In contrast, the convolutional neural network (CNN) is computationally effecti…

Cited by 0SourceScholar
2021

Decentralized Circle Formation Control for Fish-like Robots in the Real-world via Reinforcement Learning

ICRA 2021poster

In this paper, the circle formation control problem is addressed for a group of cooperative underactuated fish-like robots involving unknown nonlinear dynamics and disturbances. Based on the reinforcement learning and cognitive consistency theory, we propose a decentralized controller without the kn…

Cited by 27SourceScholar
2021

FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement Learning

ICML 2021spotlight

Value decomposition recently injects vigorous vitality into multi-agent actor-critic methods. However, existing decomposed actor-critic methods cannot guarantee the convergence of global optimum. In this paper, we present a novel multi-agent actor-critic method, FOP, which can factorize the optimal…

Cited by 102SourcePDFScholar
2021

MFVFD: A Multi-Agent Q-Learning Approach to Cooperative and Non-Cooperative Tasks

IJCAI 2021poster

Value function decomposition (VFD) methods under the popular paradigm of centralized training and decentralized execution (CTDE) have promoted multi-agent reinforcement learning progress. However, existing VFD methods proceed from a group's value function decomposition to only solve cooperative task…

Cited by 16SourcePDFScholar
2020

Stacking Networks Dynamically for Image Restoration Based on the Plug-and-Play Framework

ECCV 2020poster

Recently, stacked networks show powerful performance in Image Restoration, such as challenging motion deblurring problems. However, the number of stacking levels is a hyper-parameter fine-tuned manually, making the stacking levels static during training without theoretical explanations for optimal s…

Cited by 12SourcePDFScholar
2018

Deep Imitation Learning for Complex Manipulation Tasks from Virtual Reality Teleoperation

ICRA 2018poster

Imitation learning is a powerful paradigm for robot skill acquisition. However, obtaining demonstrations suitable for learning a policy that maps from raw pixels to actions can be challenging. In this paper we describe how consumer-grade Virtual Reality headsets and hand tracking hardware can be use…

Cited by 895SourceScholar
2018

One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning

RSS 2018poster

Humans and animals are capable of learning a new behavior by observing others perform the skill just once. We consider the problem of allowing a robot to do the same -- learning from a video of a human, even when there is domain shift in the perspective, environment, and embodiment between the robot…

Cited by 435SourcePDFScholar
2017

Learning from the hindsight plan — Episodic MPC improvement

ICRA 2017poster

Model predictive control (MPC) is a popular control method that has proved effective for robotics, among other fields. MPC performs re-planning at every time step. Re-planning is done with a limited horizon per computational and real-time constraints and often also for robustness to potential model…

Cited by 83SourceScholar
2017

PLATO: Policy learning using adaptive trajectory optimization

ICRA 2017poster

Policy search can in principle acquire complex strategies for control of robots and other autonomous systems. When the policy is trained to process raw sensory inputs, such as images and depth maps, it can also acquire a strategy that combines perception and control. However, effectively processing…

Cited by 167SourceScholar
2016

Learning deep control policies for autonomous aerial vehicles with MPC-guided policy search

ICRA 2016

Model predictive control (MPC) is an effective method for controlling robotic systems, particularly autonomous aerial vehicles such as quadcopters. However, application of MPC can be computationally demanding, and typically requires estimating the state of the system, which can be challenging in com

Cited by 442SourceScholar