← Search

Haotian Jiang

10 accepted papers

2026

The Effect of Attention Head Count on Transformer Approximation

ICLR 2026poster

Transformer has become the dominant architecture for sequence modeling, yet a detailed understanding of how its structural parameters influence expressive power remains limited. In this work, we study the approximation properties of transformers, with particular emphasis on the role of the number of…

Cited by 0SourceScholar
2025

DCIM-AVSR: Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module

ICASSP 2025accepted

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recog…

Cited by 0SourceScholar
2025

MITracker: Multi-View Integration for Visual Object Tracking

CVPR 2025highlight

Multi-view object tracking (MVOT) offers promising solutions to challenges such as occlusion and target loss, which are common in traditional single-view tracking. However, progress has been limited by the lack of comprehensive multi-view datasets and effective cross-view integration methods. To ove…

Cited by 0SourcePDFScholar
2024

Differentially Private Synthetic Data via Foundation Model APIs 2: Text

ICML 2024spotlight

Text data has become extremely valuable due to the emergence of machine learning algorithms that learn from it. A lot of high-quality text data generated in the real world is private and therefore cannot be shared or used freely due to privacy concerns. Generating synthetic replicas of private text…

2022

Decomposable Non-Smooth Convex Optimization with Nearly-Linear Gradient Oracle Complexity

NeurIPS 2022accept

Many fundamental problems in machine learning can be formulated by the convex program \[ \min_{\theta\in \mathbb{R}^d}\ \sum_{i=1}^{n}f_{i}(\theta), \] where each $f_i$ is a convex, Lipschitz function supported on a subset of $d_i$ coordinates of $\theta$. One common approach to this problem, exemp…

Cited by 3SourcePDFScholar
2022

On the approximation properties of recurrent encoder-decoder architectures

ICLR 2022spotlight

Encoder-decoder architectures have recently gained popularity in sequence to sequence modelling, featuring in state-of-the-art models such as transformers. However, a mathematical understanding of their working principles still remains limited. In this paper, we study the approximation properties of…

Cited by 7SourcePDFScholar
2021

Approximation Theory of Convolutional Architectures for Time Series Modelling

ICML 2021spotlight

We study the approximation properties of convolutional architectures applied to time series modelling, which can be formulated mathematically as a functional approximation problem. In the recurrent setting, recent results reveal an intricate connection between approximation efficiency and memory str…

Cited by 13SourcePDFScholar