← Search

Zhi Tian

34 accepted papers

2025

Scaling Diffusion Transformers Efficiently via $\mu$P

NeurIPS 2025poster

Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales. Recently, Maximal Update Parametrization ($\mu$P) was proposed for vanilla Transformers, which enables stable HP transf…

Cited by 0SourceScholar
2023

Conditional Positional Encodings for Vision Transformers

ICLR 2023poster

We propose a conditional positional encoding (CPE) scheme for vision Transformers. Unlike previous fixed or learnable positional encodings that are predefined and independent of input tokens, CPE is dynamically generated and conditioned on the local neighborhood of the input tokens. As a result, CPE…

2023

Distributed Online Learning With Adversarial Participants In An Adversarial Environment

ICASSP 2023accepted

This paper studies distributed online learning under Byzantine attacks. The performance of an online learning algorithm is characterized by (adversarial) regret, and a sublinear bound is preferred. But we prove that, even with a class of state-of-the-art robust aggregation rules, in an adversarial e…

Cited by 0SourceScholar
2023

H-nobs: Achieving Certified Fairness and Robustness in Distributed Learning on Heterogeneous Datasets

NeurIPS 2023poster

Fairness and robustness are two important goals in the design of modern distributed learning systems. Despite a few prior works attempting to achieve both fairness and robustness, some key aspects of this direction remain underexplored. In this paper, we try to answer three largely unnoticed and una…

Cited by 7SourcePDFScholar
2022

Fully Convolutional One-Stage 3D Object Detection on LiDAR Range Images

NeurIPS 2022accept

We present a simple yet effective fully convolutional one-stage 3D object detector for LiDAR point clouds of autonomous driving scenes, termed FCOS-LiDAR. Unlike the dominant methods that use the bird-eye view (BEV), our proposed detector detects objects from the range view (RV, a.k.a. range image)…

Cited by 130SourcePDFScholar
2022

Poseur: Direct Human Pose Regression with Transformers

ECCV 2022poster

"We propose a direct, regression-based approach to 2D human pose estimation from single images. We formulate the problem as a sequence prediction task, which we solve using a Transformer network. This network directly learns a regression mapping from images to the keypoint coordinates, without resor…

2022

SegViT: Semantic Segmentation with Plain Vision Transformers

NeurIPS 2022accept

We explore the capability of plain Vision Transformers (ViTs) for semantic segmentation and propose the SegViT. Previous ViT-based segmentation networks usually learn a pixel-level representation from the output of the ViT. Differently, we make use of the fundamental component—attention mechanism, t…

2021

Dynamic Neural Representational Decoders for High-Resolution Semantic Segmentation

NeurIPS 2021poster

Semantic segmentation requires per-pixel prediction for a given image. Typically, the output resolution of a segmentation network is severely reduced due to the downsampling operations in the CNN backbone. Most previous methods employ upsampling decoders to recover the spatial resolution. Various de…

Cited by 16SourcePDFScholar
2021

FCPose: Fully Convolutional Multi-Person Pose Estimation With Dynamic Instance-Aware Convolutions

CVPR 2021poster

We propose a fully convolutional multi-person pose estimation framework using dynamic instance-aware convolutions, termed FCPose. Different from existing methods, which often require ROI (Region of Interest) operations and/or grouping post-processing, FCPose eliminates the ROIs and grouping post-pro…

Cited by 81PDFScholar
2021

Twins: Revisiting the Design of Spatial Attention in Vision Transformers

NeurIPS 2021poster

Very recently, a variety of vision transformer architectures for dense prediction tasks have been proposed and they show that the design of spatial attention is critical to their success in these tasks. In this work, we revisit the design of the spatial attention and demonstrate that a carefully dev…

2020

BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation

CVPR 2020oral

Instance segmentation is one of the fundamental vision tasks. Recently, fully convolutional instance segmentation methods have drawn much attention as they are often simpler and more efficient than two-stage approaches like Mask R-CNN. To date, almost all such approaches fall behind the two-stage Ma…

Cited by 697PDFScholar
2020

Cumulant Slice Reconstruction from Compressive Measurements and Its Application to Line Spectrum Estimation

ICASSP 2020accepted

Higher-order statistics (HOS) estimation hinges on the availability of a huge amount of data records, which causes exceedingly high sampling rates and overwhelming energy consumption for the sampling devices, especially when dealing with wideband signals. To overcome these challenges, this paper dev…

Cited by 0SourceScholar
2020

Efficient Super-Resolution Two-Dimensional Harmonic Retrieval Via Enhanced Low-Rank Structured Covariance Reconstruction

ICASSP 2020accepted

This paper develops an enhanced low-rank structured covariance reconstruction (LRSCR) method based on the decoupled atomic norm minimization (D-ANM), for super-resolution two-dimensional (2D) harmonic retrieval with multiple measurement vectors. This LRSCR-D-ANM approach exploits a potential structu…

Cited by 7SourceScholar
2020

Learning and Memorizing Representative Prototypes for 3D Point Cloud Semantic and Instance Segmentation

ECCV 2020poster

3D point cloud semantic and instance segmentation are crucial and fundamental for 3D scene understanding. Due to the complex structure, point sets are distributed off-balance and diversely, appearing as both category and pattern imbalance. It has been proved that deep networks can easily forget the…

Cited by 50SourcePDFScholar
2020

NAS-FCOS: Fast Neural Architecture Search for Object Detection

CVPR 2020poster

The success of deep neural networks relies on significant architecture engineering. Recently neural architecture search (NAS) has emerged as a promise to greatly reduce manual effort in network design by automatically searching for optimal architectures, although typically such algorithms need an ex…

Cited by 281PDFScholar
2019

COLA: Communication-censored Linearized ADMM for Decentralized Consensus Optimization

ICASSP 2019accepted

This paper proposes a communication- and computation-efficient algorithm to solve a convex consensus optimization problem defined over a decentralized network. A remarkable existing algorithm to solve this problem is the alternating direction method of multipliers (ADMM), in which at every iteration…

Cited by 0SourceScholar
2019

Decoders Matter for Semantic Segmentation: Data-Dependent Decoding Enables Flexible Feature Aggregation

CVPR 2019poster

Recent semantic segmentation methods exploit encoder-decoder architectures to produce the desired pixel-wise segmentation prediction. The last layer of the decoders is typically a bilinear upsampling procedure to recover the final pixel-wise prediction. We empirically show that this oversimple and d…

Cited by 297PDFScholar
2019

Knowledge Adaptation for Efficient Semantic Segmentation

CVPR 2019poster

Both accuracy and efficiency are of significant importance to the task of semantic segmentation. Existing deep FCNs suffer from heavy computations due to a series of high-resolution feature maps for preserving the detailed knowledge in dense estimation. Although reducing the feature map resolution (…

Cited by 291PDFScholar
2018

An End-to-End TextSpotter With Explicit Alignment and Attention

CVPR 2018poster

Text detection and recognition in natural images have long been considered as two separate tasks that are processed sequentially. Jointly training two tasks is non-trivial due to significant differences in learning difficulties and convergence rates. In this work, we present a conceptually simple ye…

2017

Low-complexity optimization for two-dimensional direction-of-arrival estimation via decoupled atomic norm minimization

ICASSP 2017accepted

This paper presents an efficient optimization technique for super-resolution two-dimensional (2D) direction of arrival (DOA) estimation by introducing a new formulation of atomic norm minimization (ANM). ANM allows gridless angle estimation for correlated sources even when the number of snapshots is…

Cited by 0SourceScholar
2016

Communication-efficient weighted ADMM for decentralized network optimization

ICASSP 2016accepted

In this paper, we propose a weighted alternating direction method of multipliers (ADMM) to solve the consensus optimization problem over a decentralized network. Compared with the conventional ADMM that is popular in decentralized network optimization, the weighted ADMM is able to tune its weight ma…

Cited by 0SourceScholar
2016

Efficient channel statistics estimation for millimeter-wave MIMO systems

ICASSP 2016accepted

In millimeter-wave (mmWave) multiple-input multiple-output (MIMO) systems, channel estimation is a challenging task in terms of acquiring the instantaneous channel state information (CSI), because both the estimation complexity and the overhead required for pilot symbols and feedback grow drasticall…

Cited by 0SourceScholar