← Search

Zitian Chen

5 accepted papers

2023

Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task Learners

CVPR 2023poster

Optimization in multi-task learning (MTL) is more challenging than single-task learning (STL), as the gradient from different tasks can be contradictory. When tasks are related, it can be beneficial to share some parameters among them (cooperation). However, some tasks require additional parameters…

Cited by 107SourcePDFScholar
2023

Visual Dependency Transformers: Dependency Tree Emerges From Reversed Attention

CVPR 2023poster

Humans possess a versatile mechanism for extracting structured representations of our visual world. When looking at an image, we can decompose the scene into entities and their parts as well as obtain the dependencies between them. To mimic such capability, we propose Visual Dependency Transformers…

2021

Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study

ICLR 2021poster

This work aims to empirically clarify a recently discovered perspective that label smoothing is incompatible with knowledge distillation. We begin by introducing the motivation behind on how this incompatibility is raised, i.e., label smoothing erases relative information between teacher logits. We…

Cited by 101SourcePDFScholar
2019

Image Deformation Meta-Networks for One-Shot Learning

CVPR 2019oral

Humans can robustly learn novel visual concepts even when images undergo various deformations and loose certain information. Mimicking the same behavior and synthesizing deformed instances of new concepts may help visual recognition systems perform better one-shot learning, i.e., learning concepts f…

Cited by 303PDFcodeScholar
2018

Environment Upgrade Reinforcement Learning for Non-Differentiable Multi-Stage Pipelines

CVPR 2018poster

Recent advances in multi-stage algorithms have shown great promise, but two important problems still remain. First of all, at inference time, information can't feed back from downstream to upstream. Second, at training time, end-to-end training is not possible if the overall pipeline involves non-di…

Cited by 8SourcePDFScholar