← Search

Erjin Zhou

10 accepted papers

2026

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

ICLR 2026poster

Temporal context is essential for robotic manipulation because such tasks are inherently non-Markovian, yet mainstream VLA models typically overlook it and struggle with long-horizon, temporally dependent tasks. Cognitive science suggests that humans rely on working memory to buffer short-lived repr…

Cited by 0SourcecodeScholar
2022

W2N: Switching from Weak Supervision to Noisy Supervision for Object Detection

ECCV 2022poster

"Weakly-supervised object detection (WSOD) aims to train an object detector only requiring the image-level annotations. Recently, some works have managed to select the accurate boxes generated from a well-trained WSOD network to supervise a semi-supervised detection framework for better performance.…

2021

General Instance Distillation for Object Detection

CVPR 2021poster

In recent years, knowledge distillation has been proved to be an effective solution for model compression. This approach can make lightweight student models acquire the knowledge extracted from cumbersome teacher models. However, previous distillation methods of detection have weak generalization fo…

Cited by 269PDFcodeScholar
2021

Rethinking the Heatmap Regression for Bottom-Up Human Pose Estimation

CVPR 2021poster

Heatmap regression has become the most prevalent choice for nowadays human pose estimation methods. The ground-truth heatmaps are usually constructed by covering all skeletal keypoints by 2D gaussian kernels. The standard deviations of these kernels are fixed. However, for bottom-up methods, which n…

Cited by 219PDFcodeScholar
2021

TokenPose: Learning Keypoint Tokens for Human Pose Estimation

ICCV 2021poster

Human pose estimation deeply relies on visual clues and anatomical constraints between parts to locate keypoints. Most existing CNN-based methods do well in visual representation, however, lacking in the ability to explicitly learn the constraint relationships between keypoints. In this paper, we pr…

Cited by 386PDFcodeScholar
2020

DPGN: Distribution Propagation Graph Network for Few-Shot Learning

CVPR 2020poster

Most graph-network-based meta-learning approaches model instance-level relation of examples. We extend this idea further to explicitly model the distribution-level relation of one example to all other examples in a 1-vs-N manner. We propose a novel approach named distribution propagation graph netwo…

Cited by 285PDFcodeScholar
2020

High-Order Information Matters: Learning Relation and Topology for Occluded Person Re-Identification

CVPR 2020poster

Occluded person re-identification (ReID) aims to match occluded person images to holistic ones across dis-joint cameras. In this paper, we propose a novel framework by learning high-order relation and topology information for discriminative features and robust alignment. At first, we use a CNN backb…

Cited by 555PDFcodeScholar
2020

Learning Delicate Local Representations for Multi-Person Pose Estimation

ECCV 2020poster

In this paper, we propose a novel method called Residual Steps Network (RSN). RSN aggregates features with the same spatial size (Intra-level features) efficiently to obtain delicate local representations, which retain rich low-level spatial information and result in precise keypoint localization. A…

2018

Symmetric Variational Autoencoder and Connections to Adversarial Learning

AISTATS 2018poster

A new form of the variational autoencoder (VAE) is proposed, based on the symmetric Kullback- Leibler divergence. It is demonstrated that learn- ing of the resulting symmetric VAE (sVAE) has close connections to previously developed adversarial-learning methods. This relationship helps unify the pre…

Cited by 0SourcePDFScholar