← Search

Liheng Zhang

6 accepted papers

2026

HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models

AAAI 2026technical

3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. Howe

Cited by 0SourcePDFScholar
2019

AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations Rather Than Data

CVPR 2019oral

The success of deep neural networks often relies on a large amount of labeled examples, which can be difficult to obtain in many real scenarios. To address this challenge, unsupervised methods are strongly preferred for training neural networks without using any labeled data. In this paper, we prese…

Cited by 265PDFcodeScholar
2019

AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations

ICCV 2019poster

The learning of Transformation-Equivariant Representations (TERs), which is introduced by Hinton et al. [??], has been considered as a principle to reveal visual structures under various transformations. It contains the celebrated Convolutional Neural Networks (CNNs) as a special case that only equi…

Cited by 49PDFScholar
2018

CapProNet: Deep Feature Learning via Orthogonal Projections onto Capsule Subspaces

NeurIPS 2018poster

In this paper, we formalize the idea behind capsule nets of using a capsule vector rather than a neuron activation to predict the label of samples. To this end, we propose to learn a group of capsule subspaces onto which an input feature vector is projected. Then the lengths of resultant capsules ar…

Cited by 89SourcePDFScholar
2018

Global Versus Localized Generative Adversarial Nets

CVPR 2018poster

In this paper, we present a novel localized Generative Adversarial Net (GAN) to learn on the manifold of real data. Compared with the classic GAN that {em globally} parameterizes a manifold, the Localized GAN (LGAN) uses local coordinate charts to parameterize distinct local geometry of how data poi…

Cited by 97SourcePDFScholar