← Search

Zhenlin Xu

10 accepted papers

2026

Spherical Leech Quantization for Visual Tokenization and Generation

CVPR 2026

Lookup-free quantization has received much attention due to its efficiency on parameters and scalability to a large codebook. In this paper, we present a unified formulation of different non-parametric quantization methods through the lens of lattice coding. The geometry of lattice codes explains th

Cited by 0SourcecodeScholar
2025

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

NeurIPS 2025poster

Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns. Training LLMs on openly licensed text presents a first step towards addressing these issues, but…

Cited by 0SourceScholar
2024

Self-Supervised Multi-Object Tracking with Path Consistency

CVPR 2024highlight

In this paper we propose a novel concept of path consistency to learn robust object matching without using manual object identity supervision. Our key idea is that to track a object through frames we can obtain multiple different association results from a model by varying the frames it can observe…

2023

ScaleDet: A Scalable Multi-Dataset Object Detector

CVPR 2023poster

Multi-dataset training provides a viable solution for exploiting heterogeneous large-scale datasets without extra annotation cost. In this work, we propose a scalable multi-dataset detector (ScaleDet) that can scale up its generalization across datasets when increasing the number of training dataset…

Cited by 23SourcePDFScholar
2023

SimpleClick: Interactive Image Segmentation with Simple Vision Transformers

ICCV 2023poster

Click-based interactive image segmentation aims at extracting objects with a limited user clicking. A hierarchical backbone is the de-facto architecture for current methods. Recently, the plain, non-hierarchical Vision Transformer (ViT) has emerged as a competitive backbone for dense prediction task…

Cited by 174PDFcodeScholar
2022

Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language

NeurIPS 2022accept

Deep learning models struggle with compositional generalization, i.e. the ability to recognize or generate novel combinations of observed elementary concepts. In hopes of enabling compositional generalization, various unsupervised learning algorithms have been proposed with inductive biases that aim…

2021

Robust and Generalizable Visual Representation Learning via Random Convolutions

ICLR 2021poster

While successful for various computer vision tasks, deep neural networks have shown to be vulnerable to texture style shifts and small perturbations to which humans are robust. In this work, we show that the robustness of neural networks can be greatly improved through the use of random convolutions…

Cited by 267SourcePDFScholar
2020

Adversarial Data Augmentation via Deformation Statistics

ECCV 2020poster

Deep learning models have been successful in computer vision and medical image analysis. However, training these models frequently requires large labeled image sets whose creation is often very time and labor intensive, for example, in the context of 3D segmentations. Approaches capable of training…

Cited by 12SourcePDFScholar