← Search

Tong Tong

7 accepted papers

2026

Entropy-Aware Dynamic KV Cache Sparsification for Autoregressive Image Generation and Editing

ICML 2026poster

Autoregressive (AR) image generation has recently gained momentum as a scalable alternative to diffusion models, benefiting from unified next-token prediction paradigm and strong instruction following ability. However, AR visual generation must decode excessively long sequences of visual tokens, mak…

Cited by 0SourceScholar
2025

Contrastive Learning via Randomly Generated Deep Supervision

ICASSP 2025accepted

Unsupervised visual representation learning has gained significant attention in the computer vision community, driven by recent advancements in contrastive learning. Most existing contrastive learning frameworks rely on instance discrimination as a pretext task, treating each instance as a distinct…

Cited by 0SourceScholar
2025

MoEdit: On Learning Quantity Perception for Multi-object Image Editing

CVPR 2025poster

Multi-object images are widely present in the real world, spanning various areas of daily life. Efficient and accurate editing of these images is crucial for applications such as augmented reality, advertisement design, and medical imaging. Stable Diffusion (SD) has ushered in a new era of high-qual…

2024

A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications

ICASSP 2024accepted

Although current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propo…

Cited by 0SourceScholar
2023

MPCViT: Searching for Accurate and Efficient MPC-Friendly Vision Transformer with Heterogeneous Attention

ICCV 2023poster

Secure multi-party computation (MPC) enables computation directly on encrypted data and protects both data and model privacy in deep learning inference. However, existing neural network architectures, including Vision Transformers (ViTs), are not designed or optimized for MPC and incur significant l…

Cited by 23PDFcodeScholar
2021

Architecture Disentanglement for Deep Neural Networks

ICCV 2021poster

Understanding the inner workings of deep neural networks (DNNs) is essential to provide trustworthy artificial intelligence techniques for practical applications. Existing studies typically involve linking semantic concepts to units or layers of DNNs, but fail to explain the inference process. In th…

Cited by 25PDFcodeScholar