← Search

Zhen Song

2 accepted papers

2026

ApET: Approximation-Error Guided Token Compression for Efficient VLMs

CVPR 2026

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead and degrade inference efficiency. Prior studies typically relies on [CLS] attention or text-vision cross-attention to iden

Cited by 0SourcecodeScholar
2026

Breaking the Continuum: Discrete Distribution Learning for Structural MRI Reconstruction

CVPR 2026

Anatomical structures in MRI exhibit strong spatial priors, including well-defined boundaries, low inter-subject variability, and consistent topology. These properties naturally induce clustered patterns in the latent space, which are difficult to capture using conventional continuous generative pri

Cited by 0SourcecodeScholar