← Search

Wenyu Sun

2 accepted papers

2026

Rosetta Stone For Unified MLLMs: A Unified Tokenizer to Decipher Understanding and Generation

CVPR 2026

Major state-of-the-art unified tokenizers predominantly adopt pixel reconstruction and feature alignment as pretext tasks, they leave key domains largely unexplored such as architecture, supervised objectives and tasks interaction, potentially resulting in limited performance. We systematically inve

Cited by 0SourceScholar
2020

High-quality Single-model Deep Video Compression with Frame-Conv3D and Multi-frame Differential Modulation

ECCV 2020poster

Deep learning (DL) methods have revolutionized the paradigm of computer vision tasks and DL-based video compression is becoming a hot topic. This paper proposes a deep video compression method to simultaneously encode multiple frames with Frame-Conv3D and differential modulation. We first adopt Fram…

Cited by 15SourcePDFScholar