← Search

Changxu Cheng

5 accepted papers

2025

ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation

NeurIPS 2025poster

While humans effortlessly draw visual objects and shapes by adaptively allocating attention based on their complexity, existing multimodal large language models (MLLMs) remain constrained by rigid token representations. Bridging this gap, we propose ALTo, an adaptive length tokenizer for autoregress…

Cited by 0SourcecodeScholar
2025

HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

ICCV 2025poster

The remarkable performance of large multimodal models (LMMs) has attracted significant interest from the image segmentation community.To align with the next-token-prediction paradigm, current LMM-driven segmentation methods either use object boundary points to represent masks or introduce special se…

2024

DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing

EMNLP 2024main

Parsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding. However, previously the research on this topic has been largely hindered since most existing datasets are small-s…

2023

GeoLayoutLM: Geometric Pre-Training for Visual Information Extraction

CVPR 2023highlight

Visual information extraction (VIE) plays an important role in Document Intelligence. Generally, it is divided into two tasks: semantic entity recognition (SER) and relation extraction (RE). Recently, pre-trained models for documents have achieved substantial progress in VIE, particularly in SER. Ho…

2023

LISTER: Neighbor Decoding for Length-Insensitive Scene Text Recognition

ICCV 2023poster

The diversity in length constitutes a significant characteristic of text. Due to the long-tail distribution of text lengths, most existing methods for scene text recognition (STR) only work well on short or seen-length text, lacking the capability of recognizing longer text or performing length extr…

Cited by 29PDFcodeScholar