← Search

Kun Yin

3 accepted papers

2024

Few-shot Temporal Pruning Accelerates Diffusion Models for Text Generation

COLING 2024main

Diffusion models have achieved significant success in computer vision and shown immense potential in natural language processing applications, particularly for text generation tasks. However, generating high-quality text using these models often necessitates thousands of iterations, leading to slow…

Cited by 1SourcePDFScholar
2024

HRVDA: High-Resolution Visual Document Assistant

CVPR 2024poster

Leveraging vast training data multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However their performance in visual document understanding still leaves much room for improvement. T…

Cited by 20SourcePDFScholar
2023

Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration

ICCV 2023poster

We propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including document analysis, retrieval, and office automation. Unlike state-of-the-art approaches that rely on multi-stage techni…

Cited by 15PDFScholar