← Search

Michał Pietruszka

3 accepted papers

2023

Document Understanding Dataset and Evaluation (DUDE)

ICCV 2023poster

We call on the Document AI (DocAI) community to re-evaluate current methodologies and embrace the challenge of creating more practically-oriented benchmarks. Document Understanding Dataset and Evaluation (DUDE) seeks to remediate the halted research progress in understanding visually-rich documents…

Cited by 67PDFcodeScholar
2022

Sparsifying Transformer Models with Trainable Representation Pooling

ACL 2022long

We propose a novel method to sparsify attention in the Transformer model by learning to select the most-informative token representations during the training process, thus focusing on the task-specific parts of an input. A reduction of quadratic time and memory complexity to sublinear was achieved d…

2021

DUE: End-to-End Document Understanding Benchmark

NeurIPS 2021poster

Understanding documents with rich layouts plays a vital role in digitization and hyper-automation but remains a challenging topic in the NLP research community. Additionally, the lack of a commonly accepted benchmark made it difficult to quantify progress in the domain. To empower research in this f…

Cited by 63SourceScholar