← Search

Nishant Sankaran

3 accepted papers

2024

DocFormerv2: Local Features for Document Understanding

AAAI 2024technical

We propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU). The VDU domain entails understanding documents (beyond mere OCR predictions) e.g., extracting information from a form, VQA for documents and other tasks. VDU is challenging as it needs a model to make sense of…

Cited by 47SourcePDFScholar
2023

A Multi-Modal Multilingual Benchmark for Document Image Classification

EMNLP 2023long findings

Document image classification is different from plain-text document classification and consists of classifying a document by understanding the content and structure of documents such as forms, emails, and other such documents. We show that the only existing dataset for this task (Lewis et al., 2006)…

Cited by 0SourceScholar
2020

Moving in the Right Direction: A Regularization for Deep Metric Learning

CVPR 2020poster

Deep metric learning leverages carefully designed sampling strategies and loss functions that aid in optimizing the generation of a discriminable embedding space. While effective sampling of pairs is critical for shaping the metric space during training, the relative interactions between pairs, and…

Cited by 45PDFScholar