← Search

Josiah Poon

8 accepted papers

2025

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

IJCAI 2025

Visually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across domains such as medical, financial, and educational applications. However, form-like documents pose unique challenges d

Cited by 0SourcePDFScholar
2024

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

ACL 2024findings

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fine-grained and coarse-grained levels by facilitating a nuanced correlation betwe…

2024

Do Text-to-Vis Benchmarks Test Real Use of Visualisations?

EMNLP 2024main

Large language models are able to generate code for visualisations in response to simple user requests.This is a useful application and an appealing one for NLP research because plots of data provide grounding for language.However, there are relatively few benchmarks, and those that exist may not be…

2022

Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis

COLING 2022main

Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent studies in Document Layout Analysis usually rely on visual cues to understand documents while ignoring other information, su…

2022

Understanding Attention for Vision-and-Language Tasks

COLING 2022main

Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While attention has been widely used in VL tasks, it has not been examined the capability of different attention alignment calcula…

2020

Detect All Abuse! Toward Universal Abusive Language Detection Models

COLING 2020main

Online abusive language detection (ALD) has become a societal issue of increasing importance in recent years. Several previous works in online ALD focused on solving a single abusive language problem in a single domain, like Twitter, and have not been successfully transferable to the general ALD tas…

2020

VICTR: Visual Information Captured Text Representation for Text-to-Vision Multimodal Tasks

COLING 2020main

Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited information in order to fully describe visually realistic images. We propose a new visual contextual text representation for t…