← Search

Haofu Liao

12 accepted papers

2026

Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding

CVPR 2026

Autoregressive (AR) vision-language models (VLMs) have long dominated multimodal understanding, reasoning, and graphical user interface (GUI) grounding. Recently, discrete diffusion vision-language models (DVLMs) have shown strong performance in multimodal reasoning, offering bidirectional attention

Cited by 0SourceScholar
2025

Turbocharging Web Automation: The Impact of Compressed History States

ACL 2025finding

Language models have led to leap forward in web automation. The current web automation approaches take the current web state, history actions, and language instruction as inputs to predict the next action, overlooking the importance of history states. However, the highly verbose nature of web page s…

Cited by 0SourcePDFScholar
2024

DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models

EMNLP 2024main

Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This study aims to enhance generalizability of small VDU models by distilling knowledge from LLMs. We identify that directly p…

Cited by 0SourcePDFScholar
2023

DocTr: Document Transformer for Structured Information Extraction in Documents

ICCV 2023poster

We present a new formulation for structured information extraction (SIE) from visually rich documents. We address the limitations of existing IOB tagging and graph-based formulations, which are either overly reliant on the correct ordering of input text or struggle with decoding a complex graph. Ins…

Cited by 23PDFScholar
2021

Learning Bias-Invariant Representation by Cross-Sample Mutual Information Minimization

ICCV 2021poster

Deep learning algorithms mine knowledge from the training data and thus would likely inherit the dataset's bias information. As a result, the obtained model would generalize poorly and even mislead the decision process in real-life applications. We propose to remove the bias information misused by t…

Cited by 50PDFScholar
2021

Visual Relationship Detection Using Part-and-Sum Transformers With Composite Queries

ICCV 2021poster

Computer vision applications such as visual relationship detection and human object interaction can be formulated as a composite (structured) set detection problem in which both the parts (subject, object, and predicate) and the sum (triplet as a whole) are to be detected in a hierarchical fashion.…

Cited by 48PDFScholar
2021

XraySyn: Realistic View Synthesis From a Single Radiograph Through CT Priors

AAAI 2021technical

A radiograph visualizes the internal anatomy of a patient through the use of X-ray, which projects 3D information onto a 2D plane. Hence, radiograph analysis naturally requires physicians to relate their prior knowledge about 3D human anatomy to 2D radiographs. Synthesizing novel radiographic views…

2020

Example-Guided Image Synthesis using Masked Spatial-Channel Attention and Self-Supervision

ECCV 2020poster

Example-guided image synthesis has recently been attempted to synthesize an image from a semantic label map and an exemplary image. In the task, the additional exemplar image provides the style guidance that controls the appearance of the synthesized output. Despite the controllability advantage, th…

Cited by 22SourcePDFScholar
2020

SAINT: Spatially Aware Interpolation NeTwork for Medical Slice Synthesis

CVPR 2020poster

Deep learning-based single image super-resolution (SISR) methods face various challenges when applied to 3D medical volumetric data (i.e., CT and MR images) due to the high memory cost and anisotropic resolution, which adversely affect their performance. Furthermore, mainstream SISR methods are desi…

Cited by 63PDFScholar
2020

Structured Landmark Detection via Topology-Adapting Deep Graph Learning

ECCV 2020poster

Image landmark detection aims to automatically identify the locations of predefined fiducial points. Despite recent success in this field, higher-ordered structural modeling to capture implicit or explicit relationships among anatomical landmarks has not been adequately exploited. In this work, we p…

Cited by 121SourcePDFScholar
2019

DuDoNet: Dual Domain Network for CT Metal Artifact Reduction

CVPR 2019poster

Computed tomography (CT) is an imaging modality widely used for medical diagnosis and treatment. CT images are often corrupted by undesirable artifacts when metallic implants are carried by patients, which creates the problem of metal artifact reduction (MAR). Existing methods for reducing the artif…

Cited by 267PDFScholar
2019

Multiview 2D/3D Rigid Registration via a Point-Of-Interest Network for Tracking and Triangulation

CVPR 2019poster

We propose to tackle the problem of multiview 2D/3D rigid registration for intervention via a Point-Of-Interest Network for Tracking and Triangulation (POINT^2). POINT^2 learns to establish 2D point-to-point correspondences between the pre- and intra-intervention images by tracking a set of random P…

Cited by 65PDFScholar