← Search

Xiang Dai

10 accepted papers

2025

Can VLMs Actually See and Read? A Survey on Modality Collapse in Vision-Language Models

ACL 2025finding

Vision-language models (VLMs) integrate textual and visual information, enabling the model to process visual inputs and leverage visual information to generate predictions. Such models are demanding for tasks such as visual question answering, image captioning, and visual grounding. However, some re…

Cited by 0SourcePDFScholar
2025

RnGCam: High-speed video from rolling & global shutter measurements

ICCV 2025poster

Compressive video capture encodes a short high-speed video into a single measurement using a low-speed sensor, then computationally reconstructs the original video. Prior implementations rely on expensive hardware and are restricted to imaging sparse scenes with empty backgrounds. We propose RnGCam,…

Cited by 0SourcePDFScholar
2025

The More, The Better? A Critical Study of Multimodal Context in Radiology Report Summarization

EMNLP 2025

The Impression section of a radiology report summarizes critical findings of a radiology report and thus plays a crucial role in communication between radiologists and physicians. Research on radiology report summarization mostly focuses on generating the Impression section by summarizing informatio

Cited by 0SourcePDFScholar
2024

Born Differently Makes a Difference: Counterfactual Study of Bias in Biography Generation from a Data-to-Text Perspective

ACL 2024short

How do personal attributes affect biography generation? Addressing this question requires an identical pair of biographies where only the personal attributes of interest are different. However, it is rare in the real world. To address this, we propose a counterfactual methodology from a data-to-text…

Cited by 1SourcePDFScholar
2024

Learning a Dynamic Privacy-preserving Camera Robust to Inversion Attacks

ECCV 2024oral

"The problem of designing a privacy-preserving camera (PPC) is considered. Previous designs rely on a static point spread function (PSF), optimized to prevent detection of private visual information, such as recognizable facial features. However, the PSF can be easily recovered by measuring the came…

Cited by 0SourcePDFScholar
2024

Understanding Faithfulness and Reasoning of Large Language Models on Plain Biomedical Summaries

EMNLP 2024finding

Generating plain biomedical summaries with Large Language Models (LLMs) can enhance the accessibility of biomedical knowledge to the public. However, how faithful the generated summaries are remains an open yet critical question. To address this, we propose FaReBio, a benchmark dataset with expert-a…

2022

Revisiting Transformer-based Models for Long Document Classification

EMNLP 2022finding

The recent literature in text classification is biased towards short text sequences (e.g., sentences or paragraphs). In real-world applications, multi-page multi-paragraph documents are common and they cannot be efficiently encoded by vanilla Transformer-based models. We compare different Transforme…

2021

mDAPT: Multilingual Domain Adaptive Pretraining in a Single Model

EMNLP 2021finding

Domain adaptive pretraining, i.e. the continued unsupervised pretraining of a language model on domain-specific text, improves the modelling of text for downstream tasks within the domain. Numerous real-world applications are based on domain-specific text, e.g. working with financial or biomedical d…