← Search

Zhenglu Yang

19 accepted papers

2025

Generating Questions, Answers, and Distractors for Videos: Exploring Semantic Uncertainty of Object Motions

ACL 2025finding

Video Question-Answer-Distractors (QADs) show promising values for assessing the performance of systems in perceiving and comprehending multimedia content. Given the significant cost and labor demands of manual annotation, existing large-scale Video QADs benchmarks are typically generated automatica…

Cited by 0SourcePDFScholar
2025

Listening to Patients: Detecting and Mitigating Patient Misreport in Medical Dialogue System

ACL 2025finding

Medical Dialogue Systems (MDSs) have emerged as promising tools for automated healthcare support through patient-agent interactions. Previous efforts typically relied on an idealized assumption — patients can accurately report symptoms aligned with their actual health conditions. However, in reality…

Cited by 0SourcePDFScholar
2024

Can We Learn Question, Answer, and Distractors All from an Image? A New Task for Multiple-choice Visual Question Answering

COLING 2024main

Multiple-choice visual question answering (MC VQA) requires an answer picked from a list of distractors, based on a question and an image. This research has attracted wide interest from the fields of visual question answering, visual question generation, and visual distractor generation. However, th…

Cited by 4SourcePDFScholar
2024

Exploring Union and Intersection of Visual Regions for Generating Questions, Answers, and Distractors

EMNLP 2024main

Multiple-choice visual question answering (VQA) is to automatically choose a correct answer from a set of choices after reading an image. Existing efforts have been devoted to a separate generation of an image-related question, a correct answer, or challenge distractors. By contrast, we turn to a ho…

2023

ACROSS: An Alignment-based Framework for Low-Resource Many-to-One Cross-Lingual Summarization

ACL 2023findings

This research addresses the challenges of Cross-Lingual Summarization (CLS) in low-resource scenarios and over imbalanced multilingual data. Existing CLS studies mostly resort to pipeline frameworks or multi-task methods in bilingual settings. However, they ignore the data imbalance in multilingual…

2023

HaPPy: Harnessing the Wisdom from Multi-Perspective Graphs for Protein-Ligand Binding Affinity Prediction (Student Abstract)

AAAI 2023technical

Gathering information from multi-perspective graphs is an essential issue for many applications especially for proteinligand binding affinity prediction. Most of traditional approaches obtained such information individually with low interpretability. In this paper, we harness the rich information fr…

Cited by 0SourcePDFScholar
2023

HyperPELT: Unified Parameter-Efficient Language Model Tuning for Both Language and Vision-and-Language Tasks

ACL 2023findings

With the scale and capacity of pretrained models growing rapidly, parameter-efficient language model tuning has emerged as a popular paradigm for solving various NLP and Vision-and-Language (V&L) tasks. In this paper, we design a unified parameter-efficient multitask learning framework that works ef…

Cited by 17SourcePDFScholar
2023

Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded Dialogue

EMNLP 2023long main

Accurate knowledge selection is critical in knowledge-grounded dialogue systems. Towards a closer look at it, we offer a novel perspective to organize existing literature, i.e., knowledge selection coupled with, after, and before generation. We focus on the third under-explored category of study,…

Cited by 0SourcecodeScholar
2022

Fact-Tree Reasoning for N-ary Question Answering over Knowledge Graphs

ACL 2022findings

Current Question Answering over Knowledge Graphs (KGQA) task mainly focuses on performing answer reasoning upon KGs with binary facts. However, it neglects the n-ary facts, which contain more than two entities. In this work, we highlight a more challenging but under-explored task: n-ary KGQA, i.e.,…

Cited by 8SourcePDFScholar
2022

Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer Supervision

ICASSP 2022accepted

Recently, pioneer work finds that self-supervised pre-training methods can improve multiple downstream speech tasks, because the model utilizes bottom layers to learn speaker-related information and top layers to encode content-related information. Since the network capacity is limited, we believe t…

Cited by 0SourceScholar
2022

Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension

ACL 2022long

Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream reasoning tasks is designated as Procedural MultiModal Machine Comprehension (M3C). In this study, we approach Procedur…

2022

Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems

ACL 2022long

Empathetic dialogue assembles emotion understanding, feeling projection, and appropriate response generation. Existing work for empathetic dialogue generation concentrates on the two-party conversation scenario. Multi-party dialogues, however, are pervasive in reality. Furthermore, emotion and sensi…

Cited by 17SourcePDFScholar
2022

UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation

AAAI 2022technical

With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly…

2021

Deep Symmetric Network for Underexposed Image Enhancement With Recurrent Attentional Learning

ICCV 2021poster

Underexposed image enhancement is of importance in many research domains. In this paper, we take this problem as image feature transformation between the underexposed image and its paired enhanced version, and we propose a deep symmetric network for the issue. Our symmetric network adapts invertible…

Cited by 70PDFScholar
2021

GMH: A General Multi-hop Reasoning Model for KG Completion

EMNLP 2021main

Knowledge graphs are essential for numerous downstream natural language processing applications, but are typically incomplete with many facts missing. This results in research efforts on multi-hop reasoning task, which can be formulated as a search process and current models typically perform short…

Cited by 17SourcePDFScholar
2021

Generalized Relation Learning with Semantic Correlation Awareness for Link Prediction

AAAI 2021technical

Developing link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link prediction task have two natural problems: 1) the relation distributions in KGs are usually unbalanced, and 2) there are many un…

Cited by 18SourcePDFScholar
2021

News Content Completion with Location-Aware Image Selection

AAAI 2021technical

News, as one of the fundamental social media types, typically contains both texts and images. Image selection, which involves choosing appropriate images according to some specified contexts, is crucial for formulating good news. However, it presents two challenges: where to place images and which i…

Cited by 2SourcePDFScholar
2021

RepSum: Unsupervised Dialogue Summarization based on Replacement Strategy

ACL 2021long

In the field of dialogue summarization, due to the lack of training data, it is often difficult for supervised summary generation methods to learn vital information from dialogue context with limited data. Several attempts on unsupervised summarization for text by leveraging semantic information sol…

Cited by 15SourcePDFScholar