← Search

Zhengkun Zhang

8 accepted papers

2024

M3sum: A Novel Unsupervised Language-Guided Video Summarization

ICASSP 2024accepted

Language-guided video summarization empowers users to use natural language queries to effortlessly summarize lengthy videos into concise and relevant summaries that cater specifically to their information needs, which is more friendly to access and digest. However, most of the previous works rely on…

Cited by 0SourceScholar
2023

ACROSS: An Alignment-based Framework for Low-Resource Many-to-One Cross-Lingual Summarization

ACL 2023findings

This research addresses the challenges of Cross-Lingual Summarization (CLS) in low-resource scenarios and over imbalanced multilingual data. Existing CLS studies mostly resort to pipeline frameworks or multi-task methods in bilingual settings. However, they ignore the data imbalance in multilingual…

2023

HyperPELT: Unified Parameter-Efficient Language Model Tuning for Both Language and Vision-and-Language Tasks

ACL 2023findings

With the scale and capacity of pretrained models growing rapidly, parameter-efficient language model tuning has emerged as a popular paradigm for solving various NLP and Vision-and-Language (V&L) tasks. In this paper, we design a unified parameter-efficient multitask learning framework that works ef…

Cited by 17SourcePDFScholar
2023

Licon: A Diverse, Controllable and Challenging Linguistic Concept Learning Benchmark

EMNLP 2023long findings

Concept Learning requires learning the definition of a general category from given training examples. Most of the existing methods focus on learning concepts from images. However, the visual information cannot present abstract concepts exactly, which struggles the introduction of novel concepts rela…

Cited by 0SourceScholar
2022

Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension

ACL 2022long

Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream reasoning tasks is designated as Procedural MultiModal Machine Comprehension (M3C). In this study, we approach Procedur…

2022

Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems

ACL 2022long

Empathetic dialogue assembles emotion understanding, feeling projection, and appropriate response generation. Existing work for empathetic dialogue generation concentrates on the two-party conversation scenario. Multi-party dialogues, however, are pervasive in reality. Furthermore, emotion and sensi…

Cited by 17SourcePDFScholar
2022

UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation

AAAI 2022technical

With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly…

2021

News Content Completion with Location-Aware Image Selection

AAAI 2021technical

News, as one of the fundamental social media types, typically contains both texts and images. Image selection, which involves choosing appropriate images according to some specified contexts, is crucial for formulating good news. However, it presents two challenges: where to place images and which i…

Cited by 2SourcePDFScholar