← Search

Shaoning Xiao

3 accepted papers

2022

Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample Perspectives

EMNLP 2022main

Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA). The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at different levels of abstraction. Existing efforts mainly focus…

Cited by 6SourcePDFScholar
2021

Boundary Proposal Network for Two-stage Natural Language Video Localization

AAAI 2021technical

We aim to address the problem of Natural Language Video Localization (NLVL) — localizing the video segment corresponding to a natural language description in a long and untrimmed video. State-of-the-art NLVL methods are almost in one-stage fashion, which can be typically grouped into two categories:…

Cited by 187SourcePDFScholar
2021

Natural Language Video Localization with Learnable Moment Proposals

EMNLP 2021main

Given an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by query. To address this task, existing methods can be roughly grouped into two groups: 1) propose-and-rank models first define a set of hand-designed moment…