← Search

Dmytro Okhonko

5 accepted papers

2022

CCQA: A New Web-Scale Question Answering Dataset for Model Pre-Training

NAACL 2022findings

We propose a novel open-domain question-answering dataset based on the Common Crawl project. With a previously unseen number of around 130 million multilingual question-answer pairs (including about 60 million English data-points), we use our large-scale, natural, diverse and high-quality corpus to…

2022

HTLM: Hyper-Text Pre-Training and Prompting of Language Models

ICLR 2022poster

We introduce HTLM, a hyper-text language model trained on a large-scale web crawl. Modeling hyper-text has a number of advantages: (1) it is easily gathered at scale, (2) it provides rich document-level and end-task-adjacent supervision (e.g. 'class' and 'id' attributes often encode document categor…

Cited by 84SourcePDFScholar
2022

UniK-QA: Unified Representations of Structured and Unstructured Knowledge for Open-Domain Question Answering

NAACL 2022findings

We study open-domain question answering with structured, unstructured and semi-structured knowledge sources, including text, tables, lists and knowledge bases. Departing from prior work, we propose a unifying approach that homogenizes all sources by reducing them to text and applies the retriever-re…

2021

VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

EMNLP 2021main

We present VideoCLIP, a contrastive approach to pre-train a unified model for zero-shot video and text understanding, without using any labels on downstream tasks. VideoCLIP trains a transformer for video and text by contrasting temporally overlapping positive video-text pairs with hard negatives fr…

2020

Training ASR Models By Generation of Contextual Information

ICASSP 2020accepted

Supervised ASR models have reached unprecedented levels of accuracy, thanks in part to ever-increasing amounts of labelled training data. However, in many applications and locales, only moderate amounts of data are available, which has led to a surge in semi- and weakly-supervised learning research.…

Cited by 0SourceScholar