← Search

Jack Urbanek

7 accepted papers

2024

A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions

CVPR 2024poster

Curation methods for massive vision-language datasets trade off between dataset size and quality. However even the highest quality of available curated captions are far too short to capture the rich visual detail in an image. To show the value of dense and highly-aligned image-text pairs we collect…

2024

Text Motion Translator: A Bi-Directional Model for Enhanced 3D Human Motion Generation from Open-Vocabulary Descriptions

ECCV 2024poster

"The field of 3D human motion generation from natural language descriptions, known as Text2Motion, has gained significant attention for its potential application in industries such as film, gaming, and AR/VR. To tackle a key challenge in Text2Motion, the deficiency of 3D human motions and their corr…

Cited by 3SourcePDFScholar
2023

Breaking The Limits of Text-conditioned 3D Motion Synthesis with Elaborative Descriptions

ICCV 2023poster

Given its wide applications, there is increasing focus on generating 3D human motions from textual descriptions. Differing from the majority of previous works, which regard actions as single entities and can only generate short sequences for simple motions, we propose EMS, an elaborative motion synt…

Cited by 16PDFScholar
2022

Am I Me or You? State-of-the-Art Dialogue Models Cannot Maintain an Identity

NAACL 2022findings

State-of-the-art dialogue models still often stumble with regards to factual accuracy and self-contradiction. Anecdotally, they have been observed to fail to maintain character identity throughout discourse; and more specifically, may take on the role of their interlocutor. In this work we formalize…

Cited by 29SourcePDFScholar
2022

Reason first, then respond: Modular Generation for Knowledge-infused Dialogue

EMNLP 2022finding

Large language models can produce fluent dialogue but often hallucinate factual inaccuracies. While retrieval-augmented models help alleviate this issue, they still face a difficult challenge of both reasoning to provide correct knowledge and generating conversation simultaneously. In this work, we…

Cited by 48SourcePDFScholar
2021

How to Motivate Your Dragon: Teaching Goal-Driven Agents to Speak and Act in Fantasy Worlds

NAACL 2021long

We seek to create agents that both act and communicate with other agents in pursuit of a goal. Towards this end, we extend LIGHT (Urbanek et al. 2019)—a large-scale crowd-sourced fantasy text-game—with a dataset of quests. These contain natural language motivations paired with in-game goals and huma…

Cited by 59SourcePDFScholar
2018

Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent

ICLR 2018poster

Contrary to most natural language processing research, which makes use of static datasets, humans learn language interactively, grounded in an environment. In this work we propose an interactive learning procedure called Mechanical Turker Descent (MTD) that trains agents to execute natural language…

Cited by 32SourcePDFScholar