← Search

Alexi Gladstone

4 accepted papers

2026

Energy-Based Transformers are Scalable Learners and Thinkers

ICLR 2026oral

Inference-time computation, analogous to human System 2 Thinking, has recently become popular for improving model performance. However, most existing approaches suffer from several limitations: they are modality-specific (e.g., working only in text), problem-specific (e.g., verifiable domains like m…

Cited by 0SourcecodeScholar
2024

EQA-MX: Embodied Question Answering using Multimodal Expression

ICLR 2024spotlight

Humans predominantly use verbal utterances and nonverbal gestures (e.g., eye gaze and pointing gestures) in their natural interactions. For instance, pointing gestures and verbal information is often required to comprehend questions such as "what object is that?" Thus, this question-answering (QA) t…

Cited by 10SourcePDFScholar
2023

PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal Cues

AAAI 2023technical

Humans naturally use referring expressions with verbal utterances and nonverbal gestures to refer to objects and events. As these referring expressions can be interpreted differently from the speaker's or the observer's perspective, people effectively decide on the perspective in comprehending the e…

Cited by 4SourcePDFScholar
2022

CAESAR: An Embodied Simulator for Generating Multimodal Referring Expression Datasets

NeurIPS 2022accept

Humans naturally use verbal utterances and nonverbal gestures to refer to various objects (known as $\textit{referring expressions}$) in different interactional scenarios. As collecting real human interaction datasets are costly and laborious, synthetic datasets are often used to train models to una…

Cited by 13SourcePDFScholar