← Search

Tariq Iqbal

16 accepted papers

2026

Energy-Based Transformers are Scalable Learners and Thinkers

ICLR 2026oral

Inference-time computation, analogous to human System 2 Thinking, has recently become popular for improving model performance. However, most existing approaches suffer from several limitations: they are modality-specific (e.g., working only in text), problem-specific (e.g., verifiable domains like m…

Cited by 0SourcecodeScholar
2025

DM-Codec: Distilling Multimodal Representations for Speech Tokenization

EMNLP 2025

Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete tokens remains challenging. This process demands acoustic, semantic, and contextual

2024

EQA-MX: Embodied Question Answering using Multimodal Expression

ICLR 2024spotlight

Humans predominantly use verbal utterances and nonverbal gestures (e.g., eye gaze and pointing gestures) in their natural interactions. For instance, pointing gestures and verbal information is often required to comprehend questions such as "what object is that?" Thus, this question-answering (QA) t…

Cited by 10SourcePDFScholar
2023

PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal Cues

AAAI 2023technical

Humans naturally use referring expressions with verbal utterances and nonverbal gestures to refer to objects and events. As these referring expressions can be interpreted differently from the speaker's or the observer's perspective, people effectively decide on the perspective in comprehending the e…

Cited by 4SourcePDFScholar
2023

Representation Learning in Deep RL via Discrete Information Bottleneck

AISTATS 2023poster

Several self-supervised representation learning methods have been proposed for reinforcement learning (RL) with rich observations. For real world applications of RL, recovering underlying latent states is crucial, particularly when sensory inputs can contain irrelevant and exogenous information. In…

Cited by 11SourcePDFScholar
2022

CAESAR: An Embodied Simulator for Generating Multimodal Referring Expression Datasets

NeurIPS 2022accept

Humans naturally use verbal utterances and nonverbal gestures to refer to various objects (known as $\textit{referring expressions}$) in different interactional scenarios. As collecting real human interaction datasets are costly and laborious, synthetic datasets are often used to train models to una…

Cited by 13SourcePDFScholar
2021

Multi-GAT: A Graphical Attention-Based Hierarchical Multimodal Representation Learning Approach for Human Activity Recognition

RA-L 2021

Recognizing human activities is one of the crucial capabilities that a robot needs to have to be useful around people. Although modern robots are equipped with various types of sensors, human activity recognition (HAR) still remains a challenging problem, particularly in the presence of noisy sensor

Cited by 88SourceScholar
2019

Activity recognition in manufacturing: The roles of motion capture and sEMG+inertial wearables in detecting fine vs. gross motion

ICRA 2019poster

In safety-critical environments, robots need to reliably recognize human activity to be effective and trust-worthy partners. Since most human activity recognition (HAR) approaches rely on unimodal sensor data (e.g. motion capture or wearable sensors), it is unclear how the relationship between the s…

Cited by 66SourceScholar
2019

Fast Online Segmentation of Activities from Partial Trajectories

ICRA 2019poster

Augmenting a robot with the capacity to understand the activities of the people it collaborates with in order to then label and segment those activities allows the robot to generate an efficient and safe plan for performing its own actions. In this work, we introduce an online activity segmentation…

Cited by 22SourceScholar