← Search

Arash Eshghi

9 accepted papers

2024

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

EMNLP 2024finding

AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person view videos, neglecting the richness of egocentric perceptual experience. To address this gap,…

2024

Clarifying Completions: Evaluating How LLMs Respond to Incomplete Questions

COLING 2024main

People understand and produce language incrementally on a word by word basis. This gives rise to many characteristic conversational phenomena including long mid-sentence pauses that are followed by incremental clarification requests (iCRs) intended to recover the rest of the truncated turn (see Fig.…

2024

Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers

NAACL 2024short

An effective method for combining frozen large language models (LLM) and visual encoders involves a resampler module that creates a ‘visual prompt’ which is provided to the LLM, along with the textual prompt. While this approach has enabled impressive performance across many coarse-grained tasks lik…

2024

Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs

EMNLP 2024main

Evaluating Large Language Models (LLMs) on reasoning benchmarks demonstrates their ability to solve compositional questions. However, little is known of whether these models engage in genuine logical reasoning or simply rely on implicit cues to generate answers. In this paper, we investigate the tra…

Cited by 0SourcePDFScholar
2024

Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models

EMNLP 2024main

In dialogue, the addressee may initially misunderstand the speaker and respond erroneously, often prompting the speaker to correct the misunderstanding in the next turn with a Third Position Repair (TPR). The ability to process and respond appropriately to such repair sequences is thus crucial in co…

2024

Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling

EMNLP 2024main

This study explores replacing Transformers in Visual Language Models (VLMs) with Mamba, a recent structured state space model (SSM) that demonstrates promising performance in sequence modeling. We test models up to 3B parameters under controlled conditions, showing that Mamba-based VLMs outperforms…

2023

Multitask Multimodal Prompted Training for Interactive Embodied Task Completion

EMNLP 2023long main

Interactive and embodied tasks pose at least two fundamental challenges to existing Vision \& Language (VL) models, including 1) grounding language in trajectories of actions and observations, and 2) referential disambiguation. To tackle these challenges, we propose an Embodied MultiModal Agent (EMM…

Cited by 0SourceScholar
2023

The Dangers of trusting Stochastic Parrots: Faithfulness and Trust in Open-domain Conversational Question Answering

ACL 2023findings

Large language models are known to produce output which sounds fluent and convincing, but is also often wrong, e.g. “unfaithful” with respect to a rationale as retrieved from a knowledge base. In this paper, we show that task-based systems which exhibit certain advanced linguistic dialog behaviors,…

Cited by 34SourcePDFScholar
2020

A Comprehensive Evaluation of Incremental Speech Recognition and Diarization for Conversational AI

COLING 2020main

Automatic Speech Recognition (ASR) systems are increasingly powerful and more accurate, but also more numerous with several options existing currently as a service (e.g. Google, IBM, and Microsoft). Currently the most stringent standards for such systems are set within the context of their use in, a…