← Search

Sherzod Hakimov

9 accepted papers

2025

Playpen: An Environment for Exploring Learning From Dialogue Game Feedback

EMNLP 2025

Interaction between learner and feedback-giver has come into focus recently for post-training of Large Language Models (LLMs), through the use of reward models that judge the appropriateness of a model’s response. In this paper, we investigate whether Dialogue Games—goal-directed and rule-governed a

2025

Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models

COLING 2025main

While the situation has improved for text-only models, it again seems to be the case currently that multimodal (text and image) models develop faster than ways to evaluate them. In this paper, we bring a recently developed evaluation paradigm from text models to multimodal models, namely evaluation…

2024

M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets

COLING 2024main

In recent years, multimodal natural language processing, aimed at learning from diverse data types, has garnered significant attention. However, there needs to be more clarity when it comes to analysing multimodal tasks in multi-lingual contexts. While prior studies on sentiment analysis of tweets h…

2024

Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft

EMNLP 2024finding

In the Minecraft Collaborative Building Task, two players collaborate: an Architect (A) provides instructions to a Builder (B) to assemble a specified structure using 3D blocks. In this work, we investigate the use of large language models (LLMs) to predict the sequence of actions taken by the Build…

Cited by 2SourcePDFScholar
2024

Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies

COLING 2024main

In collaborative goal-oriented settings, the participants are not only interested in achieving a successful outcome, but do also implicitly negotiate the effort they put into the interaction (by adapting to each other). In this work, we propose a challenging interactive reference game that requires…

2023

Images in Language Space: Exploring the Suitability of Large Language Models for Vision & Language Tasks

ACL 2023findings

Large language models have demonstrated robust performance on various language tasks using zero-shot or few-shot learning paradigms. While being actively researched, multimodal models that can additionally handle images as input have yet to catch up in size and generality with language-only models.…

2023

Yes, this Way! Learning to Ground Referring Expressions into Actions with Intra-episodic Feedback from Supportive Teachers

ACL 2023findings

The ability to pick up on language signals in an ongoing interaction is crucial for future machine learning models to collaborate and interact with humans naturally. In this paper, we present an initial study that evaluates intra-episodic feedback given in a collaborative setting. We use a referenti…

2023

clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational Agents

EMNLP 2023long main

Recent work has proposed a methodology for the systematic evaluation of "Situated Language Understanding Agents" --- agents that operate in rich linguistic and non-linguistic contexts --- through testing them in carefully constructed interactive settings. Other recent work has argued that Large Lang…

Cited by 62SourcecodeScholar
2022

MM-Claims: A Dataset for Multimodal Claim Detection in Social Media

NAACL 2022findings

In recent years, the problem of misinformation on the web has become widespread across languages, countries, and various social media platforms. Although there has been much work on automated fake news detection, the role of images and their variety are not well explored. In this paper, we investiga…