← Search

Georgios Chochlakis

8 accepted papers

2025

Aggregation Artifacts in Subjective Tasks Collapse Large Language Models’ Posteriors

NAACL 2025long

In-context Learning (ICL) has become the primary method for performing natural language tasks with Large Language Models (LLMs). The knowledge acquired during pre-training is crucial for this few-shot capability, providing the model with task priors. However, recent studies have shown that ICL predo…

2025

Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts

EMNLP 2025

Modeling complex subjective tasks in Natural Language Processing, such as recognizing emotion and morality, is considerably challenging due to significant variation in human annotations. This variation often reflects reasonable differences in semantic interpretations rather than mere noise, necessit

2025

Large Language Models Do Multi-Label Classification Differently

EMNLP 2025

Multi-label classification is prevalent in real-world settings, but the behavior of Large Language Models (LLMs) in this setting is understudied. We investigate how autoregressive LLMs perform multi-label classification, focusing on subjective tasks, by analyzing the output distributions of the mode

2025

Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks

ICASSP 2025accepted

In-Context Learning (ICL) in Large Language Models (LLM) has emerged as the dominant technique for performing natural language tasks, as it does not require updating the model parameters with gradient-based methods. ICL promises to "adapt" the LLM to perform the present task at a competitive or stat…

Cited by 0SourceScholar
2024

CVAT-BWV: A Web-Based Video Annotation Platform for Police Body-Worn Video

IJCAI 2024poster

We introduce an open-source platform for annotating body-worn video (BWV) footage aimed at enhancing transparency and accountability in policing. Despite the widespread adoption of BWVs in police departments, analyzing the vast amount of footage generated has presented significant challenges. This i…

2023

Leveraging Label Correlations in a Multi-Label Setting: a Case Study in Emotion

ICASSP 2023accepted

Detecting emotions expressed in text has become critical to a range of fields. In this work, we investigate ways to exploit label correlations in multi-label emotion recognition models to improve emotion detection. First, we develop two modeling approaches to the problem in order to capture word ass…

Cited by 0SourceScholar
2023

Using Emotion Embeddings to Transfer Knowledge between Emotions, Languages, and Annotation Formats

ICASSP 2023accepted

The need for emotional inference from text continues to diversify as more and more disciplines integrate emotions into their theories and applications. These needs include inferring different emotion types, handling multiple languages, and different annotation formats. A shared model between differe…

Cited by 0SourceScholar
2022

CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks

NeurIPS 2022accept

Current state-of-the-art vision-and-language models are evaluated on tasks either individually or in a multi-task setting, overlooking the challenges of continually learning (CL) tasks as they arrive. Existing CL benchmarks have facilitated research on task adaptation and mitigating "catastrophic fo…