← Search

Marcos Zampieri

22 accepted papers

2026

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments

ICRA 2026poster

Large Vision-Language Models (VLMs) have demonstrated potential in enhancing mobile robot navigation in human-centric environments by understanding contextual cues, human intentions, and social dynamics while exhibiting reasoning capabilities. However, their computational complexity and limited sens…

2025

ARTICLE: Annotator Reliability Through In-Context Learning

AAAI 2025technical

Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to dist…

2025

ARTICLE: Annotator Reliability Through In-Context Learning (Student Abstract)

AAAI 2025technical

Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to dist…

2025

Bayelemabaga: Creating Resources for Bambara NLP

NAACL 2025long

Data curation for under-resource languages enables the development of more accurate and culturally sensitive natural language processing models. However, the scarcity of well-structured multilingual datasets remains a challenge for advancing machine translation in these languages, especially for Afr…

Cited by 0SourcePDFScholar
2025

MojoBench: Language Modeling and Benchmarks for Mojo

NAACL 2025findings

The recently introduced Mojo programming language (PL) by Modular, has received significant attention in the scientific community due to its claimed significant speed boost over Python. Despite advancements in code Large Language Models (LLMs) across various PLs, Mojo remains unexplored in this cont…

2025

Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations

EMNLP 2025

Language transfer is an important topic of research in second language acquisition and computational linguistics. The availability of suitable learner corpora is paramount for the study of second language acquisition (SLA) and language transfer. However, curating learner corpora is a challenging end

2025

mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation

NAACL 2025long

Recent advancements in large language models (LLMs) have significantly enhanced code generation from natural language prompts. The HumanEval Benchmark, developed by OpenAI, remains the most widely used code generation benchmark. However, this and other Code LLM benchmarks face critical limitations,…

2024

A Survey of Multimodal Sarcasm Detection

IJCAI 2024poster

Sarcasm is a rhetorical device that is used to convey the opposite of the literal meaning of an utterance. Sarcasm is widely used on social media and other forms of computer-mediated communication motivating the use of computational models to identify it automatically. While the clear majority of ap…

Cited by 4SourcePDFScholar
2024

Language Variety Identification with True Labels

COLING 2024main

Language identification is an important first step in many NLP applications. Most publicly available language identification datasets, however, are compiled under the assumption that the gold label of each instance is determined by where texts are retrieved from. Research has shown that this is a pr…

2024

MentalHelp: A Multi-Task Dataset for Mental Health in Social Media

COLING 2024main

Early detection of mental health disorders is an essential step in treating and preventing mental health conditions. Computational approaches have been applied to users’ social media profiles in an attempt to identify various mental health conditions such as depression, PTSD, schizophrenia, and eati…

Cited by 6SourcePDFScholar
2024

Native Language Identification in Texts: A Survey

NAACL 2024long

We present the first comprehensive survey of Native Language Identification (NLI) applied to texts. NLI is the task of automatically identifying an author’s native language (L1) based on their second language (L2) production. NLI is an important task with practical applications in second language te…

Cited by 6SourcePDFScholar
2024

Rater Cohesion and Quality from a Vicarious Perspective

EMNLP 2024finding

Human feedback is essential for building human-centered AI systems across domains where disagreement is prevalent, such as AI safety, content moderation, or sentiment analysis. Many disagreements, particularly in politically charged settings, arise because raters have opposing values or beliefs. Vic…

2023

Target-Based Offensive Language Identification

ACL 2023short

We present TBO, a new dataset for Target-based Offensive language identification. TBO contains post-level annotations regarding the harmfulness of an offensive post and token-level annotations comprising of the target and the offensive argument expression. Popular offensive language identification d…

2023

Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive

EMNLP 2023long main

Offensive speech detection is a key component of content moderation. However, what is offensive can be highly subjective. This paper investigates how machine and human moderators disagree on what is offensive when it comes to real-world social web political discourse. We show that (1) there is exten…

Cited by 0SourcecodeScholar
2022

ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification

COLING 2022main

Lexical simplification (LS) is the task of automatically replacing complex words for easier ones making texts more accessible to various target populations (e.g. individuals with low literacy, individuals with learning disabilities, second language learners). To train and test models, LS systems usu…

Cited by 15SourcePDFScholar
2021

A Computational Exploration of Pejorative Language in Social Media

EMNLP 2021finding

In this paper we study pejorative language, an under-explored topic in computational linguistics. Unlike existing models of offensive language and hate speech, pejorative language manifests itself primarily at the lexical level, and describes a word that is used with a negative connotation, making i…

Cited by 22SourcePDFScholar
2021

Handling Extreme Class Imbalance in Technical Logbook Datasets

ACL 2021long

Technical logbooks are a challenging and under-explored text type in automated event identification. These texts are typically short and written in non-standard yet technical language, posing challenges to off-the-shelf NLP pipelines. The granularity of issue types described in these datasets additi…

Cited by 14SourcePDFScholar
2021

fBERT: A Neural Transformer for Identifying Offensive Content

EMNLP 2021finding

Transformer-based models such as BERT, XLNET, and XLM-R have achieved state-of-the-art performance across various NLP tasks including the identification of offensive language and hate speech, an important problem in social media. In this paper, we present fBERT, a BERT model retrained on SOLID, the…

Cited by 62SourcePDFScholar
2020

MaintNet: A Collaborative Open-Source Library for Predictive Maintenance Language Resources

COLING 2020system demonstrations

Maintenance record logbooks are an emerging text type in NLP. An important part of them typically consist of free text with many domain specific technical terms, abbreviations, and non-standard spelling and grammar. This poses difficulties for NLP pipelines trained on standard corpora. Analyzing and…

Marcos Zampieri — accepted AI-conference papers · AIConfPaper