← Search

Estevam Hruschka

11 accepted papers

2026

Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling

ICLR 2026poster

As language models gain access to external tools through structured function calls, they become increasingly more capable of solving complex, multi-step tasks. However, existing benchmarks for tool-augmented language models (TaLMs) provide insufficient control over factors such as the number of func…

Cited by 0SourcecodeScholar
2025

Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education

NAACL 2025industry

Large Language Models (LLMs) offer the potential to automate hiring by matching job descriptions with candidate resumes, streamlining recruitment processes, and reducing operational costs. However, biases inherent in these models may lead to unfair hiring practices, reinforcing societal prejudices a…

2025

FactLens: Benchmarking Fine-Grained Fact Verification

ACL 2025finding

Large Language Models (LLMs) have shown impressive capability in language generation and understanding, but their tendency to hallucinate and produce factually incorrect information remains a key limitation. To verify LLM-generated contents and claims from other sources, traditional verification app…

2025

From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization

NAACL 2025findings

Although many studies have investigated and reduced hallucinations in large language models (LLMs) for single-document tasks, research on hallucination in multi-document summarization (MDS) tasks remains largely unexplored. Specifically, it is unclear how the challenges arising from handling multipl…

2025

Mixed Signals: Decoding VLMs’ Reasoning and Underlying Bias in Vision-Language Conflict

EMNLP 2025

Vision-language models (VLMs) have demonstrated impressive performance by effectively integrating visual and textual information to solve complex tasks. However, it is not clear how these models reason over the visual and textual data together, nor how the flow of information between modalities is s

2025

Natural Language Processing for Human Resources: A Survey

NAACL 2025industry

Advances in Natural Language Processing (NLP) have the potential to transform HR processes, from recruitment to employee management. While recent breakthroughs in NLP have generated significant interest in its industrial applications, a comprehensive overview of how NLP can be applied across HR acti…

2024

Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks

ACL 2024findings

Large language models (LLMs) are proficient at generating fluent text with minimal task-specific supervision. However, their ability to generate rationales for knowledge-intensive tasks (KITs) remains under-explored. Generating rationales for KIT solutions, such as commonsense multiple-choice QA, re…

Cited by 7SourcePDFScholar
2024

Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

NAACL 2024findings

Large Language Models (LLMs) have demonstrated remarkable capabilities in various NLP tasks. However, previous works have shown these models are sensitive towards prompt wording, and few-shot demonstrations and their order, posing challenges to fair assessment of these models. As these models become…

Cited by 196SourcePDFScholar
2022

Low-resource Entity Set Expansion: A Comprehensive Study on User-generated Text

NAACL 2022findings

Entity set expansion (ESE) aims at obtaining a more complete set of entities given a textual corpus and a seed set of entities of a concept. Although it is a critical task in many NLP applications, existing benchmarks are limited to well-formed text (e.g., Wikipedia) and well-defined concepts (e.g.,…

2022

Low-resource Interactive Active Labeling for Fine-tuning Language Models

EMNLP 2022finding

Recently, active learning (AL) methods have been used to effectively fine-tune pre-trained language models for various NLP tasks such as sentiment analysis and document classification. However, given the task of fine-tuning language models, understanding the impact of different aspects on AL methods…

Cited by 16SourcePDFScholar