← Search

Julia Hockenmaier

13 accepted papers

2025

Entailment-Preserving First-order Logic Representations in Natural Language Entailment

ACL 2025long

First-order logic (FOL) is often used to represent logical entailment, but determining natural language (NL) entailment using FOL remains a challenge. To address this, we propose the Entailment-Preserving FOL representations (EPF) task and introduce reference-free evaluation metrics for EPF (Entailm…

Cited by 0SourcePDFScholar
2025

The Power of Bullet Lists: A Simple Yet Effective Prompting Approach to Enhancing Spatial Reasoning in Large Language Models

NAACL 2025findings

While large language models (LLMs) are dominating the field of natural language processing, it remains an open question how well these models can perform spatial reasoning. Contrary to recent studies suggesting that LLMs struggle with spatial reasoning tasks, we demonstrate in this paper that a nove…

Cited by 0SourcePDFScholar
2025

Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language Models

EMNLP 2025

Sparse autoencoders (SAEs) have emerged as a powerful analytical tool in mechanistic interpretability for large language models (LLMs), with growing success in applications beyond interpretability. Building on this momentum, we present a novel approach that leverages SAEs to enhance the general in-c

2024

Analyzing the Performance of Large Language Models on Code Summarization

COLING 2024main

Large language models (LLMs) such as Llama 2 perform very well on tasks that involve both natural language and source code, particularly code summarization and code generation. We show that for the task of code summarization, the performance of these models on individual examples often depends on th…

2024

Tutor-ICL: Guiding Large Language Models for Improved In-Context Learning Performance

EMNLP 2024finding

There has been a growing body of work focusing on the in-context learning (ICL) abilities of large language models (LLMs). However, it is an open question how effective ICL can be. This paper presents Tutor-ICL, a simple prompting method for classification tasks inspired by how effective instructors…

2023

A Framework for Bidirectional Decoding: Case Study in Morphological Inflection

EMNLP 2023long findings

Transformer-based encoder-decoder models that generate outputs in a left-to-right fashion have become standard for sequence-to-sequence tasks. In this paper, we propose a framework for decoding that produces sequences from the "outside-in": at each step, the model chooses to generate a token on the…

Cited by 0SourcecodeScholar
2023

Multimedia Generative Script Learning for Task Planning

ACL 2023findings

Goal-oriented generative script learning aims to generate subsequent steps to reach a particular goal, which is an essential task to assist robots or humans in performing stereotypical activities. An important aspect of this process is the ability to capture historical states visually, which provide…

2023

SIR-ABSC: Incorporating Syntax into RoBERTa-based Sentiment Analysis Models with a Special Aggregator Token

EMNLP 2023long findings

We present a simple, but effective method to incorporate syntactic dependency information directly into transformer-based language models (e.g. RoBERTa) for tasks such as Aspect-Based Sentiment Classification (ABSC), where the desired output depends on specific input tokens. In contrast to prior app…

Cited by 0SourceScholar
2017

Phrase Localization and Visual Relationship Detection With Comprehensive Image-Language Cues

ICCV 2017poster

This paper presents a framework for localization or grounding of phrases in images using a large collection of linguistic and visual cues. We model the appearance, size, and position of entity bounding boxes, adjectives that contain attribute information, and spatial relationships between pairs of e…

Cited by 230PDFcodeScholar
2015

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

ICCV 2015poster

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains linking mentions of the same entities in images, as well as 276k manually annotated boundi…

Cited by 2475PDFScholar