← Search

Musashi Hinck

6 accepted papers

2026

LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models

AAAI 2026technical

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbo

Cited by 0SourcePDFScholar
2025

Probing Semantic Routing in Large Mixture-of-Expert Models

EMNLP 2025

In the past year, large ( >100 B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expe

Cited by 0SourcePDFScholar
2024

AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments

EMNLP 2024main

We introduce a three-part framework for constructing persuasive messages, AutoPersuade. First, we curate a large collection of arguments and gather human evaluations of their persuasiveness. Next, we introduce a novel topic model to identify the features of these arguments that influence persuasion.…

2024

Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models

ACL 2024long

Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example, politically-biased LLMs may subtly influence society when they are used…

2024

Why do LLaVA Vision-Language Models Reply to Images in English?

EMNLP 2024finding

We uncover a surprising multilingual bias occurring in a popular class of multimodal vision-language models (VLMs). Including an image in the query to a LLaVA-style VLM significantly increases the likelihood of the model returning an English response, regardless of the language of the query. This pa…

Cited by 4SourcePDFScholar
2023

Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models

NeurIPS 2023poster

In computational social science (CSS), researchers analyze documents to explain social and political phenomena. In most scenarios, CSS researchers first obtain labels for documents and then explain labels using interpretable regression analyses in the second step. One increasingly common way to anno…

Cited by 40SourcePDFScholar