← Search

Andrew Estornell

12 accepted papers

2026

Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth

ICLR 2026poster

Large Language Models (LLMs) exhibit strong but shallow alignment: they directly refuse harmful queries when a refusal is expected at the very start of an assistant turn, yet this protection collapses once a harmful continuation is underway (either through the adversarial attacks or via harmful assi…

Cited by 0SourceScholar
2026

D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning

ICML 2026poster

Test-time scaling methods such as majority vote aggregation and iterative refinement (e.g., self-reflection or multi-agent inference) improve reasoning performance by leveraging multiple solution samples. However, their efficacy depends not only on raw performance, but critically on the distribution…

Cited by 0SourceScholar
2025

ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration

ICLR 2025poster

Large language models (LLMs) have demonstrated a remarkable ability to serve as general-purpose tools for various language-based tasks. Recent works have demonstrated that the efficacy of such models can be improved through iterative dialog between multiple models. While these paradigms show…

Cited by 0SourcePDFScholar
2025

To Give or Not to Give? The Impacts of Strategically Withheld Recourse

AISTATS 2025poster

Individuals often aim to reverse undesired outcomes in interactions with automated systems, like loan denials, by either implementing system-recommended actions (recourse), or manipulating their features. While providing recourse benefits users and enhances system utility, it also provides informati…

Cited by 0SourcecodeScholar
2024

The Impact of Features Used by Algorithms on Perceptions of Fairness

IJCAI 2024poster

We investigate perceptions of fairness in the choice of features that algorithms use about individuals in a simulated gigwork employment experiment. First, a collection of experimental participants (the selectors) were asked to recommend an algorithm for making employment decisions. Second, a differ…

Cited by 0SourcePDFScholar
2024

User-Creator Feature Polarization in Recommender Systems with Dual Influence

NeurIPS 2024poster

Recommender systems serve the dual purpose of presenting relevant content to users and helping content creators reach their target audience. The dual nature of these systems naturally influences both users and creators: users' preferences are affected by the items they are recommended, while creator…

Cited by 0SourcePDFScholar
2023

Incentivizing Recourse through Auditing in Strategic Classification

IJCAI 2023poster

The increasing automation of high-stakes decisions with direct impact on the lives and well-being of individuals raises a number of important considerations. Prominent among these is strategic behavior by individuals hoping to achieve a more desirable outcome. Two forms of such behavior are commonly…

Cited by 7SourcePDFScholar
2023

Popularizing Fairness: Group Fairness and Individual Welfare

AAAI 2023technical

Group-fair learning methods typically seek to ensure that some measure of prediction efficacy for (often historically) disadvantaged minority groups is comparable to that for the majority of the population. When a principal seeks to adopt a group-fair approach to replace another, the principal may f…

Cited by 0SourcePDFScholar
2022

Manipulating Elections by Changing Voter Perceptions

IJCAI 2022poster

The integrity of elections is central to democratic systems. However, a myriad of malicious actors aspire to influence election outcomes for financial or political benefit. A common means to such ends is by manipulating perceptions of the voting public about select candidates, for example, through m…

Cited by 6SourcePDFScholar
2021

Incentivizing Truthfulness Through Audits in Strategic Classification

AAAI 2021technical

In many societal resource allocation domains, machine learning methods are increasingly used to either score or rank agents in order to decide which ones should receive either resources (e.g., homeless services) or scrutiny (e.g., child welfare investigations) from social services agencies. An agenc…

Cited by 12SourcePDFScholar
2020

Election Control by Manipulating Issue Significance

UAI 2020poster

Integrity of elections is vital to democratic systems, but it is frequently threatened by malicious actors.The study of algorithmic complexity of the problem of manipulating election outcomes by changing its structural features is known as election control Rothe [2016].One means of election control…

Cited by 5SourcePDFScholar