← Search

Yevgeniy Vorobeychik

41 accepted papers

2026

Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees

AAAI 2026technical

Ensuring safety in autonomous systems with vision-based control remains a critical challenge due to the high dimensionality of image inputs and the fact that the relationship between true system state and its visual manifestation is unknown. Existing methods for learning-based control in such setti

Cited by 0SourcePDFScholar
2026

Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated Planning

ICML 2026poster

The integration of external tools has substantially expanded the capabilities of large language model (LLM) agents, but also introduced new attack surfaces beyond prompt injection. In particular, cross-tool description poisoning can manipulate planner-visible tool metadata to steer an agent’s trajec…

Cited by 0SourceScholar
2025

Active Geospatial Search for Efficient Tenant Eviction Outreach

AAAI 2025technical

Tenant evictions threaten housing stability and are a major concern for many cities. An open question concerns whether data-driven methods enhance outreach programs that target at-risk tenants to mitigate their risk of eviction. We propose a novel active geospatial search (AGS) modeling framework fo…

Cited by 0SourcePDFScholar
2025

Active Target Discovery under Uninformative Priors: The Power of Permanent and Transient Memory

NeurIPS 2025poster

In many scientific and engineering fields, where acquiring high-quality data is expensive—such as medical imaging, environmental monitoring, and remote sensing—strategic sampling of unobserved regions based on prior observations is crucial for maximizing discovery rates within a constrained budget.…

Cited by 0SourceScholar
2025

AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

ICLR 2025spotlight

Jailbreak attacks serve as essential red-teaming tools, proactively assessing whether LLMs can behave responsibly and safely in adversarial environments. Despite diverse strategies (e.g., cipher, low-resource language, persuasions, and so on) that have been proposed and shown success, these strategi…

2025

EcoLoRA: Communication-Efficient Federated Fine-Tuning of Large Language Models

EMNLP 2025

To address data locality and privacy restrictions, Federated Learning (FL) has recently been adopted to fine-tune large language models (LLMs), enabling improved performance on various downstream tasks without requiring aggregated data. However, the repeated exchange of model updates in FL can resul

Cited by 0SourcePDFScholar
2025

Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks

ICML 2025poster

Many dynamic decision problems, such as robotic control, involve a series of tasks, many of which are unknown at training time. Typical approaches for these problems, such as multi-task and meta reinforcement learning, do not generalize well when the tasks are diverse. On the other hand, approaches…

2025

Online Feedback Efficient Active Target Discovery in Partially Observable Environments

NeurIPS 2025poster

In various scientific and engineering domains, where data acquisition is costly—such as in medical imaging, environmental monitoring, or remote sensing—strategic sampling from unobserved regions, guided by prior observations, is essential to maximize target discovery within a limited sampling budget…

Cited by 0SourceScholar
2025

To Give or Not to Give? The Impacts of Strategically Withheld Recourse

AISTATS 2025poster

Individuals often aim to reverse undesired outcomes in interactions with automated systems, like loan denials, by either implementing system-recommended actions (recourse), or manipulating their features. While providing recourse benefits users and enhances system utility, it also provides informati…

Cited by 0SourcecodeScholar
2024

Axioms for AI Alignment from Human Feedback

NeurIPS 2024spotlight

In the context of reinforcement learning from human feedback (RLHF), the reward function is generally derived from maximum likelihood estimation of a random utility model based on pairwise comparisons made by humans. The problem of learning a reward function is one of preference aggregation that, we…

Cited by 17SourcePDFScholar
2024

GOMAA-Geo: GOal Modality Agnostic Active Geo-localization

NeurIPS 2024poster

We consider the task of active geo-localization (AGL) in which an agent uses a sequence of visual cues observed during aerial navigation to find a target specified through multiple possible modalities. This could emulate a UAV involved in a search-and-rescue operation navigating through an area, obs…

2024

Providing Fair Recourse over Plausible Groups

AAAI 2024technical

Machine learning models now automate decisions in applications where we may wish to provide recourse to adversely affected individuals. In practice, existing methods to provide recourse return actions that fail to account for latent characteristics that are not captured in the model (e.g., age, sex,…

Cited by 1SourcePDFScholar
2024

RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models

ACL 2024long

Reinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment. Despite its advantages, RLHF relies on human annotators to rank the text, which can introduce potential security vuln…

2024

The Impact of Features Used by Algorithms on Perceptions of Fairness

IJCAI 2024poster

We investigate perceptions of fairness in the choice of features that algorithms use about individuals in a simulated gigwork employment experiment. First, a collection of experimental participants (the selectors) were asked to recommend an algorithm for making employment decisions. Second, a differ…

Cited by 0SourcePDFScholar
2024

Verified Safe Reinforcement Learning for Neural Network Dynamic Models

NeurIPS 2024poster

Learning reliably safe autonomous control is one of the core problems in trustworthy autonomy. However, training a controller that can be formally verified to be safe remains a major challenge. We introduce a novel approach for learning verified safe control policies in nonlinear neural dynamical sy…

2023

A Partially-Supervised Reinforcement Learning Framework for Visual Active Search

NeurIPS 2023poster

Visual active search (VAS) has been proposed as a modeling framework in which visual cues are used to guide exploration, with the goal of identifying regions of interest in a large geospatial area. Its potential applications include identifying hot spots of rare wildlife poaching activity, search-a…

2023

CodeIPPrompt: Intellectual Property Infringement Assessment of Code Language Models

ICML 2023poster

Recent advances in large language models (LMs) have facilitated their ability to synthesize programming code. However, they have also raised concerns about intellectual property (IP) rights violations. Despite the significance of this issue, it has been relatively less explored. In this paper, we ai…

Cited by 36SourcePDFScholar
2023

Exact Verification of ReLU Neural Control Barrier Functions

NeurIPS 2023poster

Control Barrier Functions (CBFs) are a popular approach for safe control of nonlinear systems. In CBF-based control, the desired safety properties of the system are mapped to nonnegativity of a CBF, and the control input is chosen to ensure that the CBF remains nonnegative for all time. Recently, ma…

2023

Incentivizing Recourse through Auditing in Strategic Classification

IJCAI 2023poster

The increasing automation of high-stakes decisions with direct impact on the lives and well-being of individuals raises a number of important considerations. Prominent among these is strategic behavior by individuals hoping to achieve a more desirable outcome. Two forms of such behavior are commonly…

Cited by 7SourcePDFScholar
2023

Neural Lyapunov Control for Discrete-Time Systems

NeurIPS 2023poster

While ensuring stability for linear systems is well understood, it remains a major challenge for nonlinear systems. A general approach in such cases is to compute a combination of a Lyapunov function and an associated control policy. However, finding Lyapunov functions for general nonlinear systems…

2023

Popularizing Fairness: Group Fairness and Individual Welfare

AAAI 2023technical

Group-fair learning methods typically seek to ensure that some measure of prediction efficacy for (often historically) disadvantaged minority groups is comparable to that for the majority of the population. When a principal seeks to adopt a group-fair approach to replace another, the principal may f…

Cited by 0SourcePDFScholar
2023

SlowLiDAR: Increasing the Latency of LiDAR-Based Detection Using Adversarial Examples

CVPR 2023poster

LiDAR-based perception is a central component of autonomous driving, playing a key role in tasks such as vehicle localization and obstacle detection. Since the safety of LiDAR-based perceptual pipelines is critical to safe autonomous driving, a number of past efforts have investigated its vulnerabil…

2022

CROP: Certifying Robust Policies for Reinforcement Learning through Functional Smoothing

ICLR 2022poster

As reinforcement learning (RL) has achieved great success and been even adopted in safety-critical domains such as autonomous vehicles, a range of empirical studies have been conducted to improve its robustness against adversarial attacks. However, how to certify its robustness with theoretical guar…

2022

Learning binary multi-scale games on networks

UAI 2022poster

Network games are a natural modeling framework for strategic interactions of agents whose actions have local impact on others. Recently, a multi-scale network game model has been proposed to capture local effects at multiple network scales, such as among both individuals and groups. We propose a fra…

2022

Manipulating Elections by Changing Voter Perceptions

IJCAI 2022poster

The integrity of elections is central to democratic systems. However, a myriad of malicious actors aspire to influence election outcomes for financial or political benefit. A common means to such ends is by manipulating perceptions of the voting public about select candidates, for example, through m…

Cited by 6SourcePDFScholar
2022

Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum

ICML 2022spotlight

Despite considerable advances in deep reinforcement learning, it has been shown to be highly vulnerable to adversarial perturbations to state observations. Recent efforts that have attempted to improve adversarial robustness of reinforcement learning can nevertheless tolerate only very small perturb…

2022

Solving structured hierarchical games using differential backward induction

UAI 2022poster

From large-scale organizations to decentralized political systems, hierarchical strategic decision making is commonplace. We introduce a novel class of structured hierarchical games (SHGs) that formally capture such hierarchical strategic interactions. In an SHG, each player is a node in a tree, and…

2021

Enhancing Robustness of Neural Networks through Fourier Stabilization

ICML 2021spotlight

Despite the considerable success of neural networks in security settings such as malware detection, such models have proved vulnerable to evasion attacks, in which attackers make slight changes to inputs (e.g., malware) to bypass detection. We propose a novel approach, Fourier stabilization, for des…

Cited by 11SourcePDFScholar
2021

FaceSec: A Fine-Grained Robustness Evaluation Framework for Face Recognition Systems

CVPR 2021poster

We present FACESEC, a framework for fine-grained robustness evaluation of face recognition systems. FACESEC evaluation is performed along four dimensions of adversarial modeling: the nature of perturbation (e.g., pixel-level or face accessories), the attacker's system knowledge (about training data…

Cited by 27PDFcodeScholar
2021

Incentivizing Truthfulness Through Audits in Strategic Classification

AAAI 2021technical

In many societal resource allocation domains, machine learning methods are increasingly used to either score or rank agents in order to decide which ones should receive either resources (e.g., homeless services) or scrutiny (e.g., child welfare investigations) from social services agencies. An agenc…

Cited by 12SourcePDFScholar
2021

Multi-Scale Games: Representing and Solving Games on Networks with Group Structure

AAAI 2021technical

Network games provide a natural machinery to compactly represent strategic interactions among agents whose payoffs exhibit sparsity in their dependence on the actions of others. Besides encoding interaction sparsity, however, real networks often exhibit a multi-scale structure, in which agents can b…

Cited by 4SourcePDFScholar
2020

Defending Against Physically Realizable Attacks on Image Classification

ICLR 2020spotlight

We study the problem of defending deep neural network approaches for image classification from physically realizable attacks. First, we demonstrate that the two most scalable and effective methods for learning robust models, adversarial training with PGD attacks and randomized smoothing, exhibit ver…

Cited by 149SourcecodeScholar
2020

Election Control by Manipulating Issue Significance

UAI 2020poster

Integrity of elections is vital to democratic systems, but it is frequently threatened by malicious actors.The study of algorithmic complexity of the problem of manipulating election outcomes by changing its structural features is known as election control Rothe [2016].One means of election control…

Cited by 5SourcePDFScholar
2020

Robust Spatial-Temporal Incident Prediction

UAI 2020poster

Spatio-temporal incident prediction is a central issue in law enforcement, with applications in fighting crimes like poaching, human trafficking, illegal fishing, burglaries and smuggling. However, state of the art approaches fail to account for evasion in response to predictive models, a common fo…

Cited by 6SourcePDFScholar
2016

Data Poisoning Attacks on Factorization-Based Collaborative Filtering

NeurIPS 2016poster

Recommendation and collaborative filtering systems are important in modern information and e-commerce applications. As these systems are becoming increasingly popular in industry, their outputs could affect business decision making, introducing incentives for an adversarial party to compromise the…

Cited by 444SourcePDFScholar
2015

Scalable Optimization of Randomized Operational Decisions in Adversarial Classification Settings

AISTATS 2015poster

When learning, such as classification, is used in adversarial settings, such as intrusion detection, intelligent adversaries will attempt to evade the resulting policies. The literature on adversarial machine learning aims to develop learning algorithms which are robust to such adversarial evasion,…

Cited by 78SourcePDFScholar