← Search

Ateret Anaby Tavor

8 accepted papers

2025

Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In

NAACL 2025findings

Following the advancement of large language models (LLMs), the development of LLM-based autonomous agents has become prevalent.As a result, the need to understand the security vulnerabilities of these agents has become a critical task. We examine how ReAct agents can be exploited using a straightfor…

Cited by 4SourcePDFScholar
2025

Effective Red-Teaming of Policy-Adherent Agents

EMNLP 2025

Task-oriented LLM-based agents are increasingly used in domains with strict policies, such as refund eligibility or cancellation rules. The challenge lies in ensuring that the agent consistently adheres to these rules and policies, appropriately refusing any request that would violate them, while st

Cited by 0SourcePDFScholar
2025

Exploring Straightforward Methods for Automatic Conversational Red-Teaming

NAACL 2025industry

Large language models (LLMs) are increasingly used in business dialogue systems but they also pose security and ethical risks. Multi-turn conversations, in which context influences the model’s behavior, can be exploited to generate undesired responses. In this paper, we investigate the use of off-th…

Cited by 0SourcePDFScholar
2024

A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios

EMNLP 2024finding

We evaluate the robustness of several large language models on multiple datasets. Robustness here refers to the relative insensitivity of the model’s answers to meaning-preserving variants of their input. Benchmark datasets are constructed by introducing naturally-occurring, non-malicious perturbati…

2023

Reliable and Interpretable Drift Detection in Streams of Short Texts

ACL 2023industry

Data drift is the change in model input data that is one of the key factors leading to machine learning models performance degradation over time. Monitoring drift helps detecting these issues and preventing their harmful consequences. Meaningful drift interpretation is a fundamental step towards eff…

Cited by 16SourcePDFScholar
2023

Text Augmentation Using Dataset Reconstruction for Low-Resource Classification

ACL 2023findings

In the deployment of real-world text classification models, label scarcity is a common problem and as the number of classes increases, this problem becomes even more complex. An approach to addressing this problem is by applying text augmentation methods. One of the more prominent methods involves u…

Cited by 10SourcePDFScholar
2022

Gaining Insights into Unrecognized User Utterances in Task-Oriented Dialog Systems

EMNLP 2022industry

The rapidly growing market demand for automatic dialogue agents capable of goal-oriented behavior has caused many tech-industry leaders to invest considerable efforts into task-oriented dialog systems. The success of these systems is highly dependent on the accuracy of their intent identification –…

Cited by 7SourcePDFScholar
2021

We’ve had this conversation before: A Novel Approach to Measuring Dialog Similarity

EMNLP 2021main

Dialog is a core building block of human natural language interactions. It contains multi-party utterances used to convey information from one party to another in a dynamic and evolving manner. The ability to compare dialogs is beneficial in many real world use cases, such as conversation analytics…

Cited by 7SourcePDFScholar