← Search

Aleksandar Petrov

12 accepted papers

2025

Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models

ICLR 2025poster

Recent research shows that fine-tuning on benign instruction-following data can inadvertently undo the safety alignment process and increase a model's propensity to comply with harmful queries. While instruction-following fine-tuning is important, task-specific fine-tuning-where models are trained o…

Cited by 1SourcePDFScholar
2025

Language-Models-as-a-Service: Overview of a New Paradigm and its Challenges

AAAI 2025technical

Some of the most powerful language models currently are proprietary systems, accessible only via (typically restrictive) web or software programming interfaces. This is the LanguageModels-as-a-Service (LMaaS) paradigm. In contrast with scenarios where full model access is available, as in the case…

Cited by 18SourcePDFScholar
2025

On the Coexistence and Ensembling of Watermarks

NeurIPS 2025poster

Watermarking, the practice of embedding imperceptible information into media such as images, videos, audio, and text, is essential for intellectual property protection, content provenance and attribution. The growing complexity of digital ecosystems necessitates watermarks for different uses to be e…

Cited by 0SourceScholar
2024

Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AI

ICML 2024oral

In the next few years, applications of Generative AI are expected to revolutionize a number of different areas, ranging from science & medicine to education. The potential for these seismic changes has triggered a lively debate about potential risks and resulted in calls for tighter regulation, in p…

Cited by 9SourcePDFScholar
2024

Universal In-Context Approximation By Prompting Fully Recurrent Models

NeurIPS 2024poster

Zero-shot and in-context learning enable solving tasks without model fine-tuning, making them essential for developing generative model solutions. Therefore, it is crucial to understand whether a pretrained model can be prompted to approximate any function, i.e., whether it is a universal in-context…

2024

When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations

ICLR 2024poster

Context-based fine-tuning methods, including prompting, in-context learning, soft prompting (also known as prompt tuning), and prefix-tuning, have gained popularity due to their ability to often match the performance of full fine-tuning with a fraction of the parameters. Despite their empirical succ…

2023

Certifying Ensembles: A General Certification Theory with S-Lipschitzness

ICML 2023poster

Improving and guaranteeing the robustness of deep learning models has been a topic of intense research. Ensembling, which combines several classifiers to provide a better model, has been shown to be beneficial for generalisation, uncertainty estimation, calibration, and mitigating the effects of con…

Cited by 2SourcePDFScholar
2023

Language Model Tokenizers Introduce Unfairness Between Languages

NeurIPS 2023poster

Recent language models have shown impressive multilingual performance, even when not explicitly trained for it. Despite this, there are concerns about the quality of their outputs across different languages. In this paper, we show how disparity in the treatment of different languages arises at the t…

2022

HiddenGems: Efficient safety boundary detection with active learning

IROS 2022poster

Evaluating safety performance in a resource-efficient way is crucial for the development of autonomous systems. Simulation of parameterized scenarios is a popular testing strategy but parameter sweeps can be prohibitively expensive. To address this, we propose HiddenGems: a sample-efficient method f…

Cited by 2SourceScholar
2020

Integrated Benchmarking and Design for Reproducible and Accessible Evaluation of Robotic Agents

IROS 2020poster

As robotics matures and increases in complexity, it is more necessary than ever that robot autonomy research be reproducible. Compared to other sciences, there are specific challenges to benchmarking autonomy, such as the complexity of the software stacks, the variability of the hardware and the rel…

Cited by 17SourceScholar
2020

Learning Camera Miscalibration Detection

ICRA 2020poster

Self-diagnosis and self-repair are some of the key challenges in deploying robotic platforms for long-term real-world applications. One of the issues that can occur to a robot is miscalibration of its sensors due to aging, environmental transients, or external disturbances. Precise calibration lies…

Cited by 20SourcecodeScholar