← Search

Sophie Fellenz

17 accepted papers

2026

Formally Exploring Visual Anomaly Detection Evaluation Metrics

ICML 2026poster

Inaccurate Visual Anomaly Detection (VAD) can lead to critical failures in safety-sensitive domains, including autonomous navigation and industrial surveillance. With the increasing abundance and rapid proliferation of VAD algorithms, their reliable evaluation has become increasingly important and c…

Cited by 0SourceScholar
2026

Heavy-tailed Physics-Informed Neural Networks

ICML 2026poster

Physics-informed neural networks (PINNs) enforce physical laws by minimizing partial differential equation (PDE) residuals and auxiliary constraints. Standard training relies on a mean-squared error (MSE) objective, which implicitly assumes independent Gaussian residuals with a fixed global variance…

Cited by 0SourceScholar
2026

Landmark-Guided Policy Optimization for Multi-Objective Language Model Selection

ICML 2026poster

Selecting a pretrained large language model (LLM) to fine-tune for a task-specific dataset can be time-consuming and costly. With several candidate models available to choose from, varying in size, architecture, and pretraining data, finding the best model for a specific task often involves extensiv…

Cited by 0SourceScholar
2026

Physics-Constrained Fine-Tuning of Flow-Matching Models for Generation and Inverse Problems

ICLR 2026poster

We present a framework for fine-tuning flow-matching generative models to enforce physical constraints and solve inverse problems in scientific systems. Starting from a model trained on low-fidelity or observational data, we apply a differentiable post-training procedure that minimizes weak-form res…

Cited by 0SourceScholar
2026

Physics-Informed Residual Flows

ICML 2026poster

Physics-Informed Neural Networks (PINNs) embed physical laws into deep learning models. However, conventional PINNs often suffer from failure modes leading to inaccurate solutions. We trace these failure modes to two structural pathologies: gradient shattering, where gradients degrade with depth and…

Cited by 0SourceScholar
2026

Reimagining Anomalies: What If Anomalies Were Normal?

AAAI 2026technical

Deep learning-based methods have achieved a breakthrough in image anomaly detection, but their complexity introduces a considerable challenge to understanding why an instance is predicted to be anomalous. We introduce a novel explanation method that generates multiple alternative modifications for e

Cited by 0SourcePDFScholar
2026

Skipping the Zeros in Diffusion Models for Sparse Data Generation

ICML 2026poster

Diffusion models (DMs) excel on dense continuous data, but are not designed for sparse continuous data. They do not model exact zeros that represent the deliberate absence of a signal. As a result, they erase sparsity patterns and perform unnecessary computation on mostly zero entries. With Sparsity…

Cited by 0SourceScholar
2026

TORA: Train Once, Realign Anytime for Offline Multi-Objective Reinforcement Learning

AAAI 2026technical

Intelligent agents in real-world applications must adapt their behavior to changing contexts and user preferences. For example, planning a road trip requires considering both travel time and cost. Multi-objective reinforcement learning (MORL) provides a principled approach to navigate such trade-of

Cited by 0SourcePDFScholar
2025

Mitigating Spurious Features in Contrastive Learning with Spectral Regularization

NeurIPS 2025poster

Neural networks generally prefer simple and easy-to-learn features. When these features are spuriously correlated with the labels, the network's performance can suffer, particularly for underrepresented classes or concepts. Self-supervised representation learning methods, such as contrastive learnin…

Cited by 0SourcecodeScholar
2025

NoBOOM: Chemical Process Datasets for Industrial Anomaly Detection

NeurIPS 2025poster

Monitoring chemical processes is essential to prevent catastrophic failures, optimize costs and profits, and ensure the safety of employees and the environment. A key component of modern monitoring systems is the automated detection of anomalies in sensor data over time, called time series, enablin…

Cited by 0SourceScholar
2025

Tethering Broken Themes: Aligning Neural Topic Models with Labels and Authors

NAACL 2025findings

Topic models are a popular approach for extracting semantic information from large document collections. However, recent studies suggest that the topics generated by these models often do not align well with human intentions. Although metadata such as labels and authorship information are available,…

2024

Characterizing Text Datasets with Psycholinguistic Features

EMNLP 2024finding

Fine-tuning pretrained language models on task-specific data is a common practice in Natural Language Processing (NLP) applications. However, the number of pretrained models available to choose from can be very large, and it remains unclear how to select the optimal model without spending considerab…

2024

Ethics in Action: Training Reinforcement Learning Agents for Moral Decision-making In Text-based Adventure Games

AISTATS 2024poster

Reinforcement Learning (RL) has demonstrated its potential in solving goal-oriented sequential tasks. However, with the increasing capabilities of RL agents, ensuring morally responsible agent behavior is becoming a pressing concern. Previous approaches have included moral considerations by statical…

2024

Evaluating Dynamic Topic Models

ACL 2024long

There is a lack of quantitative measures to evaluate the progression of topics through time in dynamic topic models (DTMs). Filling this gap, we propose a novel evaluation measure for DTMs that analyzes the changes in the quality of each topic over time. Additionally, we propose an extension combini…

Cited by 2SourcePDFScholar
2024

Text Style Transfer Evaluation Using Large Language Models

COLING 2024main

Evaluating Text Style Transfer (TST) is a complex task due to its multi-faceted nature. The quality of the generated text is measured based on challenging factors, such as style transfer accuracy, content preservation, and overall fluency. While human evaluation is considered to be the gold standard…

Cited by 11SourcePDFScholar
2023

A Call for Standardization and Validation of Text Style Transfer Evaluation

ACL 2023findings

Text Style Transfer (TST) evaluation is, in practice, inconsistent. Therefore, we conduct a meta-analysis on human and automated TST evaluation and experimentation that thoroughly examines existing literature in the field. The meta-analysis reveals a substantial standardization gap in human and auto…

Cited by 11SourcePDFScholar