← Search

Maximilian Dax

5 accepted papers

2025

Reparameterized LLM Training via Orthogonal Equivalence Transformation

NeurIPS 2025poster

While large language models (LLMs) are driving the rapid advancement of artificial intelligence, effectively and reliably training these large models remains one of the field's most significant challenges. To address this challenge, we propose POET, a novel reParameterized training algorithm that us…

Cited by 0SourceScholar
2023

Flow Matching for Scalable Simulation-Based Inference

NeurIPS 2023poster

Neural posterior estimation methods based on discrete normalizing flows have become established tools for simulation-based inference (SBI), but scaling them to high-dimensional problems can be challenging. Building on recent advances in generative modeling, we here present flow matching posterior es…

2022

Group equivariant neural posterior estimation

ICLR 2022poster

Simulation-based inference with conditional neural density estimators is a powerful approach to solving inverse problems in science. However, these methods typically treat the underlying forward model as a black box, with no way to exploit geometric properties such as equivariances. Equivariances ar…

2021

Explicitly Modeled Attention Maps for Image Classification

AAAI 2021technical

Self-attention networks have shown remarkable progress in computer vision tasks such as image classification. The main benefit of the self-attention mechanism is the ability to capture long-range feature interactions in attention-maps. However, the computation of attention-maps requires a learnable…

Cited by 13SourcePDFScholar
2019

DeepUSPS: Deep Robust Unsupervised Saliency Prediction via Self-supervision

NeurIPS 2019poster

Deep neural network (DNN) based salient object detection in images based on high-quality labels is expensive. Alternative unsupervised approaches rely on careful selection of multiple handcrafted saliency methods to generate noisy pseudo-ground-truth labels. In this work, we propose a two-stage mech…

Cited by 172SourcePDFScholar