← Search

Puria Radmard

5 accepted papers

2026

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

ICML 2026spotlight

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may internalise corresponding behavioural priors, giving rise to…

Cited by 20SourceScholar
2026

Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors

ICLR 2026poster

Discovering the neural mechanisms underpinning cognition is one of the grand challenges of neuroscience. Addressing this challenge greatly benefits from specific hypotheses about the underlying neural network dynamics. However, previous approaches bridging neural network dynamics and cognitive behav…

Cited by 0SourceScholar
2025

Large language models can learn and generalize steganographic chain-of-thought under process supervision

NeurIPS 2025poster

Chain-of-thought (CoT) reasoning not only enhances large language model performance but also provides critical insights into decision-making processes, marking it as a useful tool for monitoring model intent and planning. By proactively preventing models from acting on CoT indicating misaligned or h…

Cited by 0SourceScholar
2024

Recurrent neural network dynamical systems for biological vision

NeurIPS 2024spotlight

In neuroscience, recurrent neural networks (RNNs) are modeled as continuous-time dynamical systems to more accurately reflect the dynamics inherent in biological circuits. However, convolutional neural networks (CNNs) remain the preferred architecture in vision neuroscience due to their ability to e…

Cited by 1SourcePDFScholar