AAAI 2026technical0 citations
Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
Sabrina Sadiekh, Elena Ericheva, Chirag Agarwal
Abstract
Advances in unsupervised probes like Contrast‑Consistent Search (CCS), which reveal latent beliefs without token outputs, raise the question of whether they can reliably assess model alignment. We investigate this by examining CCS
BibTeX
@inproceedings{aaai2026_polarityawarepro,
title = {Polarity-Aware Probing for Quantifying Latent Alignment in Language Models},
author = {Sabrina Sadiekh and Elena Ericheva and Chirag Agarwal},
booktitle = {AAAI 2026},
year = {2026}
}