AAAI 2026technical0 citations

Polarity-Aware Probing for Quantifying Latent Alignment in Language Models

Sabrina Sadiekh, Elena Ericheva, Chirag Agarwal

Abstract

Advances in unsupervised probes like Contrast‑Consistent Search (CCS), which reveal latent beliefs without token outputs, raise the question of whether they can reliably assess model alignment. We investigate this by examining CCS

BibTeX
@inproceedings{aaai2026_polarityawarepro,
  title = {Polarity-Aware Probing for Quantifying Latent Alignment in Language Models},
  author = {Sabrina Sadiekh and Elena Ericheva and Chirag Agarwal},
  booktitle = {AAAI 2026},
  year = {2026}
}
Polarity-Aware Probing for Quantifying Latent Alignment in Language Models · AAAI 2026