2026
Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
AAAI 2026technical
Advances in unsupervised probes like Contrast‑Consistent Search (CCS), which reveal latent beliefs without token outputs, raise the question of whether they can reliably assess model alignment. We investigate this by examining CCS