← Search

Marcello Bullo

2 accepted papers

2026

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

AAAI 2026technical

Uniform-reward reinforcement learning from human feedback (RLHF), which trains a single reward model to represent the preferences of all annotators, fails to capture the diversity of opinions across sub-populations, inadvertently favoring dominant groups. The state-of-the-art, MaxMin-RLHF, addresses

Cited by 0SourcePDFScholar
2026

Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality

ICLR 2026poster

While test-time scaling with verification has shown promise in improving the performance of large language models (LLMs), role of the verifier and its imperfections remain underexplored. The effect of verification manifests through interactions of three quantities: (i) the generator’s *coverage*, (i…

Cited by 0SourcecodeScholar