← Search

Tarun Ram Menta

3 accepted papers

2026

ALPHA: Action-Based Learning for Pluralistic Human Alignment in Large Language Models

AAAI 2026technical

Large language models are widely used, but aligning them with societal values remains challenging. Current approaches often rely on human annotations, which are hard to scale, or synthetic data produced by models that may themselves be misaligned, making it difficult to capture genuine public opinio

Cited by 0SourcePDFScholar
2025

Analyzing Memorization in Large Language Models through the Lens of Model Attribution

NAACL 2025long

Large Language Models (LLMs) are prevalent in modern applications but often memorize training data, leading to privacy breaches and copyright issues. Existing research has mainly focused on post-hoc analyses—such as extracting memorized content or developing memorization metrics—without exploring th…

2022

Improving Attribution Methods by Learning Submodular Functions

AISTATS 2022poster

This work explores the novel idea of learning a submodular scoring function to improve the specificity/selectivity of existing feature attribution methods. Submodular scores are natural for attribution as they are known to accurately model the principle of diminishing returns. A new formulation for…