← Search

Thien Q. Tran

2 accepted papers

2024

Stepwise Alignment for Constrained Language Model Policy Optimization

NeurIPS 2024poster

Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs). This paper formulates human value alignment as an optimization problem of the language model policy to maximize reward under a safety constraint, and then proposes…

2022

Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model Explanation

AAAI 2022technical

We aim to explain a black-box classifier with the form: "data X is classified as class Y because X has A, B and does not have C" in which A, B, and C are high-level concepts. The challenge is that we have to discover in an unsupervised manner a set of concepts, i.e., A, B and C, that is useful for e…

Cited by 10SourcePDFScholar