← Search

Hanyu Zhang

5 accepted papers

2025

Controlling Large Language Models Through Concept Activation Vectors

AAAI 2025technical

As large language models (LLMs) are widely deployed across various domains, the ability to control their generated outputs has become more critical. This control involves aligning LLMs outputs with human values and ethical principles or customizing LLMs on specific topics or styles for individual us…

Cited by 1SourcePDFScholar
2025

Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models

ACL 2025long

Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful technique for aligning large language models (LLMs) with human preferences. However, effectively aligning LLMs with diverse human preferences remains a significant challenge, particularly when they are conflict. To address t…

2024

Consistency of Dictionary-Based Manifold Learning

AISTATS 2024poster

We analyze a paradigm for interpretable Manifold Learning for scientific data analysis, whereby one parametrizes a manifold with d smooth functions from a scientist-provided dictionary of meaningful, domain-related functions. When such a parametrization exists, we provide an algorithm for finding it…

Cited by 2SourcePDFScholar