← Search

Jiahai Feng

7 accepted papers

2026

Learning a Generative Meta-Model of LLM Activations

ICML 2026poster

Existing approaches for manipulating neural network activations, such as PCA and SAEs, rely on strong assumptions about activation structure. We develop a generative approach that models activations with diffusion, that makes minimal assumptions and improves with data and model scale. We use this ac…

Cited by 0SourceScholar
2025

Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts

ICML 2025poster

Pretrained language models (LMs) can generalize to implications of facts that they are finetuned on. For example, if finetuned on "John Doe lives in Tokyo," LMs correctly answer "What language do the people in John Doe's city speak?'' with "Japanese''. However, little is known about the mechanisms t…

2025

Monitoring Latent World States in Language Models with Propositional Probes

ICLR 2025spotlight

Language models (LMs) are susceptible to bias, sycophancy, backdoors, and other tendencies that lead to unfaithful responses to the input context. Interpreting internal states of LMs could help monitor and correct unfaithful behavior. We hypothesize that LMs faithfully represent their input contexts…

2024

Learning Grounded Action Abstractions from Language

ICLR 2024poster

Effective planning in the real world requires not only world knowledge, but the ability to leverage that knowledge to build the right representation of the task at hand. Decades of hierarchical planning techniques have used domain-specific temporal action abstractions to support efficient and accura…

Cited by 5SourcePDFScholar
2020

AI Feynman 2.0: Pareto-optimal symbolic regression exploiting graph modularity

NeurIPS 2020oral

We present an improved method for symbolic regression that seeks to fit data to formulas that are Pareto-optimal, in the sense of having the best accuracy for a given complexity. It improves on the previous state-of-the-art by typically being orders of magnitude more robust toward noise and bad data…

Cited by 272SourcePDFScholar