← Search

Andrew Wang

12 accepted papers

2025

Better Training Data Attribution via Better Inverse Hessian-Vector Products

NeurIPS 2025poster

Training data attribution (TDA) provides insights into which training data is responsible for a learned model behavior. Gradient-based TDA methods such as influence functions and unrolled differentiation both involve a computation that resembles an inverse Hessian-vector product (iHVP), which is dif…

Cited by 0SourceScholar
2025

FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External Feedback

NeurIPS 2025poster

Recent studies have shown LLMs possess some ability to improve their responses when given external feedback. However, it remains unclear how effectively and thoroughly these models can incorporate extrinsic feedback. In an ideal scenario, if LLMs receive near-perfect and complete feedback, we would…

Cited by 0SourceScholar
2025

Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data

NeurIPS 2025poster

Grounding language in perception and action is a key challenge when building situated agents that can interact with humans, or other agents, via language. In the past, addressing this challenge has required manually designing the language grounding or curating massive datasets that associate languag…

Cited by 0SourceScholar
2025

InFlux: A Benchmark for Self-Calibration of Dynamic Intrinsics of Video Cameras

NeurIPS 2025poster

Accurately tracking camera intrinsics is crucial for achieving 3D understanding from 2D video. However, most 3D algorithms assume that camera intrinsics stay constant throughout a video, which is often not true for many real-world in-the-wild videos. A major obstacle in this field is a lack of dynam…

Cited by 0SourceScholar
2025

Learning Extrapolative Sequence Transformations from Markov Chains

ICML 2025poster

Most successful applications of deep learning involve similar training and test conditions. However, tasks such as biological sequence design involve searching for sequences that improve desirable properties beyond previously known values, which requires novel hypotheses that \emph{extrapolate} beyo…

2025

RATIONALYST: Pre-training Process-Supervision for Improving Reasoning

ACL 2025long

The reasoning steps generated by LLMs might be incomplete, as they mimic logical leaps common in everyday communication found in their pre-training data: underlying rationales are frequently left implicit (unstated). To address this challenge, we introduce RATIONALYST, a model for process-supervisio…

2024

AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies

EMNLP 2024main

Humans regularly engage in analogical thinking, relating personal experiences to current situations (X is analogous to Y because of Z). Analogical thinking allows humans to solve problems in creative ways, grasp difficult concepts, and articulate ideas more effectively. Can language models (LMs) do…

Cited by 6SourcePDFScholar
2024

Identifying the Risks of LM Agents with an LM-Emulated Sandbox

ICLR 2024spotlight

Recent advances in Language Model (LM) agents and tool use, exemplified by applications like ChatGPT Plugins, enable a rich set of capabilities but also amplify potential risks—such as leaking private data or causing financial losses. Identifying these risks is labor-intensive, necessitating impleme…

2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2023

Learning Belief Representations for Partially Observable Deep RL

ICML 2023poster

Many important real-world Reinforcement Learning (RL) problems involve partial observability and require policies with memory. Unfortunately, standard deep RL algorithms for partially observable settings typically condition on the full history of interactions and are notoriously difficult to train.…

Cited by 12SourcePDFScholar
2022

Robust Classification with Flexible Discriminant Analysis in Heterogeneous Data

ICASSP 2022accepted

Linear and Quadratic Discriminant Analysis are well-known classical methods but can heavily suffer from non-Gaussian distributions and/or contaminated datasets, mainly because of the underlying Gaussian assumption that is not robust. To fill this gap, this paper presents a new robust discriminant an…

Cited by 0SourceScholar