← Search

Lilian Weng

5 accepted papers

2025

First-Person Fairness in Chatbots

ICLR 2025spotlight

Evaluating chatbot fairness is crucial given their rapid proliferation, yet typical chatbot tasks (e.g., resume writing, entertainment) diverge from the institutional decision-making tasks (e.g., resume screening) which have traditionally been central to discussion of algorithmic fairness. The open-…

Cited by 8SourcePDFScholar
2025

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

ICLR 2025oral

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test real-world ML engineering skills such as training models, pre…

2024

Rule Based Rewards for Language Model Safety

NeurIPS 2024poster

Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cases related to safety, without precise instructions to human annotators, the data collected may cause the model to beco…

2023

A Holistic Approach to Undesired Content Detection in the Real World

AAAI 2023technical

We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed steps, including the design of content taxonomies and labeling instructions, data…

2020

Automatic Curriculum Learning For Deep RL: A Short Survey

IJCAI 2020poster

Automatic Curriculum Learning (ACL) has become a cornerstone of recent successes in Deep Reinforcement Learning (DRL). These methods shape the learning trajectories of agents by challenging them with tasks adapted to their capacities. In recent years, they have been used to improve sample efficiency…

Cited by 0SourcePDFScholar