← Search

Haolun Wu

9 accepted papers

2026

Group Verification-based Policy Optimization for Interactive Coding Agents

ICLR 2026poster

Recent advancements in reinforcement learning from verifiable rewards (RLVR), particularly through Group Relative Policy Optimization (GRPO), have significantly improved the capabilities of large language models (LLMs) for interactive coding agents. However, these methods overlook process-verifiable…

Cited by 0SourceScholar
2026

Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization

ICML 2026poster

Offline black-box optimization aims to discover novel designs with high property scores using only a static dataset, a task fundamentally challenged by the out-of-distribution (OOD) extrapolation problem. Existing approaches typically bifurcate into inverse methods, which struggle with the ill-posed…

Cited by 0SourceScholar
2026

Training Diffusion Language Models for Black-Box Optimization

ICML 2026spotlight

We study offline black-box optimization (BBO), aiming to discover improved designs from an offline dataset of designs and labels, a problem common in robotics, DNA, and materials science with limited labeled samples. While recent work applies autoregressive LLMs to BBO by formatting tasks as natural…

Cited by 0SourceScholar
2025

Aligning Compound AI Systems via System-level DPO

NeurIPS 2025poster

Compound AI systems, comprising multiple interacting components such as LLMs, foundation models, and external tools, have demonstrated remarkable improvements compared to single models in various tasks. To ensure their effective deployment in real-world applications, aligning these systems with huma…

Cited by 0SourceScholar
2025

Logits are All We Need to Adapt Closed Models

ICML 2025poster

Many commercial Large Language Models (LLMs) are often closed-source, limiting developers to prompt tuning for aligning content generation with specific applications. While these models currently do not provide access to token logits, we argue that if such access were available, it would enable more…

2024

Density-based User Representation using Gaussian Process Regression for Multi-interest Personalized Retrieval

NeurIPS 2024poster

Accurate modeling of the diverse and dynamic interests of users remains a significant challenge in the design of personalized recommender systems. Existing user modeling methods, like single-point and multi-point representations, have limitations w.r.t.\ accuracy, diversity, and adaptability. To ove…

Cited by 0SourcePDFScholar
2024

Learning to Extract Structured Entities Using Language Models

EMNLP 2024main

Recent advances in machine learning have significantly impacted the field of information extraction, with Language Models (LMs) playing a pivotal role in extracting structured information from unstructured text. Prior works typically represent information extraction as triplet-centric and use classi…

2024

Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation

ICLR 2024poster

Knowledge distillation aims to train a compact student network using soft supervision from a larger teacher network and hard supervision from ground truths. However, determining an optimal knowledge fusion ratio that balances these supervisory signals remains challenging. Prior methods generally res…

Cited by 3SourcePDFScholar
2021

Knowledge-Enhanced Top-K Recommendation in Poincaré Ball

AAAI 2021technical

Personalized recommender systems are increasingly important as more content and services become available and users struggle to identify what might interest them. Thanks to the ability for providing rich information, knowledge graphs (KGs) are being incorporated to enhance the recommendation perform…

Cited by 42SourcePDFScholar