← Search

Zaiwen Wen

12 accepted papers

2026

Constructing Industrial-Scale Optimization Modeling Benchmark

ICML 2026poster

Optimization modeling underpins decision-making in logistics, manufacturing, energy, and finance, yet translating natural-language requirements into correct optimization formulations and solver-executable code remains labor-intensive. Although large language models (LLMs) have been explored for this…

Cited by 0SourceScholar
2026

LMask: Learn to Solve Constrained Routing Problems with Lazy Masking

ICLR 2026poster

Routing problems are canonical combinatorial optimization tasks with wide-ranging applications in logistics, transportation, and supply chain management. However, solving these problems becomes significantly more challenging when complex constraints are involved. In this paper, we propose LMask, a n…

Cited by 0SourcecodeScholar
2026

OptProver: Bridging Olympiad and Optimization through Continual Training in Formal Theorem Proving

ICML 2026poster

Recent advances in formal theorem proving have focused on Olympiad-level mathematics, leaving undergraduate domains largely unexplored. Optimization, fundamental to machine learning, operations research, and scientific computing, remains underserved by existing provers. Its reliance on domain-specif…

Cited by 0SourceScholar
2026

SITA: A Framework for Structure-to-Instance Theorem Autoformalization

AAAI 2026technical

While large language models (LLMs) have shown progress in mathematical reasoning, they still face challenges in formalizing theorems that arise from instantiating abstract structures in concrete settings. With the goal of auto-formalizing mathematical results at the research level, we develop a fram

Cited by 0SourcePDFScholar
2025

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

ICML 2025poster

The memory challenges associated with training Large Language Models (LLMs) have become a critical concern, particularly when using the Adam optimizer. To address this issue, numerous memory-efficient techniques have been proposed, with GaLore standing out as a notable example designed to reduce the…

Cited by 1SourcePDFScholar
2025

Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures

ICLR 2025poster

Parameter-efficient fine-tuning (PEFT) significantly reduces memory costs when adapting large language models (LLMs) for downstream applications. However, traditional first-order (FO) fine-tuning algorithms incur substantial memory overhead due to the need to store activation values for back-propaga…

2025

OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling

ICML 2025poster

Despite the rapid development of large language models (LLMs), a fundamental challenge persists: the lack of high-quality optimization modeling datasets hampers LLMs' robust modeling of practical optimization problems from natural language descriptions (NL). This data scarcity also contributes to th…

2024

An Improved Finite-time Analysis of Temporal Difference Learning with Deep Neural Networks

ICML 2024poster

Temporal difference (TD) learning algorithms with neural network function parameterization have well-established empirical success in many practical large-scale reinforcement learning tasks. However, theoretical understanding of these algorithms remains challenging due to the nonlinearity of the act…

Cited by 1SourcePDFScholar
2021

Enhance Curvature Information by Structured Stochastic Quasi-Newton Methods

CVPR 2021poster

In this paper, we consider stochastic second-order methods for minimizing a finite summation of nonconvex functions. One important key is to find an ingenious but cheap scheme to incorporate local curvature information. Since the true Hessian matrix is often a combination of a cheap part and an expe…

Cited by 10PDFScholar