← Search

Xiyuan Wei

6 accepted papers

2026

A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization

ICML 2026poster

This paper studies optimization for a family of problems termed **compositional entropic risk minimization**, in which each data's loss is formulated as a Log-Expectation-Exponential (Log-E-Exp) function. The Log-E-Exp formulation serves as an abstraction of the Log-Sum-Exponential (LogSumExp) funct…

Cited by 0SourceScholar
2026

NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization

ICLR 2026poster

Accurately estimating the normalization term (also known as the partition function) in the contrastive loss is a central challenge for training Contrastive Language-Image Pre-training (CLIP) models. Conventional methods rely on large batches for approximation, demanding substantial computational res…

Cited by 0SourcecodeScholar
2026

Statistical Consistency and Generalization of Contrastive Representation Learning

ICML 2026poster

Contrastive representation learning (CRL) underpins many modern foundation models. Despite recent theoretical progress, existing analyses suffer from several key limitations: (i) the statistical consistency of CRL remains poorly understood; (ii) available generalization bounds deteriorate as the num…

Cited by 0SourceScholar
2025

Advancing Interpretability of CLIP Representations with Concept Surrogate Model

NeurIPS 2025poster

Contrastive Language-Image Pre-training (CLIP) generates versatile multimodal embeddings for diverse applications, yet the specific information captured within these representations is not fully understood. Current explainability techniques often target specific tasks, overlooking the rich, general…

Cited by 0SourceScholar
2025

Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws

ICML 2025spotlight

This paper formalizes an emerging learning paradigm that uses a trained model as a reference to guide and enhance the training of a target model through strategic data selection or weighting, named **model steering**. While ad-hoc methods have been used in various contexts, including the training of…

2024

Stability and Generalization of Stochastic Compositional Gradient Descent Algorithms

ICML 2024poster

Many machine learning tasks can be formulated as a stochastic compositional optimization (SCO) problem such as reinforcement learning, AUC maximization and meta-learning, where the objective function involves a nested composition associated with an expectation. Although many studies have been devote…

Cited by 2SourcePDFScholar