← Search

Adel Javanmard

14 accepted papers

2026

Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models

ICML 2026poster

Large Language Models (LLMs) are pretrained on massive datasets and later instruction-tuned via supervised fine-tuning (SFT) or reinforcement learning (RL). Best practices emphasize large, diverse pretraining data, whereas post-training operates differently: SFT relies on smaller, high-quality datas…

Cited by 0SourceScholar
2025

DeepCrossAttention: Supercharging Transformer Residual Connections

ICML 2025poster

Transformer networks have achieved remarkable success across diverse domains, leveraging a variety of architectural innovations, including residual connections. However, traditional residual connections, which simply sum the outputs of previous layers, can dilute crucial information. This work intro…

Cited by 0SourcePDFScholar
2025

Improving the Variance of Differentially Private Randomized Experiments through Clustering

ICML 2025poster

Estimating causal effects from randomized experiments is only possible if participants are willing to disclose their potentially sensitive responses. Differential privacy, a widely used framework for ensuring an algorithm’s privacy guarantees, can encourage participants to share their responses with…

Cited by 0SourcePDFScholar
2025

Integer Programming for Generalized Causal Bootstrap Designs

ICML 2025poster

In experimental causal inference, we distinguish between two sources of uncertainty: design uncertainty, due to the treatment assignment mechanism, and sampling uncertainty, when the sample is drawn from a super-population. This distinction matters in settings with small fixed samples and heterogene…

Cited by 0SourcePDFScholar
2025

Retraining with Predicted Hard Labels Provably Increases Model Accuracy

ICML 2025poster

The performance of a model trained with noisy labels is often improved by simply *retraining* the model with its *own predicted hard labels* (i.e., $1$/$0$ labels). Yet, a detailed theoretical characterization of this phenomenon is lacking. In this paper, we theoretically analyze retraining in a lin…

Cited by 2SourcePDFScholar
2025

Robust Feature Learning for Multi-Index Models in High Dimensions

ICLR 2025poster

Recently, there have been numerous studies on feature learning with neural networks, specifically on learning single- and multi-index models where the target is a function of a low-dimensional projection of the input. Prior works have shown that in high dimensions, the majority of the compute and da…

2025

Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing

NeurIPS 2025poster

Retraining a model using its own predictions together with the original, potentially noisy labels is a well-known strategy for improving the model’s performance. While prior works have demonstrated the benefits of specific heuristic retraining schemes, the question of how to optimally combine the mo…

Cited by 0SourceScholar
2024

Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions

ICLR 2024poster

Due to the rise of privacy concerns, in many practical applications, the training data is aggregated before being shared with the learner to protect the privacy of users' sensitive responses. In an aggregate learning framework, the dataset is grouped into bags of samples, where each bag is available…

Cited by 2SourcePDFScholar
2024

PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses

ICML 2024spotlight

This work studies algorithms for learning from aggregate responses. We focus on the construction of aggregation sets (called *bags* in the literature) for event-level loss functions. We prove for linear regression and generalized linear models (GLMs) that the optimal bagging problem reduces to one-d…

Cited by 2SourcePDFScholar
2023

Anonymous Learning via Look-Alike Clustering: A Precise Analysis of Model Generalization

NeurIPS 2023poster

While personalized recommendations systems have become increasingly popular, ensuring user data protection remains a top concern in the development of these learning systems. A common approach to enhancing privacy involves training models using anonymous data rather than individual data. In this pap…

Cited by 0SourcePDFScholar
2023

Learning Rate Schedules in the Presence of Distribution Shift

ICML 2023poster

We design learning rate schedules that minimize regret for SGD-based online learning in the presence of a changing data distribution. We fully characterize the optimal learning rate schedule for online linear regression via a novel analysis with stochastic differential equations. For general convex…

2021

Fundamental Tradeoffs in Distributionally Adversarial Training

ICML 2021spotlight

Adversarial training is among the most effective techniques to improve robustness of models against adversarial perturbations. However, the full effect of this approach on models is not well understood. For example, while adversarial training can reduce the adversarial risk (prediction error against…

Cited by 29SourcePDFScholar
2019

Dynamic Incentive-Aware Learning: Robust Pricing in Contextual Auctions

NeurIPS 2019poster

Motivated by pricing in ad exchange markets, we consider the problem of robust learning of reserve prices against strategic buyers in repeated contextual second-price auctions. Buyers' valuations \new{for} an item depend on the context that describes the item. However, the seller is not aware…

Cited by 113SourcePDFScholar