← Search

Peng Cui

62 accepted papers

2026

Error Slice Discovery via Manifold Compactness

AAAI 2026technical

Despite the great performance of deep learning models in many areas, they still make mistakes and underperform on certain subsets of data, i.e. error slices. Given a trained model, it is important to identify its semantically coherent error slices that are easy to interpret, which is referred to as

Cited by 0SourcePDFScholar
2026

Generating Risky Samples with Conformity Constraints via Diffusion Models

AAAI 2026technical

Although neural networks achieve promising performance in many tasks, they may still fail when encountering some examples and bring about risks to applications. To discover risky samples, previous literature attempts to search for patterns of risky samples within existing datasets or inject perturba

Cited by 0SourcePDFScholar
2026

Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification

ICML 2026poster

As LLM-powered agents have been used for high-stakes decision-making, such as clinical diagnosis, it becomes critical to develop reliable verification of their decisions to facilitate trustworthy deployment. Yet, existing verifiers usually underperform owing to a lack of domain knowledge and limited…

Cited by 0SourceScholar
2026

MiniX: Mitigating Low-Rank Collapse and Attention Bottlenecks in Tabular Foundation Models

ICML 2026poster

Recent tabular foundation models routinely match or surpass strong tree ensembles and specialized deep architectures, yet their numeric embeddings remain a bottleneck. We diagnose a low-rank collapse induced by the prevalent linear+ID scheme and introduce RaBEL, a compact Radial Basis Embedding Laye…

Cited by 0SourceScholar
2026

Unveiling Prior-data Fitted Networks on Causal Effect Estimation: Pre-training or Finetuning?

ICML 2026poster

Amortized causal inference via Prior-data Fitted Networks (PFNs) has emerged as a promising paradigm, enabling zero-shot estimation of causal effects without the need for dataset-specific model tuning. However, the principled effectiveness of unified pre-training across general interventional regime…

Cited by 0SourceScholar
2025

COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts

CVPR 2025highlight

Current object detectors often suffer significant performance degradation in real-world applications when encountering distributional shifts, posing serious risks in high-stakes domains such as autonomous driving and medical diagnosis. Consequently, the out-of-distribution (OOD) generalization capab…

Cited by 0SourcePDFScholar
2025

Environment Inference for Learning Generalizable Dynamical System

NeurIPS 2025spotlight

Data-driven methods offer efficient and robust solutions for analyzing complex dynamical systems but rely on the assumption of I.I.D. data, driving the development of generalization techniques for handling environmental differences. These techniques, however, are limited by their dependence on envir…

Cited by 0SourceScholar
2025

Going Beyond Static: Understanding Shifts with Time-Series Attribution

ICLR 2025poster

Distribution shifts in time-series data are complex due to temporal dependencies, multivariable interactions, and trend changes. However, robust methods often rely on structural assumptions that lack thorough empirical validation, limiting their practical applicability. In order to support an empi…

Cited by 0SourcePDFScholar
2025

Grammar Control in Dialogue Response Generation for Language Learning Chatbots

NAACL 2025long

Chatbots based on large language models offer cheap conversation practice opportunities for language learners. However, they are hard to control for linguistic forms that correspond to learners’ current needs, such as grammar. We control grammar in chatbot conversation practice by grounding a dialog…

2025

Improving Accuracy and Calibration via Differentiated Deep Mutual Learning

CVPR 2025poster

Deep Neural Networks (DNNs) have achieved remarkable success in a variety of tasks, particularly in terms of prediction accuracy. However, in real-world scenarios, especially in safety-critical applications, accuracy alone is insufficient; reliable uncertainty estimates are essential. Modern DNNs, o…

Cited by 0SourcePDFScholar
2025

Investigating the Zone of Proximal Development of Language Models for In-Context Learning

NAACL 2025findings

In this paper, we introduce a learning analytics framework to analyze the in-context learning (ICL) behavior of large language models (LLMs) through the lens of the Zone of Proximal Development (ZPD), an established theory in educational psychology. ZPD delineates the range of tasks a learner can ac…

2025

ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction

ICCV 2025poster

Recently, there has been gradually more attention paid to Out-of-Distribution (OOD) performance prediction, whose goal is to predict the performance of trained models on unlabeled OOD test datasets, so that we could better leverage and deploy off-the-shelf trained models in risk-sensitive scenarios.…

2025

On the Out-Of-Distribution Generalization of Large Multimodal Models

CVPR 2025poster

We investigate the generalization boundaries of current Large Multimodal Models (LMMs) via comprehensive evaluation under out-of-distribution scenarios and domain-specific tasks. We evaluate their zero-shot generalization across synthetic images, real-world distributional shifts, and specialized dat…

2025

Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph

ICML 2025poster

Graph Neural Networks (GNNs) are widely used for node classification tasks but often fail to generalize when training and test nodes come from different distributions, limiting their practicality. To address this challenge, recent approaches have adopted invariant learning and sample reweighting tec…

Cited by 0SourcePDFScholar
2025

Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study

ICLR 2025poster

Large language models (LLMs) like GPT-4 and LLaMA-3 utilize the powerful in-context learning (ICL) capability of Transformer architecture to learn on the fly from limited examples. While ICL underpins many LLM applications, its full potential remains hindered by a limited understanding of its genera…

2024

Bridging Multicalibration and Out-of-distribution Generalization Beyond Covariate Shift

NeurIPS 2024poster

We establish a new model-agnostic optimization framework for out-of-distribution generalization via multicalibration, a criterion that ensures a predictor is calibrated across a family of overlapping groups. Multicalibration is shown to be associated with robustness of statistical inference under co…

Cited by 1SourcePDFScholar
2024

Debiased Collaborative Filtering with Kernel-Based Causal Balancing

ICLR 2024spotlight

Collaborative filtering builds personalized models from the collected user feedback. However, the collected data is observational rather than experimental, leading to various biases in the data, which can significantly affect the learned model. To address this issue, many studies have focused on pro…

2024

Domain-wise Data Acquisition to Improve Performance under Distribution Shift

ICML 2024poster

Despite notable progress in enhancing the capability of machine learning against distribution shifts, training data quality remains a bottleneck for cross-distribution generalization. Recently, from a data-centric perspective, there have been considerable efforts to improve model performance through…

2024

Enhancing Distributional Stability among Sub-populations

AISTATS 2024poster

Enhancing the stability of machine learning algorithms under distributional shifts is at the heart of the Out-of-Distribution (OOD) Generalization problem. Derived from causal learning, recent works of invariant learning pursue strict invariance with multiple training environments. Although intuitiv…

2024

Geometry-Calibrated DRO: Combating Over-Pessimism with Free Energy Implications

ICML 2024poster

Machine learning algorithms minimizing average risk are susceptible to distributional shifts. Distributionally Robust Optimization (DRO) addresses this issue by optimizing the worst-case risk within an uncertainty set. However, DRO suffers from over-pessimism, leading to low-confidence predictions,…

Cited by 2SourcePDFScholar
2024

How to Engage your Readers? Generating Guiding Questions to Promote Active Reading

ACL 2024long

Using questions in written text is an effective strategy to enhance readability. However, what makes an active reading question good, what the linguistic role of these questions is, and what is their impact on human reading remains understudied. We introduce GuidingQ, a dataset of 10K in-text questi…

2024

Rethinking the Evaluation Protocol of Domain Generalization

CVPR 2024poster

Domain generalization aims to solve the challenge of Out-of-Distribution (OOD) generalization by leveraging common knowledge learned from multiple training domains to generalize to unseen test domains. To accurately evaluate the OOD generalization ability it is required that test data information is…

2024

Stability Evaluation through Distributional Perturbation Analysis

ICML 2024poster

The performance of learning models often deteriorates when deployed in out-of-sample environments. To ensure reliable deployment, we propose a stability evaluation criterion based on distributional perturbations. Conceptually, our stability evaluation criterion is defined as the minimal perturbation…

Cited by 0SourcePDFScholar
2023

Competing for Shareable Arms in Multi-Player Multi-Armed Bandits

ICML 2023poster

Competitions for shareable and limited resources have long been studied with strategic agents. In reality, agents often have to learn and maximize the rewards of the resources at the same time. To design an individualized competing policy, we model the competition between agents in a novel multi-pla…

2023

Covariate-Shift Generalization via Random Sample Weighting

AAAI 2023technical

Shifts in the marginal distribution of covariates from training to the test phase, named covariate-shifts, often lead to unstable prediction performance across agnostic testing data, especially under model misspecification. Recent literature on invariant learning attempts to learn an invariant predi…

Cited by 6SourcePDFScholar
2023

Flatness-Aware Minimization for Domain Generalization

ICCV 2023poster

Domain generalization (DG) seeks to learn robust models that generalize well under unknown distribution shifts. As a critical aspect of DG, optimizer selection has not been explored in depth. Currently, most DG methods follow the widely used benchmark, DomainBed, and utilize Adam as the default opti…

Cited by 30PDFScholar
2023

Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves Generalization

CVPR 2023highlight

Recently, flat minima are proven to be effective for improving generalization and sharpness-aware minimization (SAM) achieves state-of-the-art performance. Yet the current definition of flatness discussed in SAM and its follow-ups are limited to the zeroth-order flatness (i.e., the worst-case loss w…

2023

Learning Sample Difficulty from Pre-trained Models for Reliable Prediction

NeurIPS 2023poster

Large-scale pre-trained models have achieved remarkable success in many applications, but how to leverage them to improve the prediction reliability of downstream models is undesirably under-explored. Moreover, modern neural networks have been found to be poorly calibrated and make overconfident pre…

Cited by 16SourcePDFScholar
2023

NICO++: Towards Better Benchmarking for Domain Generalization

CVPR 2023poster

Despite the remarkable performance that modern deep neural networks have achieved on independent and identically distributed (I.I.D.) data, they can crash under distribution shifts. Most current evaluation methods for domain generalization (DG) adopt the leave-one-out strategy as a compromise on the…

2023

On the Need for a Language Describing Distribution Shifts: Illustrations on Tabular Datasets

NeurIPS 2023poster

Different distribution shifts require different algorithmic and operational interventions. Methodological research must be grounded by the specific shifts they address. Although nascent benchmarks provide a promising empirical foundation, they \emph{implicitly} focus on covariate shifts, an…

2023

Propensity Matters: Measuring and Enhancing Balancing for Recommendation

ICML 2023poster

Propensity-based weighting methods have been widely studied and demonstrated competitive performance in debiased recommendations. Nevertheless, there are still many questions to be addressed. How to estimate the propensity more conducive to debiasing performance? Which metric is more reasonable to m…

Cited by 50SourcePDFScholar
2023

Provably Invariant Learning without Domain Information

ICML 2023poster

Typical machine learning applications always assume the data follows independent and identically distributed (IID) assumptions. In contrast, this assumption is frequently violated in real-world circumstances, leading to the Out-of-Distribution (OOD) generalization problem and a major drop in model r…

Cited by 15SourcePDFScholar
2023

Stable Learning via Sparse Variable Independence

AAAI 2023technical

The problem of covariate-shift generalization has attracted intensive research attention. Previous stable learning algorithms employ sample reweighting schemes to decorrelate the covariates when there is no explicit domain information about training data. However, with finite samples, it is difficul…

Cited by 17SourcePDFScholar
2022

A Theoretical Analysis on Independence-driven Importance Weighting for Covariate-shift Generalization

ICML 2022spotlight

Covariate-shift generalization, a typical case in out-of-distribution (OOD) generalization, requires a good performance on the unknown test distribution, which varies from the accessible training distribution in the form of covariate shift. Recently, independence-driven importance weighting algorith…

2022

Counterfactual Prediction for Outcome-Oriented Treatments

ICML 2022spotlight

Large amounts of efforts have been devoted into learning counterfactual treatment outcome under various settings, including binary/continuous/multiple treatments. Most of these literature aims to minimize the estimation error of counterfactual outcome for the whole treatment space. However, in most…

Cited by 9SourcePDFScholar
2022

Model Agnostic Sample Reweighting for Out-of-Distribution Learning

ICML 2022spotlight

Distributionally robust optimization (DRO) and invariant risk minimization (IRM) are two popular methods proposed to improve out-of-distribution (OOD) generalization performance of machine learning models. While effective for small models, it has been observed that these methods can be vulnerable to…

2022

Product Ranking for Revenue Maximization with Multiple Purchases

NeurIPS 2022accept

Product ranking is the core problem for revenue-maximizing online retailers. To design proper product ranking algorithms, various consumer choice models are proposed to characterize the consumers' behaviors when they are provided with a list of products. However, existing works assume that each cons…

2022

ZIN: When and How to Learn Invariance Without Environment Partition?

NeurIPS 2022accept

It is commonplace to encounter heterogeneous data, of which some aspects of the data distribution may vary but the underlying causal mechanisms remain constant. When data are divided into distinct environments according to the heterogeneity, recent invariant learning methods have proposed to learn…

2021

Deep Stable Learning for Out-of-Distribution Generalization

CVPR 2021poster

Approaches based on deep neural networks have achieved striking performance when testing data and training data share similar distribution, but can significantly fail otherwise. Therefore, eliminating the impact of distribution shifts between training and testing data is crucial for building perform…

Cited by 347PDFcodeScholar
2021

Integrated Latent Heterogeneity and Invariance Learning in Kernel Space

NeurIPS 2021poster

The ability to generalize under distributional shifts is essential to reliable machine learning, while models optimized with empirical risk minimization usually fail on non-$i.i.d$ testing data. Recently, invariant learning methods for out-of-distribution (OOD) generalization propose to find causall…

Cited by 9SourcePDFScholar
2021

Reinforcement Learning with a Disentangled Universal Value Function for Item Recommendation

AAAI 2021technical

In recent years, there are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems (RS). In this paper, we summarize three key practical challenges of large-scale RL-based recommender systems: massive state and action spaces, high-variance environ…

2021

Sliding Selector Network with Dynamic Memory for Extractive Summarization of Long Documents

NAACL 2021long

Neural-based summarization models suffer from the length limitation of text encoder. Long documents have to been truncated before they are sent to the model, which results in huge loss of summary-relevant contents. To address this issue, we propose the sliding selector network with dynamic memory fo…

2021

Stable Adversarial Learning under Distributional Shifts

AAAI 2021technical

Machine learning algorithms with empirical risk minimization are vulnerable under distributional shifts due to the greedy adoption of all the correlations found in training data. Recently, there are robust learning methods aiming at this problem by minimizing the worst-case risk over an uncertainty…

Cited by 34SourcePDFScholar
2020

Counterfactual Prediction for Bundle Treatment

NeurIPS 2020poster

Estimating counterfactual outcome of different treatments from observational data is an important problem to assist decision making in a variety of fields. Among the various forms of treatment specification, bundle treatment has been widely adopted in many scenarios, such as recommendation systems…

2019

Learning Disentangled Representations for Recommendation

NeurIPS 2019poster

User behavior data in recommender systems are driven by the complex interactions of many latent factors behind the users’ decision making processes. The factors are highly entangled, and may range from high-level ones that govern user intentions, to low-level ones that characterize a user’s preferen…

Cited by 422SourcePDFScholar