← Search

Yu Bai

63 accepted papers

2026

CE-VFAL: A Novel Framework for Communication-Efficient Vertical Federated Adversarial Learning

IJCAI 2026

Vertical Federated Learning (VFL) involves multiple participants collaborating to train machine learning models on distinct feature sets from the same data samples. This training paradigm with distributed updating focuses on secure and efficient communication. Nevertheless, the trained models exhibi

Cited by 0Scholar
2026

Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification

CVPR 2026

Recent pseudo-bag augmentation methods for Multiple Instance Learning (MIL)-based Whole Slide Image (WSI) classification sample instances from a limited number of bags, resulting in constrained diversity. To address this issue, we propose Contrastive Cross-Bag Augmentation (C2Aug) to sample instance

Cited by 0SourcecodeScholar
2026

GenJAPNet: A Generalizable Joint Angle Prediction Network with Non-Redundant Muscle Synergy Features for Lower-Limb Exoskeletons

ICRA 2026poster

Lower-limb exoskeleton robots play a significant role in both rehabilitation and assisted walking, where accurate prediction of lower-limb joint angles is crucial for achieving natural gait. However, due to inter-subject variability and differences across locomotion modes, achieving cross-task gener…

Cited by 0Scholar
2026

Identifying and Analyzing Performance-Critical Tokens in Large Language Models

AAAI 2026technical

In-context learning (ICL) has emerged as an effective solution for few-shot learning with large language models (LLMs). However, how LLMs leverage demonstrations to specify a task and learn a corresponding computational function through ICL is underexplored. Drawing from the way humans learn from c

Cited by 0SourcePDFScholar
2026

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization

ICLR 2026poster

Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-world long-context information, primarily due to insufficient long-context alignment caused by data quality issues, training inefficiencies, and the lack of w…

Cited by 0SourcecodeScholar
2026

Three Forward, One Backward: Memory-Efficient Full-Rank Fine-Tuning of Large Models via Extra Forward Passes

ICLR 2026poster

Fine-tuning large language models (LLMs) has achieved significant success in downstream tasks. However, as the model size continues to grow, traditional fine-tuning methods have become increasingly impractical due to their high computational and memory costs. This has motivated researchers to explor…

Cited by 0SourcecodeScholar
2025

Accelerated Vertical Federated Adversarial Learning through Decoupling Layer-Wise Dependencies

NeurIPS 2025poster

Vertical Federated Learning (VFL) enables participants to collaboratively train models on aligned samples while keeping their heterogeneous features private and distributed. Despite their utility, VFL models remain vulnerable to adversarial attacks during inference. Adversarial Training (AT), which…

Cited by 0SourceScholar
2025

Excluding the Impossible for Open Vocabulary Semantic Segmentation

AAAI 2025technical

Open vocabulary semantic segmentation is a hot topic in research, focusing on segmenting and recognizing a diverse array of categories in varied environments, including those previously unknown, thereby holding significant practical value. Mainstream studies utilize the CLIP model for direct semanti…

2025

Text2Data: Low-Resource Data Generation with Textual Control

AAAI 2025technical

Natural language serves as a common and straightforward control signal for humans to interact seamlessly with machines. Recognizing the importance of this interface, the machine learning community is investing considerable effort in generating data that is semantically coherent with textual instruct…

2025

TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model

CVPR 2025poster

Vision-Language Models (VLMs) demand substantial computational resources during inference, largely due to the extensive visual input tokens for representing visual information. Previous studies have noted that visual tokens tend to receive less attention than text tokens, suggesting their lower impo…

Cited by 3SourcePDFScholar
2025

Transcending Cost-Quality Tradeoff in Agent Serving via Session-Awareness

NeurIPS 2025poster

Large Language Model (LLM) agents are capable of task execution across various domains by autonomously interacting with environments and refining LLM responses based on feedback. However, existing model serving systems are not optimized for the unique demands of serving agents. Compared to classic m…

Cited by 0SourceScholar
2024

CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling

EMNLP 2024main

Long sequence modeling has gained broad interest as large language models (LLMs) continue to advance. Recent research has identified that a large portion of hidden states within the key-value caches of Transformer models can be discarded (also termed evicted) withoutaffecting the perplexity performa…

2024

Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning

AAAI 2024technical

Open-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the adv…

2024

DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot Learning

CVPR 2024poster

Open-World Few-Shot Learning (OFSL) is a critical field of research concentrating on the precise identification of target samples in environments with scarce data and unreliable labels thus possessing substantial practical significance. Recently the evolution of foundation models like CLIP has revea…

2024

Fundamental Capabilities of Large Language Models and their Applications in Domain Scenarios: A Survey

ACL 2024long

Large Language Models (LLMs) demonstrate significant value in domain-specific applications, benefiting from their fundamental capabilities. Nevertheless, it is still unclear which fundamental capabilities contribute to success in specific domains. Moreover, the existing benchmark-based evaluation ca…

Cited by 4SourcePDFScholar
2024

How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with Representations

ICLR 2024poster

While large language models based on the transformer architecture have demonstrated remarkable in-context learning (ICL) capabilities, understandings of such capabilities are still in an early stage, where existing theory and mechanistic understanding focus mostly on simple scenarios such as learnin…

Cited by 61SourcePDFScholar
2024

How Far Can In-Context Alignment Go? Exploring the State of In-Context Alignment

EMNLP 2024finding

Recent studies have demonstrated that In-Context Learning (ICL), through the use of specific demonstrations, can align Large Language Models (LLMs) with human preferences known as In-Context Alignment (ICA), indicating that models can comprehend human instructions without requiring parameter adjustm…

2024

Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical Perspective

ICML 2024poster

Inverse Reinforcement Learning (IRL)---the problem of learning reward functions from demonstrations of an *expert policy*---plays a critical role in developing intelligent systems. While widely used in applications, theoretical understandings of IRL present unique challenges and remain less develope…

Cited by 6SourcePDFScholar
2024

Norma: A Noise Robust Memory-Augmented Framework for Whole Slide Image Classification

ECCV 2024poster

"In recent years, the Whole Slide Image (WSI) classification task has achieved great advancement due to the success of Multiple Instance Learning (MIL). However, the MIL-based studies usually consider instances within each bag as unordered, potentially resulting in the missing of local and global co…

2024

Sample-Efficient Learning of POMDPs with Multiple Observations In Hindsight

ICLR 2024poster

This paper studies the sample-efficiency of learning in Partially Observable Markov Decision Processes (POMDPs), a challenging problem in reinforcement learning that is known to be exponentially hard in the worst-case. Motivated by real-world settings such as loading in game playing, we propose an e…

Cited by 8SourcePDFScholar
2024

Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining

ICLR 2024poster

Large transformer models pretrained on offline reinforcement learning datasets have demonstrated remarkable in-context reinforcement learning (ICRL) capabilities, where they can make good decisions when prompted with interaction trajectories from unseen environments. However, when and how transforme…

2023

Efficient RL with Impaired Observability: Learning to Act with Delayed and Missing State Observations

NeurIPS 2023poster

In real-world reinforcement learning (RL) systems, various forms of {\it impaired observability} can complicate matters. These situations arise when an agent is unable to observe the most recent state of the system due to latency or lossy channels, yet the agent must still make real-time decisions.…

Cited by 9SourcePDFScholar
2023

Fine-Grained Blind Face Inpainting with 3D Face Component Disentanglement

ICASSP 2023accepted

Inpainting is a task to restore occlusion or other corruption on images. However, previous works require mask of the occluded area to restore the occluded image, which is inconvenient for application. Blind face inpainting aims to automatically restore the occluded face without position information…

Cited by 0SourceScholar
2023

Improved Online Conformal Prediction via Strongly Adaptive Online Learning

ICML 2023poster

We study the problem of uncertainty quantification via prediction sets, in an online setting where the data distribution may vary arbitrarily over time. Recent work develops *online conformal prediction* techniques that leverage regret minimization algorithms from the online learning literature to l…

2023

Partially Observable RL with B-Stability: Unified Structural Condition and Sharp Sample-Efficient Algorithms

ICLR 2023top-25%

Partial Observability---where agents can only observe partial information about the true underlying state of the system---is ubiquitous in real-world applications of Reinforcement Learning (RL). Theoretically, learning a near-optimal policy under partial observability is known to be hard in the wors…

Cited by 31SourcePDFScholar
2023

The Role of Coverage in Online Reinforcement Learning

ICLR 2023top-5%

Coverage conditions---which assert that the data logging distribution adequately covers the state space---play a fundamental role in determining the sample complexity of offline reinforcement learning. While such conditions might seem irrelevant to online reinforcement learning at first glance, we e…

Cited by 89SourcePDFScholar
2023

Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection

NeurIPS 2023oral

Neural sequence models based on the transformer architecture have demonstrated remarkable \emph{in-context learning} (ICL) abilities, where they can perform new tasks when prompted with training and test examples, without any parameter update to the model. This work first provides a comprehensive st…

2023

What can a Single Attention Layer Learn? A Study Through the Random Features Lens

NeurIPS 2023poster

Attention layers---which map a sequence of inputs to a sequence of outputs---are core building blocks of the Transformer architecture which has achieved significant breakthroughs in modern artificial intelligence. This paper presents a rigorous theoretical study on the learning and generalization of…

Cited by 34SourcePDFScholar
2022

Conformal Predictor for Improving Zero-Shot Text Classification Efficiency

EMNLP 2022main

Pre-trained language models (PLMs) have been shown effective for zero-shot (0shot) text classification. 0shot models based on natural language inference (NLI) and next sentence prediction (NSP) employ cross-encoder architecture and infer by making a forward pass through the model for each label-text…

Cited by 3SourcePDFScholar
2022

Efficient Phi-Regret Minimization in Extensive-Form Games via Online Mirror Descent

NeurIPS 2022accept

A conceptually appealing approach for learning Extensive-Form Games (EFGs) is to convert them to Normal-Form Games (NFGs). This approach enables us to directly translate state-of-the-art techniques and analyses in NFGs to learning EFGs, but typically suffers from computational intractability due to…

Cited by 25SourcePDFScholar
2022

Efficient and Differentiable Conformal Prediction with General Function Classes

ICLR 2022poster

Quantifying the data uncertainty in learning tasks is often done by learning a prediction interval or prediction set of the label given the input. Two commonly desired properties for learned prediction sets are \emph{valid coverage} and \emph{good efficiency} (such as low length or low cardinality).…

2022

Identifying good directions to escape the NTK regime and efficiently learn low-degree plus sparse polynomials

NeurIPS 2022accept

A recent goal in the theory of deep learning is to identify how neural networks can escape the “lazy training,” or Neural Tangent Kernel (NTK) regime, where the network is coupled with its first order Taylor expansion at initialization. While the NTK is minimax optimal for learning dense polynomials…

2022

Local calibration: metrics and recalibration

UAI 2022poster

Probabilistic classifiers output confidence scores along with their predictions, and these confidence scores should be calibrated, i.e., they should reflect the reliability of the prediction. Confidence scores that minimize standard metrics such as the expected calibration error (ECE) accurately mea…

Cited by 23SourcePDFScholar
2022

Near-Optimal Learning of Extensive-Form Games with Imperfect Information

ICML 2022spotlight

This paper resolves the open question of designing near-optimal algorithms for learning imperfect-information extensive-form games from bandit feedback. We present the first line of algorithms that require only $\widetilde{\mathcal{O}}((XA+YB)/\varepsilon^2)$ episodes of play to find an $\varepsilon…

Cited by 36SourcePDFScholar
2022

PSP: Pre-trained Soft Prompts for Few-Shot Abstractive Summarization

COLING 2022main

Few-shot abstractive summarization has become a challenging task in natural language generation. To support it, we developed a novel soft prompts architecture coupled with a prompt pre-training plus prompt fine-tuning paradigm, which is effective and tunes only extremely light parameters. To meet th…

Cited by 26SourcePDFScholar
2022

Policy Optimization for Markov Games: Unified Framework and Faster Convergence

NeurIPS 2022accept

This paper studies policy optimization algorithms for multi-agent reinforcement learning. We begin by proposing an algorithm framework for two-player zero-sum Markov Games in the full-information setting, where each iteration consists of a policy update step at each state using a certain matrix game…

Cited by 33SourcePDFScholar
2022

Stage-wise Stylistic Headline Generation: Style Generation and Summarized Content Insertion

IJCAI 2022poster

A quality headline with a high click-rate should not only summarize the content of an article, but also reflect a style that attracts users. Such demand has drawn rising attention to the task of stylistic headline generation (SHG). An intuitive method is to first generate plain headlines leveraged b…

2022

When Can We Learn General-Sum Markov Games with a Large Number of Players Sample-Efficiently?

ICLR 2022poster

Multi-agent reinforcement learning has made substantial empirical progresses in solving games with a large number of players. However, theoretically, the best known sample complexity for finding a Nash equilibrium in general-sum games scales exponentially in the number of players due to the size of…

Cited by 124SourcePDFScholar
2021

A Sharp Analysis of Model-based Reinforcement Learning with Self-Play

ICML 2021spotlight

Model-based algorithms—algorithms that explore the environment through building and utilizing an estimated model—are widely used in reinforcement learning practice and theoretically shown to achieve optimal sample efficiency for single-agent reinforcement learning in Markov Decision Processes (MDPs)…

Cited by 169SourcePDFScholar
2021

Don’t Just Blame Over-parametrization for Over-confidence: Theoretical Analysis of Calibration in Binary Classification

ICML 2021spotlight

Modern machine learning models with high accuracy are often miscalibrated—the predicted top probability does not reflect the actual accuracy, and tends to be \emph{over-confident}. It is commonly believed that such over-confidence is mainly due to \emph{over-parametrization}, in particular when the…

Cited by 64SourcePDFScholar
2021

Exact Gap between Generalization Error and Uniform Convergence in Random Feature Models

ICML 2021spotlight

Recent work showed that there could be a large gap between the classical uniform convergence bound and the actual test error of zero-training-error predictors (interpolators) such as deep neural networks. To better understand this gap, we study the uniform convergence in the nonlinear random feature…

Cited by 27SourcePDFScholar
2021

Exploring Explainable Selection to Control Abstractive Summarization

AAAI 2021technical

Like humans, document summarization models can interpret a document’s contents in a number of ways. Unfortunately, the neural models of today are largely black boxes that provide little explanation of how or why they generated a summary in the way they did. Therefore, to begin prying open the black…

2021

How Important is the Train-Validation Split in Meta-Learning?

ICML 2021spotlight

Meta-learning aims to perform fast adaptation on a new task through learning a “prior” from multiple existing tasks. A common practice in meta-learning is to perform a train-validation split (\emph{train-val method}) where the prior adapts to the task on one split of the data, and the resulting pred…

Cited by 92SourcePDFScholar
2021

Near-Optimal Provable Uniform Convergence in Offline Policy Evaluation for Reinforcement Learning

AISTATS 2021poster

The problem of \emph{Offline Policy Evaluation} (OPE) in Reinforcement Learning (RL) is a critical step towards applying RL in real life applications. Existing work on OPE mostly focus on evaluating a \emph{fixed} target policy $\pi$, which does not provide useful bounds for offline policy learning…

Cited by 84SourcePDFScholar
2021

Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement Learning

NeurIPS 2021poster

Recent theoretical work studies sample-efficient reinforcement learning (RL) extensively in two settings: learning interactively in the environment (online RL), or learning from an offline dataset (offline RL). However, existing algorithms and theories for learning near-optimal policies in these two…

Cited by 196SourcePDFScholar
2021

Sample-Efficient Learning of Stackelberg Equilibria in General-Sum Games

NeurIPS 2021poster

Real world applications such as economics and policy making often involve solving multi-agent games with two unique features: (1) The agents are inherently *asymmetric* and partitioned into leaders and followers; (2) The agents have different reward functions, thus the game is *general-sum*. The maj…

Cited by 83SourcePDFScholar
2021

Understanding the Under-Coverage Bias in Uncertainty Estimation

NeurIPS 2021spotlight

Estimating the data uncertainty in regression tasks is often done by learning a quantile function or a prediction interval of the true label conditioned on the input. It is frequently observed that quantile regression---a vanilla algorithm for learning quantiles with asymptotic guarantees---tends to…

Cited by 18SourcePDFScholar
2020

Towards Understanding Hierarchical Learning: Benefits of Neural Representations

NeurIPS 2020poster

Deep neural networks can empirically perform efficient hierarchical learning, in which the layers learn useful representations of the data. However, how they make use of the intermediate representations are not explained by recent theories that relate them to ``shallow learners'' such as kernels. In…

Cited by 63SourcePDFScholar
2019

Compressing Deep Neural Networks Using Toeplitz Matrix: Algorithm Design and Fpga Implementation

ICASSP 2019accepted

Deep neural networks (DNNs) have emerged as an important artificial intelligence technique. However, the computation-intensive and storage-intensive DNNs pose severe challenges on efficient execution over the underlying hardware platform. In this paper we propose to impose Toeplitz structure on DNN…

Cited by 0SourceScholar