← Search

Ji Cheng

10 accepted papers

2026

Beyond Myopic Alignment: Lookahead Optimization for Online Class-Incremental Learning

CVPR 2026

Rehearsal-based methods are the cornerstone of modern online class-incremental learning (OCIL), yet they face a fundamental challenge: the gradient of the current task often conflicts with that of the rehearsal data from the memory buffer, leading to catastrophic forgetting. Recent works have implic

Cited by 0SourceScholar
2026

Offline Multi-Objective Bandits: From Logged Data to Pareto-Optimal Policies

AAAI 2026technical

Offline policy learning from logged data is a critical paradigm for enabling effective decision-making without costly online exploration. However, its application has been largely confined to single-objective problems, a stark contrast to real-world scenarios where decision-making inherently involve

Cited by 0SourcePDFScholar
2025

Multi-objective Linear Reinforcement Learning with Lexicographic Rewards

ICML 2025poster

Reinforcement Learning (RL) with linear transition kernels and reward functions has recently attracted growing attention due to its computational efficiency and theoretical advancements. However, prior theoretical research in RL has primarily focused on single-objective problems, resulting in limite…

Cited by 0SourcePDFScholar
2024

Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits

AAAI 2024technical

Multi-objective Stochastic Linear bandit (MOSLB) plays a critical role in the sequential decision-making paradigm, however, most existing methods focus on the Pareto dominance among different objectives without considering any priority. In this paper, we study bandit algorithms under mixed Pareto-le…

2024

Multiobjective Lipschitz Bandits under Lexicographic Ordering

AAAI 2024technical

This paper studies the multiobjective bandit problem under lexicographic ordering, wherein the learner aims to simultaneously maximize ? objectives hierarchically. The only existing algorithm for this problem considers the multi-armed bandit model, and its regret bound is O((KT)^(2/3)) under a metri…

Cited by 3SourcePDFScholar
2024

Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples

ICML 2024poster

Neural Network-based active learning (NAL) is a cost-effective data selection technique that utilizes neural networks to select and train on a small subset of samples. While existing work successfully develops various effective or theory-justified NAL algorithms, the understanding of the two commonl…

Cited by 3SourcePDFScholar