← Search

Ruijia Zhang

5 accepted papers

2026

Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in Reinforcement Learning

ICML 2026poster

Low-Rank Adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tuning (SFT) paradigm. However, their efficacy and behavior under Reinforcement Learning with Verifiable Rewards (RLVR) are less well understood. In particular, two s…

Cited by 0SourceScholar
2026

SafeSieve: From Heuristics to Experience in Progressive Pruning for LLM-based Multi-Agent Communication

AAAI 2026technical

LLM-based multi-agent systems exhibit strong collaborative capabilities but often suffer from redundant communication and excessive token overhead. Existing methods typically enhance efficiency through pretrained GNNs or greedy algorithms, but often isolate pre- and post-task optimization, lacking a

Cited by 0SourcePDFScholar
2026

Topology-Enhanced and Label Correlation-Aware Model for Protein-Protein Interaction Prediction

AAAI 2026technical

Protein-Protein Interactions (PPIs) prediction is crucial for understanding cellular functions and disease mechanisms. Existing deep learning–based methods primarily rely on direct interaction within the PPI network to update protein representations. However, (1) such networks overlook the potential

Cited by 0SourcePDFScholar
2025

Improved Rates of Differentially Private Nonconvex-Strongly-Concave Minimax Optimization

AAAI 2025technical

In this paper, we study the problem of (finite sum) minimax optimization in the Differential Privacy (DP) model. Unlike most of the previous studies on the (strongly) convex-concave settings or loss functions satisfying the Polyak-Lojasiewicz condition, here we mainly focus on the nonconvex-strongly…

Cited by 0SourcePDFScholar
2025

Understanding Inverse Reinforcement Learning under Overparameterization: Non-Asymptotic Analysis and Global Optimality

AISTATS 2025poster

The goal of the Inverse reinforcement learning (IRL) task is to identify the underlying reward function and the corresponding optimal policy from a set of expert demonstrations. While most IRL algorithms’ theoretical guarantees rely on a linear reward structure, we aim to extend the theoretical unde…

Cited by 0SourceScholar