← Search

Zishi Zhang

5 accepted papers

2026

Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer

ICLR 2026oral

The probabilistic diffusion model (DM), generating content by inferencing through a recursive chain structure, has emerged as a powerful framework for visual generation. After pre-training on enormous data, the model needs to be properly aligned to meet requirements for downstream applications. How…

Cited by 0SourcecodeScholar
2026

RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-Training

ICLR 2026poster

Reinforcement learning with verifiable reward has recently emerged as a central paradigm for post-training large language models (LLMs); however, prevailing mean-based methods, such as Group Relative Policy Optimization (GRPO), suffer from entropy collapse and limited reasoning gains. We argue that…

Cited by 0SourcecodeScholar
2025

FLOPS: Forward Learning with OPtimal Sampling

ICLR 2025poster

Given the limitations of backpropagation, perturbation-based gradient computation methods have recently gained focus for learning with only forward passes, also referred to as queries. Conventional forward learning consumes enormous queries on each data point for accurate gradient estimation through…

2025

Robust Dwell Time Allocation for Multiple Ballistic Reentry Target Tracking in Phased Array Radar

ICASSP 2025accepted

Phased array radar (PAR) is shown to provide an enhanced performance for target tracking due to its beam agility and ability for time resource allocation. Existing algorithms for PAR resource allocation often consider standard and simplified dynamic models for target motion, which are not suitable f…

Cited by 0SourceScholar