← Search

Maowei Jiang

2 accepted papers

2026

TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization

AAAI 2026technical

Reinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-

Cited by 0SourcePDFScholar
2025

DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series

NeurIPS 2025poster

Medical time-series data play a vital role in disease diagnosis but suffer from limited labeled samples and single-center bias, which hinder model generalization and lead to overfitting. To address these challenges, we propose DAAC (Discrepancy-Aware Adaptive Contrastive learning), a learnable multi…

Cited by 0SourcecodeScholar