← Search

Zhiyuan He

11 accepted papers

2026

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

ICLR 2026poster

Reinforcement learning (RL) has become a cornerstone for enhancing the reasoning capabilities of large language models (LLMs), with recent innovations such as Group Relative Policy Optimization (GRPO) demonstrating exceptional effectiveness. In this study, we identify a critical yet underexplored is…

Cited by 0SourcecodeScholar
2026

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

ICML 2026poster

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high computational costs and limit scalability. We propose World-R1, a fra…

Cited by 0SourceScholar
2025

A Synchronous-Optimized and Safety-Improved Framework for Human-Robot Interaction in Robot-Assisted Knee Arthroplasty

RA-L 2025

Shared control is the most commonly used physical human-robot interaction (pHRI) in robot-assisted orthopedic surgery. However, challenges persist in tasks such as robotassisted knee arthroplasty, particularly in terms of motion delays and inadequate three-dimensional (3D) constraints. In this lette

Cited by 0SourceScholar
2025

A sEMG-Based Active-Passive Fusion Rehabilitation Method for Ankle Fracture Rehabilitation Robot after Surgery

RA-L 2025

In this letter, an active-passive fusion rehabilitation training method for the ankle fracture rehabilitation robot after surgery is developed. A subject-independent continuous estimation model of ankle torque is proposed based on surface electromyography (sEMG), an online adaptive algorithm for pas

Cited by 7SourceScholar
2025

LeanK: Learnable K Cache Channel Pruning for Efficient Decoding

EMNLP 2025

Large language models (LLMs) enable long-context tasks but face efficiency challenges due to the growing key-value (KV) cache. We propose LeanK, a learning-based method that prunes unimportant key (K) cache channels by leveraging static channel sparsity. LeanK reduces GPU memory and accelerates deco

Cited by 0SourcePDFScholar
2024

Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised Learning

ICML 2024poster

Deep Neural Networks (DNNs) are vulnerable to Adversarial Examples (AEs), hindering their use in safety-critical systems. In this paper, we present **BEYOND**, an innovative AE detection framework designed for reliable predictions. BEYOND identifies AEs by distinguishing the AE’s abnormal relation w…

Cited by 7SourcePDFScholar
2024

Position Engineering: Boosting Large Language Models through Positional Information Manipulation

EMNLP 2024main

The performance of large language models (LLMs) is significantly influenced by the quality of the prompts provided. In response, researchers have developed enormous prompt engineering strategies aimed at modifying the prompt text to enhance task performance. In this paper, we introduce a novel techn…

Cited by 4SourcePDFScholar
2018

Unsupervised Discovery of Object Landmarks as Structural Representations

CVPR 2018poster

Deep neural networks can model images with rich latent representations, but they cannot naturally conceptualize structures of object categories in a human-perceptible way. This paper addresses the problem of learning object structures in an image modeling process without supervision. We propose an a…

Cited by 232SourcePDFScholar
2017

Discriminative Bimodal Networks for Visual Localization and Detection With Natural Language Queries

CVPR 2017spotlight

Associating image regions with text queries has been recently explored as a new way to bridge visual and linguistic representations. A few pioneering approaches have been proposed based on recurrent neural language models trained generatively (e.g., generating captions), but achieving somewhat limit…

Cited by 61PDFScholar