← Search

Liangyu Huo

10 accepted papers

2026

Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks

AAAI 2026technical

Large Language Models (LLMs) excel in reasoning tasks requiring a single correct answer, but they perform poorly in multi-solution tasks that require generating comprehensive and diverse answers. We attribute this limitation to reasoning overconfidence: a tendency to express undue certainty in an in

Cited by 0SourcePDFScholar
2026

MRACL: Multi-Reward Space Guided Adaptive Curriculum Reinforcement Learning for LLMs

AAAI 2026technical

Reinforcement learning (RL) has recently become a powerful yet resource-intensive approach for post-training large language models (LLMs). Incorporating curriculum learning (CL) into RL has been shown to significantly improve training efficiency, particularly in reasoning tasks. However, existing CL

Cited by 0SourcePDFScholar
2026

Norm$\times$Direction: Restoring the Missing Query Norm in Vision Linear Attention

ICML 2026poster

Linear attention mitigates the quadratic complexity of softmax attention but suffers from a critical loss of expressiveness. We identify two primary causes: (1) The normalization operation cancels the query norm, which breaks the correlation between a query's norm and the spikiness (entropy) of the …

Cited by 0SourceScholar
2025

Beyond Excess and Deficiency: Adaptive Length Bias Mitigation in Reward Models for RLHF

NAACL 2025findings

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models (LLMs) with human values. However, it has been noted that reward models in RLHF often exhibit unintended biases, such as an overemphasis on response length based on the erroneous assumption that longer re…

Cited by 0SourcePDFScholar
2024

How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers

NeurIPS 2024poster

Pre-trained language models have been proven to possess strong base capabilities, which not only excel in in-distribution language modeling but also show powerful abilities in out-of-distribution language modeling, transfer learning and few-shot learning. Unlike existing work focusing on the influen…

Cited by 0SourcePDFScholar
2024

MoGU: A Framework for Enhancing Safety of LLMs While Preserving Their Usability

NeurIPS 2024poster

Large Language Models (LLMs) are increasingly deployed in various applications. As their usage grows, concerns regarding their safety are rising, especially in maintaining harmless responses when faced with malicious instructions. Many defense strategies have been developed to enhance the safety of…

Cited by 4SourcePDFScholar
2023

Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation Learning

AAAI 2023technical

Imitation learning (IL) has recently shown impressive performance in training a reinforcement learning agent with human demonstrations, eliminating the difficulty of designing elaborate reward functions in complex environments. However, most IL methods work under the assumption of the optimality of…

Cited by 0SourcePDFScholar
2020

Anti-Jamming Routing For Internet of Satellites: a Reinforcement Learning Approach

ICASSP 2020accepted

The anti-jamming routing for the Internet of Satellites (IoS) has drawn increasing attentions due to the unknown interrupts, unexpected congestion and smart jamming. This paper investigates anti-jamming routing scheme for heterogeneous IoS, with the aim of minimizing anti-jamming routing cost. First…

Cited by 0SourceScholar
2020

Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions

ICASSP 2020accepted

Automatic sub-policy discovery has recently received much attention in hierarchical reinforcement learning (HRL). The conventional approaches to learning sub-policies suffer from collapsing into just one sub-policy dominating the whole task, lacking techniques to ensure the diversity of different su…

Cited by 0SourceScholar
2019

Optimizing QoE of Multiple Users over DASH: A Meta-learning Approach

ICASSP 2019accepted

Dynamic adaptive video streaming over HTTP (DASH) plays a key role in video transmission over the Internet. The conventional DASH adaptation approaches concentrate on optimizing the overall quality of experience (QoE) for all client sides, neglecting the QoE diversity of different users. In this pap…

Cited by 0SourceScholar