← Search

Huijie Guo

4 accepted papers

2026

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction

AAAI 2026technical

External reasoning systems combine language models with process reward models (PRMs) to select high-quality reasoning paths for complex tasks such as mathematical problem solving. However, these systems are prone to reward hacking, where high-scoring but logically incorrect paths are assigned high s

Cited by 0SourcePDFScholar
2026

Exploring Transferability of Self-Supervised Learning by Task Conflict Calibration

AAAI 2026technical

In this paper, we explore the transferability of SSL by addressing two central questions: (i) what is the representation transferability of SSL, and (ii) how can we effectively model this transferability? Transferability is defined as the ability of a representation learned from one task to support

Cited by 0SourcePDFScholar
2024

Self-Supervised Representation Learning with Meta Comprehensive Regularization

AAAI 2024technical

Self-Supervised Learning (SSL) methods harness the concept of semantic invariance by utilizing data augmentation strategies to produce similar representations for different deformations of the same input. Essentially, the model captures the shared information among multiple augmented views of sample…

Cited by 6SourcePDFScholar