← Search

Haocheng Luo

5 accepted papers

2026

Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization

ICLR 2026poster

Direct Preference Optimization (DPO) has emerged as a popular algorithm for aligning pretrained large language models with human preferences, owing to its simplicity and training stability. However, DPO suffers from the recently identified squeezing effect (also known as likelihood displacement), wh…

Cited by 0SourcecodeScholar
2025

Spatial-Temporal Graph Contrastive Learning with Decreasing Masks for Traffic Flow Forecasting

IROS 2025

In recent years, Contrastive learning has shown great potential in traffic flow prediction tasks. However, existing contrastive learning methods have difficulties in dealing with missing data and noise, and it is difficult to fully capture local and global correlations by relying on a single contras

Cited by 0SourceScholar
2025

Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise

NeurIPS 2025poster

Sharpness‐aware minimization (SAM) has emerged as a highly effective technique for improving model generalization, but its underlying principles are not fully understood. We investigated the phenomenon known as m-sharpness, where the performance of SAM improves monotonically as the micro-batch size…

Cited by 0SourceScholar
2024

Explicit Eigenvalue Regularization Improves Sharpness-Aware Minimization

NeurIPS 2024poster

Sharpness-Aware Minimization (SAM) has attracted significant attention for its effectiveness in improving generalization across various tasks. However, its underlying principles remain poorly understood. In this work, we analyze SAM’s training dynamics using the maximum eigenvalue of the Hessian as…

2023

Re-weighting Tokens: A Simple and Effective Active Learning Strategy for Named Entity Recognition

EMNLP 2023short findings

Active learning, a widely adopted technique for enhancing machine learning models in text and image classification tasks with limited annotation resources, has received relatively little attention in the domain of Named Entity Recognition (NER). The challenge of data imbalance in NER has hindered th…

Cited by 0SourceScholar