← Search

Qinghai Miao

4 accepted papers

2025

Rank-Awareness and Angular Constraints: A New Perspective on Learning Sentence Embeddings from NLI Data

EMNLP 2025

Learning high-quality sentence embeddings from Natural Language Inference (NLI) data is often challenged by a critical signal conflict between discrete labels and the continuous spectrum of semantic similarity, as well as information loss from discarded neutral sentence pairs during training. To add

2025

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

ICLR 2025poster

A significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks.…

2024

DimA: A Parameter-efficient Fine-tuning Method with Knowledge Transfer Based on Transformer

COLING 2024main

Fine-tuning is a widely used technique for leveraging pre-trained language models (PLMs) in downstream tasks, but it can be computationally expensive and storage-intensive. To address this challenge, researchers have developed parameter-efficient methods that balance performance and resource cost. H…

2024

RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences

ICML 2024spotlight

Preference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal. However, current PbRL methods excessively depend on high-quality feedback from domain experts, which results in a lack of robustness. In this paper, we pre…