← Search

Wentao Shi

12 accepted papers

2025

Exploring Stiffness Gradient Effects in Magnetically Induced Metamorphic Materials via Continuum Simulation and Validation

IROS 2025

Magnetic soft continuum robots are capable of bending with remote control in confined space environments, and they have been applied in various bioengineering contexts. As one type of ferromagnetic soft continuums, the Magnetically Induced Metamorphic Materials (MIMMs)-based continuum (MC) exhibits

Cited by 0SourceScholar
2025

Fine-grained List-wise Alignment for Generative Medication Recommendation

NeurIPS 2025spotlight

Accurate and safe medication recommendations are critical for effective clinical decision-making, especially in multimorbidity cases. However, existing systems rely on point-wise prediction paradigms that overlook synergistic drug effects and potential adverse drug-drug interactions (DDIs). We prop…

Cited by 0SourcecodeScholar
2025

K-order Ranking Preference Optimization for Large Language Models

ACL 2025finding

To adapt large language models (LLMs) to ranking tasks, existing list-wise methods, represented by list-wise Direct Preference Optimization (DPO), focus on optimizing partial-order or full-order list ranking consistency for LLMs to enhance their ranking abilities.However, we argue that optimizing to…

2025

Leveraging Unpaired Feedback for Long-Term LLM-based Recommendation Tuning

EMNLP 2025

Most recommender systems focus on short-term objectives such as click-through rate, often at the expense of long-term user satisfaction. This can lead to echo chambers, where users are repeatedly exposed to redundant content. While recent efforts integrate Large Language Models (LLMs) into recommend

2025

Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment

ACL 2025finding

Multi-Objective Alignment (MOA) aims to align LLMs’ responses with multiple human preference objectives, with Direct Preference Optimization (DPO) emerging as a prominent approach. However, we find that DPO-based MOA approaches suffer from widespread preference conflicts in the data, where different…

2024

Direct Multi-Turn Preference Optimization for Language Agents

EMNLP 2024main

Adapting Large Language Models (LLMs) for agent tasks is critical in developing language agents. Direct Preference Optimization (DPO) is a promising technique for this adaptation with the alleviation of compounding errors, offering a means to directly optimize Reinforcement Learning (RL) objectives.…

2024

SCRN: A Spectrogram Convolutional Recurrent Network for AoA Estimation Using Bluetooth 5

ICASSP 2024accepted

Bluetooth 5 employs Constant Tone Extension (CTE) signals to simulate arrival times of a common signal at multiple antennas, but incurs limited accuracy due to inevitable frequency offsets caused by Gaussian frequency shift keying (GFSK) and time offsets caused by antenna switching. To address this…

Cited by 0SourceScholar
2023

Discriminative-Invariant Representation Learning for Unbiased Recommendation

IJCAI 2023poster

Selection bias hinders recommendation models from learning unbiased user preference. Recent works empirically reveal that pursuing invariant user and item representation across biased and unbiased data is crucial for counteracting selection bias. However, our theoretical analysis reveals that simply…

2023

Understanding Contrastive Learning via Distributionally Robust Optimization

NeurIPS 2023poster

This study reveals the inherent tolerance of contrastive learning (CL) towards sampling bias, wherein negative samples may encompass similar semantics (\eg labels). However, existing theories fall short in providing explanations for this phenomenon. We bridge this research gap by analyzing CL throug…

2022

Transient Analysis of Clustered Multitask Diffusion RLS Algorithm

ICASSP 2022accepted

In this paper, we propose a novel clustered multitask diffusion RLS (MT-DRLS) algorithm over network to further improve the performance of its counterpart, the multitask diffusion LMS (MT-DLMS) algorithm. Its transient behavior is investigated, in the mean and mean-square error sense. Simulation res…

Cited by 0SourceScholar