← Search

Yuanyuan Shi

14 accepted papers

2026

Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMs

AAAI 2026technical

Medical Vision-Language Models (MedVLMs) show immense promise in clinical applicability. However, their reliability is hindered by hallucinations, where models often fail to derive answers from visual evidence, instead relying on learned textual priors. Existing mitigation strategies for MedVLMs hav

Cited by 0SourcePDFScholar
2026

DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients

AAAI 2026technical

Real-world control systems require policies that are not only high-performing but also interpretable and robust. A promising direction toward this goal is model-based control, which learns system dynamics and cost functions from historical data and then uses these models to inform decision-making. B

Cited by 0SourcePDFScholar
2026

LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts

ICLR 2026poster

Recent studies have shown that combining parameter-efficient fine-tuning (PEFT) with mixture-of-experts (MoE) is an effective strategy for adapting large language models (LLMs) to the downstream tasks. However, most existing approaches rely on conventional TopK routing, which requires careful hyperp…

Cited by 0SourcecodeScholar
2026

RN-D: Discretized Categorical Actors with Regularized Networks for On-Policy Reinforcement Learning

ICML 2026poster

On-policy deep reinforcement learning remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy and policy updates must be conservative. In this paper, w…

Cited by 0SourceScholar
2025

Analytical Lyapunov Function Discovery: An RL-based Generative Approach

ICML 2025poster

Despite advances in learning-based methods, finding valid Lyapunov functions for nonlinear dynamical systems remains challenging. Current neural network approaches face two main issues: challenges in scalable verification and limited interpretability. To address these, we propose an end-to-end fram…

2025

Uncovering the Bigger Picture: Comprehensive Event Understanding Via Diverse News Retrieval

EMNLP 2025

Access to diverse perspectives is essential for understanding real-world events, yet most news retrieval systems prioritize textual relevance, leading to redundant results and limited viewpoint exposure. We propose NEWSCOPE, a two-stage framework for diverse news retrieval that enhances event covera

2024

Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation

NeurIPS 2024poster

Training a policy in a source domain for deployment in the target domain under a dynamics shift can be challenging, often resulting in performance degradation. Previous work tackles this challenge by training on the source domain with modified rewards derived by matching distributions between the so…

2023

Structured Neural-PI Control with End-to-End Stability and Output Tracking Guarantees

NeurIPS 2023poster

We study the optimal control of multiple-input and multiple-output dynamical systems via the design of neural network-based controllers with stability and output tracking guarantees. While neural network-based nonlinear controllers have shown superior performance in various applications, their lack…

2023

SustainGym: Reinforcement Learning Environments for Sustainable Energy Systems

NeurIPS 2023poster

The lack of standardized benchmarks for reinforcement learning (RL) in sustainability applications has made it difficult to both track progress on specific domains and identify bottlenecks for researchers to focus their efforts. In this paper, we present SustainGym, a suite of five environments desi…

2022

Stable and Efficient Shapley Value-Based Reward Reallocation for Multi-Agent Reinforcement Learning of Autonomous Vehicles

ICRA 2022poster

With the development of sensing and communication technologies in networked cyber-physical systems (CPSs), multi-agent reinforcement learning (MARL)-based methodologies are integrated into the control process of physical systems and demonstrate prominent performance in a wide array of CPS domains, s…

Cited by 33SourceScholar
2021

Training Certifiably Robust Neural Networks with Efficient Local Lipschitz Bounds

NeurIPS 2021poster

Certified robustness is a desirable property for deep neural networks in safety-critical applications, and popular training algorithms can certify robustness of a neural network by computing a global bound on its Lipschitz constant. However, such a bound is often loose: it tends to over-regularize t…

2020

Robust Reinforcement Learning for Continuous Control with Model Misspecification

ICLR 2020poster

We provide a framework for incorporating robustness -- to perturbations in the transition dynamics which we refer to as model misspecification -- into continuous control Reinforcement Learning (RL) algorithms. We specifically focus on incorporating robustness into a state-of-the-art continuous contr…

Cited by 135SourceScholar