← Search

Wei Ni

5 accepted papers

2026

BOLT: Decision‑Aligned Distillation and Budget-Aware Routing for Constrained Multimodal QA on Robots

ICLR 2026poster

Robotic systems can require multimodal reasoning under stringent constraints of latency, memory, and energy. Standard instruction tuning and token-level distillation fail to deliver decision quality, reliability, and interpretability under these constraints. We introduce BOLT, a decision-aligned dis…

Cited by 0SourceScholar
2026

SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning

ICML 2026poster

Actor-critic (AC) methods are a cornerstone of reinforcement learning (RL) but offer limited interpretability. Current explainable RL methods seldom use *state attributions* to assist training. Rather, they treat all state features equally, thereby neglecting the heterogeneous impacts of individual …

Cited by 0SourceScholar
2025

RobustLight: Improving Robustness via Diffusion Reinforcement Learning for Traffic Signal Control

ICML 2025poster

Reinforcement Learning (RL) optimizes Traffic Signal Control (TSC) to reduce congestion and emissions, but real-world TSC systems face challenges like adversarial attacks and missing data, leading to incorrect signal decisions and increased congestion. Existing methods, limited to offline data predi…

Cited by 0SourcePDFScholar
2025

Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning

NeurIPS 2025poster

Multi-agent reinforcement learning (MARL), as a thriving field, explores how multiple agents independently make decisions in a shared dynamic environment. Due to environmental uncertainties, policies in MARL must remain robust to tackle the sim-to-real gap. We focus on robust two-player zero-sum Mar…

Cited by 0SourceScholar
2025

pFedRAG: A Personalized Federated Retrieval-Augmented Generation System with Depth-Adaptive Tiered Embedding Tuning

EMNLP 2025

Large Language Models (LLMs) can undergo hallucinations in specialized domains, and standard Retrieval-Augmented Generation (RAG) often falters due to general-purpose embeddings ill-suited for domain-specific terminology. Though domain-specific fine-tuning enhances retrieval, centralizing data intro

Cited by 0SourcePDFScholar