← Search

Yifei Ma

11 accepted papers

2026

DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs

AAAI 2026technical

Large Language Models (LLMs) have shown impressive capabilities in multi-step reasoning and problem-solving. Recent works introduce multi-agent reflection frameworks where multiple LLM agents critique and refine each other’s outputs using reinforcement learning (RL). However, these approaches often

Cited by 0SourcePDFScholar
2026

Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI

CVPR 2026

Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorealistic sensor rendering with geometrically grounded interaction in complex, open-world urban environments. Although rece

Cited by 0SourcecodeScholar
2024

Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms

NeurIPS 2024poster

Modern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results in certain aspects, e.g., visual aesthetic, preferred style, and responsibility. In this paper, we target the realm of vi…

Cited by 3SourcePDFScholar
2024

Optimal Design for Human Preference Elicitation

NeurIPS 2024poster

Learning of preference models from human feedback has been central to recent advances in artificial intelligence. Motivated by the cost of obtaining high-quality human annotations, we study efficient human preference elicitation for learning preference models. The key idea in our work is to generali…

Cited by 5SourcePDFScholar
2023

Fixed-Budget Best-Arm Identification with Heterogeneous Reward Variances

UAI 2023poster

We study the problem of best-arm identification (BAI) in the fixed-budget setting with heterogeneous reward variances. We propose two variance-adaptive BAI algorithms for this setting: SHVar for known reward variances and SHAdaVar for unknown reward variances. Our algorithms rely on non-uniform budg…

Cited by 9SourcePDFScholar
2022

Context Uncertainty in Contextual Bandits with Applications to Recommender Systems

AAAI 2022technical

Recurrent neural networks have proven effective in modeling sequential user feedbacks for recommender systems. However, they usually focus solely on item relevance and fail to effectively explore diverse items for users, therefore harming the system performance in the long run. To address this probl…

Cited by 8SourcePDFScholar
2019

Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling

NeurIPS 2019poster

Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) --- the problem of evaluating a new policy using the historical data obtained by different behavior policies --- under the model o…

Cited by 206SourcePDFScholar
2018

Trajectory-Optimized Sensing for Active Search of Tissue Abnormalities in Robotic Surgery

ICRA 2018poster

In this work, we develop an approach for guiding robots to automatically localize and find the shapes of tumors and other stiff inclusions present in the anatomy. Our approach uses Gaussian processes to model the stiffness distribution and active learning to direct the palpation path of the robot. T…

Cited by 27SourceScholar