← Search

Yi Ouyang

13 accepted papers

2026

VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning

AAAI 2026technical

Large language model (LLM) agents have emerged as a promising solution for enhancing recommendation systems via user simulation. However, existing studies predominantly resort to prompt-based simulation using frozen LLMs, which frequently results in suboptimal item modeling and user preference learn

Cited by 0SourcePDFScholar
2026

When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation

AAAI 2026technical

Existing video recommendation systems, relying mainly on ID-based embedding mapping and collaborative filtering, often fail to capture in-depth video content semantics. Moreover, most struggle to address biased user behaviors (e.g., accidental clicks, fast skips), leading to inaccurate interest mode

Cited by 0SourcePDFScholar
2024

Cooper: Coordinating Specialized Agents towards a Complex Dialogue Goal

AAAI 2024technical

In recent years, there has been a growing interest in exploring dialogues with more complex goals, such as negotiation, persuasion, and emotional support, which go beyond traditional service-focused dialogue systems. Apart from the requirement for much more sophisticated strategic reasoning and comm…

2024

MedJourney: Benchmark and Evaluation of Large Language Models over Patient Clinical Journey

NeurIPS 2024poster

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding and generation, leading to their widespread adoption across various fields. Among these, the medical field is particularly well-suited for LLM applications, as many medical tasks can be enhanced by LLMs.…

Cited by 1SourcePDFScholar
2024

Pre-trained Online Contrastive Learning for Insurance Fraud Detection

AAAI 2024technical

Medical insurance fraud has always been a crucial challenge in the field of healthcare industry. Existing fraud detection models mostly focus on offline learning scenes. However, fraud patterns are constantly evolving, making it difficult for models trained on past data to detect newly emerging frau…

2024

Safeguarding Fraud Detection from Attacks: A Robust Graph Learning Approach

IJCAI 2024poster

Financial fraud is one of the most significant social issues and has caused tremendous property losses. Graph neural networks (GNNs) have been applied to anti-fraud practices and achieved decent results. However, recent researches have discovered flaws in the robustness of fraud-detection models bas…

Cited by 7SourcePDFScholar
2023

Fighting against Organized Fraudsters Using Risk Diffusion-based Parallel Graph Neural Network

IJCAI 2023poster

Medical insurance plays a vital role in modern society, yet organized healthcare fraud causes billions of dollars in annual losses, severely harming the sustainability of the social welfare system. Existing works mostly focus on detecting individual fraud entities or claims, ignoring hidden conspira…

Cited by 14SourcePDFScholar
2023

Semi-supervised Credit Card Fraud Detection via Attribute-Driven Graph Representation

AAAI 2023technical

Credit card fraud incurs a considerable cost for both cardholders and issuing banks. Contemporary methods apply machine learning-based classifiers to detect fraudulent behavior from labeled transaction records. But labeled data are usually a small proportion of billions of real transactions due to e…

2022

Training a Resilient Q-network against Observational Interference

AAAI 2022technical

Deep reinforcement learning (DRL) has demonstrated impressive performance in various gaming simulators and real-world applications. In practice, however, a DRL agent may receive faulty observation by abrupt interferences such as black-out, frozen-screen, and adversarial perturbation. How to design…

2020

Enhanced Adversarial Strategically-Timed Attacks Against Deep Reinforcement Learning

ICASSP 2020accepted

Recent deep neural networks based techniques, especially those equipped with the ability of self-adaptation in the system level such as deep reinforcement learning (DRL), are shown to possess many advantages of optimizing robot learning systems (e.g., autonomous navigation and continuous robot arm c…

Cited by 0SourceScholar
2020

Regret Bounds for Decentralized Learning in Cooperative Multi-Agent Dynamical Systems

UAI 2020poster

Regret analysis is challenging in Multi-Agent Reinforcement Learning (MARL) primarily due to the dynamical environments and the decentralized information among agents. We attempt to solve this challenge in the context of decentralized learning in multi-agent linear-quadratic (LQ) dynamical systems.…

Cited by 12SourcePDFScholar
2018

A theory on the absence of spurious solutions for nonconvex and nonsmooth optimization

NeurIPS 2018poster

We study the set of continuous functions that admit no spurious local optima (i.e. local minima that are not global minima) which we term global functions. They satisfy various powerful properties for analyzing nonconvex and nonsmooth optimization problems. For instance, they satisfy a theorem akin…

Cited by 54SourcePDFScholar
2017

Learning Unknown Markov Decision Processes: A Thompson Sampling Approach

NeurIPS 2017poster

We consider the problem of learning an unknown Markov Decision Process (MDP) that is weakly communicating in the infinite horizon setting. We propose a Thompson Sampling-based reinforcement learning algorithm with dynamic episodes (TSDE). At the beginning of each episode, the algorithm generates a s…

Cited by 167SourcePDFScholar