← Search

Yue Yao

14 accepted papers

2026

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

CVPR 2026

Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks are often inconsistent and show unsatisfying performance against the rapidly evolving MLLMs, failing to exploit non-con

Cited by 0SourcecodeScholar
2026

Bipartite Mode Matching for Vision Training Set Search from a Hierarchical Data Server

AAAI 2026technical

We explore a situation in which the target domain is accessible, but real-time data annotation is not feasible. Instead, we would like to construct an alternative training set from a large-scale data server so that a competitive model can be obtained. For this problem, because the target domain usua

Cited by 0SourcePDFScholar
2026

EP-Diffuser: An Efficient Diffusion Model for Traffic Scene Generation and Prediction Via Polynomial Representations

ICRA 2026poster

As the prediction horizon increases, predicting the future evolution of traffic scenes becomes increasingly difficult due to the multi-modal nature of agent motion. Most state-of-the-art (SotA) prediction models primarily focus on forecasting the most likely future. However, for the safe operation o…

2026

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

AAAI 2026technical

Smartphones bring significant convenience to users but also enable devices to extensively record various types of personal information. Existing smartphone agents powered by Multimodal Large Language Models (MLLMs) have achieved remarkable performance in automating different tasks. However, as the c

Cited by 0SourcePDFScholar
2025

AD4CD: Causal-Guided Anomaly Detection for Enhancing Cognitive Diagnosis

AAAI 2025technical

Cognitive diagnosis is a key task in computer-aided education, aimed at assessing a students' proficiency in specific knowledge concepts based on their responses to exercises. However, existing cognitive diagnosis models often overlook anomalies in students and exercises. For instance, some students…

2025

EP-Diffuser: An Efficient Diffusion Model for Traffic Scene Generation and Prediction via Polynomial Representations

RA-L 2025

As the prediction horizon increases, predicting the future evolution of traffic scenes becomes increasingly difficult due to the multi-modal nature of agent motion. Most state-of-the-art (SotA) prediction models primarily focus on forecasting the most likely future. However, for the safe operation o

Cited by 1SourcecodeScholar
2025

Unsupervised Search for Ethnic Minorities' Medical Segmentation Training Set

ICASSP 2025accepted

This paper investigates the critical issue of dataset bias in medical imaging, with a particular emphasis on racial disparities caused by uneven population distribution in dataset collection. Our analysis reveals that medical segmentation datasets are significantly biased, primarily influenced by th…

Cited by 0SourceScholar
2024

Alice Benchmarks: Connecting Real World Re-Identification with the Synthetic

ICLR 2024poster

For object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap betwee…

Cited by 0SourcePDFScholar
2024

Improving Out-of-Distribution Generalization of Trajectory Prediction for Autonomous Driving via Polynomial Representations

IROS 2024poster

Robustness against Out-of-Distribution (OoD) samples is a key performance indicator of a trajectory prediction model. However, the development and ranking of state-of-the-art (SotA) models are driven by their In-Distribution (ID) performance on individual competition datasets. We present an OoD test…

Cited by 5SourcecodeScholar
2024

Learning-Aided Warmstart of Model Predictive Control in Uncertain Fast-Changing Traffic

ICRA 2024poster

Model Predictive Control lacks the ability to escape local minima in nonconvex problems. Furthermore, in fast-changing, uncertain environments, the conventional warmstart, using the optimal trajectory from the last timestep, often falls short of providing an adequately close initial guess for the cu…

Cited by 4SourceScholar
2024

Open-Set Facial Expression Recognition

AAAI 2024technical

Facial expression recognition (FER) models are typically trained on datasets with a fixed number of seven basic classes. However, recent research works (Cowen et al. 2021; Bryant et al. 2022; Kollias 2023) point out that there are far more expressions than the basic ones. Thus, when these models are…

Cited by 4SourcePDFScholar
2022

MOS Predictor for Synthetic Speech with I-Vector Inputs

ICASSP 2022accepted

Based on deep learning technology, non-intrusive methods have received increasing attention for synthetic speech quality assessment since it does not need reference signals. Meanwhile, i-vector has been widely used in paralinguistic speech attribute recognition such as speaker and emotion recognitio…

Cited by 0SourceScholar
2020

Simulating Content Consistent Vehicle Datasets with Attribute Descent

ECCV 2020poster

This paper uses a graphic engine to simulate a large amount of training data with free annotations. Between synthetic and real data, there is a two-level domain gap, i.e., content level and appearance level. While the latter has been widely studied, we focus on reducing the content gap in attributes…