← Search

Yu Yao

32 accepted papers

2026

Early Warning of Intraoperative Adverse Events via Transformer-Driven Multi-Label Learning

AAAI 2026technical

Early warning of intraoperative adverse events plays a vital role in reducing surgical risk and improving patient safety. While deep learning has shown promise in predicting the single adverse event, several key challenges remain: overlooking adverse event dependencies, underutilizing heterogeneous

Cited by 0SourcePDFScholar
2025

A Conditional Independence Test in the Presence of Discretization

ICLR 2025poster

Testing conditional independence (CI) has many important applications, such as Bayesian network learning and causal discovery. Although several approaches have been developed for learning CI structures for observed variables, those existing methods generally fail to work when the variables of intere…

2025

A Lens into Interpretable Transformer Mistakes via Semantic Dependency

ICML 2025poster

Semantic Dependency refers to the relationship between words in a sentence where the meaning of one word depends on another, which is important for natural language understanding. In this paper, we investigate the role of semantic dependencies in answering questions for transformer models, which is…

Cited by 0SourcePDFScholar
2025

A Robust Method to Discover Causal or Anticausal Relation

ICLR 2025poster

Understanding whether the data generative process follows causal or anticausal relations is important for many applications. Existing causal discovery methods struggle with high-dimensional perceptual data such as images. Moreover, they require well-labeled data, which may not be feasible due to mea…

Cited by 0SourcePDFScholar
2025

A Sample Efficient Conditional Independence Test in the Presence of Discretization

ICML 2025poster

Conditional independence (CI) test is a fundamental concept in statistics. In many real-world scenarios, some variables may be difficult to measure accurately, often leading to data being represented as discretized values. Applying CI tests directly to discretized data, however, can lead to incorrec…

2025

AGENTS-LLM: Augmentative GENeration of Challenging Traffic Scenarios with an Agentic LLM Framework

IROS 2025

Rare, yet critical, scenarios pose a significant challenge in testing and evaluating autonomous driving planners. Relying solely on real-world driving scenes requires collecting massive datasets to capture these scenarios. While automatic generation of traffic scenarios appears promising, data-drive

Cited by 4SourceScholar
2025

Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation

NeurIPS 2025poster

Text-to-Audio-Video (T2AV) generation aims to produce temporally and semantically aligned visual and auditory content from natural language descriptions. While recent progress in text-to-audio and text-to-video models has improved generation quality within each modality, jointly modeling them remain…

Cited by 0SourceScholar
2025

Can Dependencies Induced by LLM-Agent Workflows Be Trusted?

NeurIPS 2025poster

LLM-agent systems often decompose high-level objectives into subtask dependency graphs, assuming that each subtask’s output is reliable and conditionally independent of others given its parent responses. However, this assumption frequently breaks during execution, as ground-truth responses are inac…

Cited by 0SourcecodeScholar
2025

Chain-of-Focus Prompting: Leveraging Sequential Visual Cues to Prompt Large Autoregressive Vision Models

ICLR 2025poster

In-context learning (ICL) has revolutionized natural language processing by enabling models to adapt to diverse tasks with only a few illustrative examples. However, the exploration of ICL within the field of computer vision remains limited. Inspired by Chain-of-Thought (CoT) prompting in the langua…

Cited by 0SourcePDFScholar
2025

Decentralized Cooperative Localization: A Communication-Efficient Dual-Fusion Consistent Approach

RA-L 2025

Decentralized cooperative localization poses significant challenges in managing inter-robot correlations, especially in environments with limited communication capacity and unreliable network connectivity. In this letter, we propose a communication-efficient decentralized consistent cooperative loca

Cited by 3SourceScholar
2025

Flow: Modularized Agentic Workflow Automation

ICLR 2025poster

Multi-agent frameworks powered by large language models (LLMs) have demonstrated great success in automated planning and task execution. However, the effective adjustment of agentic workflows during execution has not been well studied. An effective workflow adjustment is crucial in real-world scenar…

2025

Ranked from Within: Ranking Large Multimodal Models Without Labels

ICML 2025poster

Can the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of givin…

Cited by 0SourcePDFScholar
2025

SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models

ICML 2025poster

Traditional autonomous driving systems often struggle to connect high-level reasoning with low-level control, leading to suboptimal and sometimes unsafe behaviors. Recent advances in multimodal large language models (MLLMs), which process both visual and textual data, offer an opportunity to unify p…

2025

Seeking Rational Demonstrations for Large Language Models: A Domain Generalization Approach to Unsupervised Cross-Domain Keyphrase Generation

ACL 2025short

Unsupervised cross-domain keyphrase generation is crucial in real-world natural language processing scenarios. However, the accuracy of up-to-date approaches is limited by the distribution shift between source and target domain, which stems from the cross-domain field. Large language models (LLMs) o…

2025

SmartCLIP: Modular Vision-language Alignment with Identification Guarantees

CVPR 2025highlight

Contrastive Language-Image Pre-training (CLIP) \citep radford2021learning has emerged as a pivotal model in computer vision and multimodal learning, achieving state-of-the-art performance at aligning visual and textual representations through contrastive learning. However, CLIP struggles with poten…

2024

Enhancing Contrastive Learning for Ordinal Regression via Ordinal Content Preserved Data Augmentation

ICLR 2024poster

Contrastive learning, while highly effective for a lot of tasks, shows limited improvement in ordinal regression. We find that the limitation comes from the predefined strong data augmentations employed in contrastive learning. Intuitively, for ordinal regression datasets, the discriminative inform…

Cited by 9SourcePDFScholar
2024

Identifying Latent State-Transition Processes for Individualized Reinforcement Learning

NeurIPS 2024poster

The application of reinforcement learning (RL) involving interactions with individuals has grown significantly in recent years. These interactions, influenced by factors such as personal preferences and physiological differences, causally influence state transitions, ranging from health conditions i…

Cited by 3SourcePDFScholar
2024

Improving Non-Transferable Representation Learning by Harnessing Content and Style

ICLR 2024spotlight

Non-transferable learning (NTL) aims to restrict the generalization of models toward the target domain(s). To this end, existing works learn non-transferable representations by reducing statistical dependence between the source and target domain. However, such statistical methods essentially neglect…

Cited by 24SourcePDFScholar
2024

Rethinking the Paradigm of Content Constraints in Unpaired Image-to-Image Translation

AAAI 2024technical

In an unpaired setting, lacking sufficient content constraints for image-to-image translation (I2I) tasks, GAN-based approaches are usually prone to model collapse. Current solutions can be divided into two categories, reconstruction-based and Siamese network-based. The former requires that the tran…

2023

CS-Isolate: Extracting Hard Confident Examples by Content and Style Isolation

NeurIPS 2023poster

Label noise widely exists in large-scale image datasets. To mitigate the side effects of label noise, state-of-the-art methods focus on selecting confident examples by leveraging semi-supervised learning. Existing research shows that the ability to extract hard confident examples, which are close to…

2023

KD-EKF: A Consistent Cooperative Localization Estimator Based on Kalman Decomposition

IROS 2023poster

In this paper, we revisit the inconsistency problem of EKF-based cooperative localization (CL) from the perspective of system decomposition. By transforming the linearized system used by the standard EKF into its Kalman observable canonical form, the observable and unobservable components of the sys…

Cited by 5SourceScholar
2023

Obstacle-Aware Topological Planning over Polyhedral Representation for Quadrotors

ICRA 2023poster

In this paper, we propose a novel mapping-planning framework for autonomous quadrotor navigation. First, a polyhedron-based mapping algorithm is presented to fully exploit the information of the onboard sensor data. Polyhedra are generated to approximate the segmented clusters of occupied voxels. Th…

Cited by 5SourceScholar
2023

Which is Better for Learning with Noisy Labels: The Semi-supervised Method or Modeling Label Noise?

ICML 2023poster

In real life, accurately annotating large-scale datasets is sometimes difficult. Datasets used for training deep learning models are likely to contain label noise. To make use of the dataset containing label noise, two typical methods have been proposed. One is to employ the semi-supervised method b…

Cited by 10SourcePDFScholar
2022

Rethinking Class-Prior Estimation for Positive-Unlabeled Learning

ICLR 2022poster

Given only positive (P) and unlabeled (U) data, PU learning can train a binary classifier without any negative data. It has two building blocks: PU class-prior estimation (CPE) and PU classification; the latter has been well studied while the former has received less attention. Hitherto, the distrib…

Cited by 26SourcePDFScholar
2021

BiTraP: Bi-Directional Pedestrian Trajectory Prediction With Multi-Modal Goal Estimation

RA-L 2021

Pedestrian trajectory prediction is an essential task in robotic applications such as autonomous driving and robot navigation. State-of-the-art trajectory predictors use a conditional variational autoencoder (CVAE) with recurrent neural networks (RNNs) to encode observed trajectories and decode mult

Cited by 185SourcecodeScholar
2021

Coupling Intent and Action for Pedestrian Crossing Behavior Prediction

IJCAI 2021poster

Accurate prediction of pedestrian crossing behaviors by autonomous vehicles can significantly improve traffic safety. Existing approaches often model pedestrian behaviors using trajectories or poses but do not offer a deeper semantic interpretation of a person's actions or how actions influence a pe…

2021

Instance-dependent Label-noise Learning under a Structural Causal Model

NeurIPS 2021poster

Label noise generally degenerates the performance of deep learning algorithms because deep neural networks easily overfit label errors. Let $X$ and $Y$ denote the instance and clean label, respectively. When $Y$ is a cause of $X$, according to which many datasets have been constructed, e.g., \text…

Cited by 90SourcePDFScholar
2020

Dual T: Reducing Estimation Error for Transition Matrix in Label-noise Learning

NeurIPS 2020poster

The transition matrix, denoting the transition relationship from clean labels to noisy labels, is essential to build statistically consistent classifiers in label-noise learning. Existing methods for estimating the transition matrix rely heavily on estimating the noisy class posterior. However, the…

Cited by 300SourcePDFScholar
2019

Egocentric Vision-based Future Vehicle Localization for Intelligent Driving Assistance Systems

ICRA 2019poster

Predicting the future location of vehicles is essential for safety-critical applications such as advanced driver assistance systems (ADAS) and autonomous driving. This paper introduces a novel approach to simultaneously predict both the location and scale of target vehicles in the first-person (egoc…

Cited by 172SourceScholar
2019

Unsupervised Traffic Accident Detection in First-Person Videos

IROS 2019poster

Recognizing abnormal events such as traffic violations and accidents in natural driving scenes is essential for successful autonomous driving and advanced driver assistance systems. However, most work on video anomaly detection suffers from two crucial drawbacks. First, they assume cameras are fixed…

Cited by 209SourcecodeScholar