← Search

Qi Yu

54 accepted papers

2026

AvAtar: Learning to Align via Active Optimal Transport

ICML 2026poster

Alignment plays a fundamental role in many machine learning problems, such as multi-network analysis, multimodal learning, and point cloud registration. Recent works increasingly leverage optimal transport (OT) for distributional alignment, whose effectiveness largely depends on sparse supervision t…

Cited by 0SourceScholar
2026

Calibrated Knowledge Aggregation in Bayesian Mixture-of-Experts for Continual VQA

ICML 2026poster

Continual learning for visual question answering (VQA) is typically implemented by training one expert per task and routing each query using task-ID supervision. Yet continual VQA tasks overlap substantially: on the VQA-v2 task stream, a non-native expert outperforms the task’s own expert on $49.9\%…

Cited by 0SourceScholar
2026

Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question Answering

ICLR 2026poster

Recent advances in large language models (LLMs) have improved the accuracy of visual question answering (VQA) systems. However, directly applying LLMs to VQA still presents several challenges: (a) suboptimal performance when handling questions from specialized domains, (b) higher computational costs…

Cited by 0SourceScholar
2026

Mixing Expertise with Confidence: A Mixture of Expert Framework for Robust Multi-Modal Continual Learner

ICML 2026poster

The Mixture of Experts (MoE) framework is widely used in continual learning to mitigate catastrophic forgetting. MoEs typically combine a small inter-task shared parameter space with largely independent expert parameters. However, as the number of tasks increases, the shared space becomes a bottlene…

Cited by 0SourceScholar
2026

PLANETALIGN: A Comprehensive Python Library for Benchmarking Network Alignment

ICLR 2026poster

Network alignment (NA) aims to identify node correspondence across different networks and serves as a critical cornerstone behind various downstream multi-network learning tasks. Despite growing research in NA, there lacks a comprehensive library that facilitates the systematic development and bench…

Cited by 0SourcecodeScholar
2026

The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection

CVPR 2026

Weakly supervised learning (WSL) provides a cost-effective learning paradigm for video anomaly detection (VAD) from data with video-level annotation instead of requiring costly fine-grained segment-level annotation. Although contemporary methods have shown promising results on challenging real-world

Cited by 0SourceScholar
2025

GLEN: Generalized Focal Loss Ensemble of Low-Rank Networks for Calibrated Visual Question Answering

AAAI 2025technical

Deep learning models with large-scale backbones have been increasingly adopted to tackle complex visual question answering (VQA) problems in real settings. While providing powerful learning capacities to handle the high-dimensional and multimodal VQA data, these models tend to suffer from the memori…

Cited by 0SourcePDFScholar
2025

Learning State-Based Node Representations from a Class Hierarchy for Fine-Grained Open-Set Detection

ICML 2025poster

Fine-Grained Openset Detection (FGOD) poses a fundamental challenge due to the similarity between the openset classes and those closed-set ones. Since real-world objects/entities tend to form a hierarchical structure, the fine-grained relationship among the closed-set classes as captured by the hier…

Cited by 0SourcePDFScholar
2025

Looking into User’s Long-term Interests through the Lens of Conservative Evidential Learning

ICLR 2025poster

Reinforcement learning (RL) provides an effective means to capture users' evolving preferences, leading to improved recommendation performance over time. However, existing RL approaches primarily rely on standard exploration strategies, which are less effective for a large item space with sparse rew…

Cited by 1SourcePDFScholar
2024

Adaptive Important Region Selection with Reinforced Hierarchical Search for Dense Object Detection

NeurIPS 2024poster

Existing state-of-the-art dense object detection techniques tend to produce a large number of false positive detections on difficult images with complex scenes because they focus on ensuring a high recall. To improve the detection accuracy, we propose an Adaptive Important Region Selection (AIRS) fr…

Cited by 0SourcePDFScholar
2024

Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection

ICML 2024poster

Subset or core-set selection offers a data-efficient way for training deep learning models. One-shot subset selection poses additional challenges as subset selection is only performed once and full set data become unavailable after the selection. However, most existing methods tend to choose either…

Cited by 1SourcePDFScholar
2024

Be Confident in What You Know: Bayesian Parameter Efficient Fine-Tuning of Vision Foundation Models

NeurIPS 2024poster

Large transformer-based foundation models have been commonly used as pre-trained models that can be adapted to different challenging datasets and settings with state-of-the-art generalization performance. Parameter efficient fine-tuning ($\texttt{PEFT}$) provides promising generalization performance…

Cited by 3SourcePDFScholar
2024

Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning Sensitivity

NeurIPS 2024poster

Multi-label active learning is a crucial yet challenging area in contemporary machine learning, often complicated by a large and sparse label space. This challenge is further exacerbated in active learning scenarios where labeling resources are constrained. Drawing inspiration from existing mixture…

Cited by 0SourcePDFScholar
2024

Evidential Stochastic Differential Equations for Time-Aware Sequential Recommendation

NeurIPS 2024poster

Sequential recommender systems are designed to capture users' evolving interests over time. Existing methods typically assume a uniform time interval among consecutive user interactions and may not capture users' continuously evolving behavior in the short and long term. In reality, the actual time…

Cited by 0SourcePDFScholar
2024

GRIT: A Dataset of Group Reference Recognition in Italian

COLING 2024main

For the analysis of political discourse a reliable identification of group references, i.e., linguistic components that refer to individuals or groups of people, is useful. However, the task of automatically recognizing group references has not yet gained much attention within NLP. To address this g…

Cited by 1SourcePDFScholar
2023

Actively Testing Your Model While It Learns: Realizing Label-Efficient Learning in Practice

NeurIPS 2023poster

In active learning (AL), we focus on reducing the data annotation cost from the model training perspective. However, "testing'', which often refers to the model evaluation process of using empirical risk to estimate the intractable true generalization risk, also requires data annotations. The annota…

2023

Deep Temporal Sets with Evidential Reinforced Attentions for Unique Behavioral Pattern Discovery

ICML 2023poster

Machine learning-driven human behavior analysis is gaining attention in behavioral/mental healthcare, due to its potential to identify behavioral patterns that cannot be recognized by traditional assessments. Real-life applications, such as digital behavioral biomarker identification, often require…

Cited by 8SourcePDFScholar
2023

Discover-Then-Rank Unlabeled Support Vectors in the Dual Space for Multi-Class Active Learning

ICML 2023poster

We propose to approach active learning (AL) from a novel perspective of discovering and then ranking potential support vectors by leveraging the key properties of the dual space of a sparse kernel max-margin predictor. We theoretically analyze the change of a hinge loss in the dual form and provide…

Cited by 1SourcePDFScholar
2023

Distributionally Robust Ensemble of Lottery Tickets Towards Calibrated Sparse Network Training

NeurIPS 2023poster

The recently developed sparse network training methods, such as Lottery Ticket Hypothesis (LTH) and its variants, have shown impressive learning capacity by finding sparse sub-networks from a dense one. While these methods could largely sparsify deep networks, they generally focus more on realizing…

2023

Figurative Language Processing: A Linguistically Informed Feature Analysis of the Behavior of Language Models and Humans

ACL 2023findings

Recent years have witnessed a growing interest in investigating what Transformer-based language models (TLMs) actually learn from the training data. This is especially relevant for complex tasks such as the understanding of non-literal meaning. In this work, we probe the performance of three black-b…

2023

Knowledge Acquisition for Human-In-The-Loop Image Captioning

AISTATS 2023poster

Image captioning offers a computational process to understand the semantics of images and convey them using descriptive language. However, automated captioning models may not always generate satisfactory captions due to the complex nature of the images and the quality/size of the training data. We p…

2023

STARS: Spatial-Temporal Active Re-sampling for Label-Efficient Learning from Noisy Annotations

AAAI 2023technical

Active learning (AL) aims to sample the most informative data instances for labeling, which makes the model fitting data efficient while significantly reducing the annotation cost. However, most existing AL models make a strong assumption that the annotated data instances are always assigned correct…

Cited by 0SourcePDFScholar
2023

Scaling Up Dynamic Graph Representation Learning via Spiking Neural Networks

AAAI 2023technical

Recent years have seen a surge in research on dynamic graph representation learning, which aims to model temporal graphs that are dynamic and evolving constantly over time. However, current work typically models graph dynamics with recurrent neural networks (RNNs), making them suffer seriously from…

2023

Sparse Maximum Margin Learning from Multimodal Human Behavioral Patterns

AAAI 2023technical

We propose a multimodal data fusion framework to systematically analyze human behavioral data from specialized domains that are inherently dynamic, sparse, and heterogeneous. We develop a two-tier architecture of probabilistic mixtures, where the lower tier leverages parametric distributions from th…

2022

A Dynamic Meta-Learning Model for Time-Sensitive Cold-Start Recommendations

AAAI 2022technical

We present a novel dynamic recommendation model that focuses on users who have interactions in the past but turn relatively inactive recently. Making effective recommendations to these time-sensitive cold-start users is critical to maintain the user base of a recommender system. Due to the sparse re…

2022

Dual-Level Adaptive Information Filtering for Interactive Image Segmentation

AISTATS 2022poster

Image segmentation can be performed interactively by accepting user annotations to refine the segmentation. It seeks frequent feedback from humans, and the model is updated with a smaller batch of data in each iteration of the feedback loop. Such a training paradigm requires effective information fi…

Cited by 1SourcePDFScholar
2022

Spiking Graph Convolutional Networks

IJCAI 2022poster

Graph Convolutional Networks (GCNs) achieve an impressive performance due to the remarkable representation ability in learning the graph information. However, GCNs, when implemented on a deep network, require expensive computation power, making them difficult to be deployed on battery-powered device…

2021

A Continual Learning Framework for Uncertainty-Aware Interactive Image Segmentation

AAAI 2021technical

Deep learning models have achieved state-of-the-art performance in semantic image segmentation, but the results provided by fully automatic algorithms are not always guaranteed satisfactory to users. Interactive segmentation offers a solution by accepting user annotations on selective areas of the i…

2021

A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning

NeurIPS 2021poster

Multi-label classification (MLC) allows complex dependencies among labels, making it more suitable to model many real-world problems. However, data annotation for training MLC models becomes much more labor-intensive due to the correlated (hence non-exclusive) labels and a potential large and sparse…

Cited by 10SourcePDFScholar
2021

Distributionally Robust Optimization for Deep Kernel Multiple Instance Learning

AISTATS 2021poster

Multiple Instance Learning (MIL) provides a promising solution to many real-world problems, where labels are only available at the bag level but missing for instances due to a high labeling cost. As a powerful Bayesian non-parametric model, Gaussian Processes (GP) have been extended from classical s…

2020

Dynamic Fusion of Eye Movement Data and Verbal Narrations in Knowledge-rich Domains

NeurIPS 2020poster

We propose to jointly analyze experts' eye movements and verbal narrations to discover important and interpretable knowledge patterns to better understand their decision-making processes. The discovered patterns can further enhance data-driven statistical models by fusing experts' domain knowledge t…

Cited by 2SourcePDFScholar
2020

Multifaceted Uncertainty Estimation for Label-Efficient Deep Learning

NeurIPS 2020poster

We present a novel multi-source uncertainty prediction approach that enables deep learning (DL) models to be actively trained with much less labeled data. By leveraging the second-order uncertainty representation provided by subjective logic (SL), we conduct evidence-based theoretical analysis and f…

Cited by 41SourcePDFScholar
2019

Fast Direct Search in an Optimally Compressed Continuous Target Space for Efficient Multi-Label Active Learning

ICML 2019oral

Active learning for multi-label classification poses fundamental challenges given the complex label correlations and a potentially large and sparse label space. We propose a novel CS-BPCA process that integrates compressed sensing and Bayesian principal component analysis to perform a two-level labe…

Cited by 8SourcePDFScholar
2019

Integrating Bayesian and Discriminative Sparse Kernel Machines for Multi-class Active Learning

NeurIPS 2019poster

We propose a novel active learning (AL) model that integrates Bayesian and discriminative kernel machines for fast and accurate multi-class data sampling. By joining a sparse Bayesian model and a maximum margin machine under a unified kernel machine committee (KMC), the proposed model is able to ide…

Cited by 23SourcePDFScholar
2018

Cortico-Muscular Coherence Enhancement Via Sparse Signal Representation

ICASSP 2018accepted

Identifiction of specific cortico-muscular interactions is essential for understanding sensorimotor control. These interactions are commonly studied by analyzing cortico-muscular coherence (CMC) between electroencephalogram (EEG) and surface electromyogram (sEMG) recorded synchronously under a motor…

Cited by 0SourceScholar