← Search

Yong Rui

20 accepted papers

2026

Collaborative Dual Representations for Semi-Supervised Partial Label Learning

AAAI 2026technical

Semi-supervised partial label learning (SSPLL) aims to improve the generalization performance of partial label (PL) classifiers by effectively leveraging unlabeled data. Nevertheless, the inherent ambiguity in supervision, where the ground-truth label of a PL example is hidden within a set of candid

Cited by 0SourcePDFScholar
2026

DivControl: Knowledge Diversion for Controllable Image Generation

AAAI 2026technical

Diffusion models have advanced from text-to-image (T2I) to image-to-image (I2I) generation by incorporating structured inputs such as depth maps, enabling fine-grained spatial control. However, existing methods either train separate models for each condition or rely on unified architectures with ent

Cited by 0SourcePDFScholar
2026

FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models

CVPR 2026

The training of diffusion models is computationally intensive, making effective pre-training essential. However, real-world deployments often demand models of variable sizes due to diverse memory and computational constraints, posing challenges when corresponding pre-trained versions are unavailable

Cited by 0SourceScholar
2026

Self-Supervised Weight Templates for Scalable Vision Model Initialization

ICML 2026poster

The increasing scale and complexity of modern model parameters underscore the importance of pre-trained models. However, deployment often demands architectures of varying sizes, exposing limitations of conventional pre-training and fine-tuning. To address this, we propose SWEET, a self-supervised fr…

Cited by 0SourceScholar
2025

Implicit Relative Labeling-Importance Aware Multi-Label Metric Learning

AAAI 2025technical

Multi-label metric learning, as an extension of metric learning to multi-label scenarios, aims to learn better similarity metrics for objects with rich semantics. Existing multi-label metric learning approaches employ the common assumption of equal labeling-importance, i.e., all associated labels ar…

Cited by 0SourcePDFScholar
2025

KIND: Knowledge Integration and Diversion for Training Decomposable Models

ICML 2025poster

Pre-trained models have become the preferred backbone due to the increasing complexity of model parameters. However, traditional pre-trained models often face deployment challenges due to their fixed sizes, and are prone to negative transfer when discrepancies arise between training tasks and target…

2021

What If We Could Not See? Counterfactual Analysis for Egocentric Action Anticipation

IJCAI 2021poster

Egocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacit…

Cited by 16SourcePDFScholar
2020

Label Distribution Learning on Auxiliary Label Space Graphs for Facial Expression Recognition

CVPR 2020poster

Many existing studies reveal that annotation inconsistency widely exists among a variety of facial expression recognition (FER) datasets. The reason might be the subjectivity of human annotators and the ambiguous nature of the expression labels. One promising strategy tackling such a problem is a re…

Cited by 244PDFScholar
2016

Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette Consistency

CVPR 2016poster

In this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouett…

Cited by 7PDFScholar
2016

Jointly Modeling Embedding and Translation to Bridge Video and Language

CVPR 2016oral

Automatically describing video content with natural language is a fundamental challenge of computer vision. Recurrent Neural Networks (RNNs), which models sequence dynamics, has attracted increasing attention on visual interpretation. However, most existing approaches generate a word locally with th…

Cited by 716PDFScholar
2015

MeshStereo: A Global Stereo Model With Mesh Alignment Regularization for View Interpolation

ICCV 2015oral

We present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D tria…

Cited by 206PDFScholar
2015

Query Adaptive Similarity Measure for RGB-D Object Recognition

ICCV 2015poster

This paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object…

Cited by 18PDFScholar
2015

Relaxing From Vocabulary: Robust Weakly-Supervised Deep Learning for Vocabulary-Free Image Tagging

ICCV 2015poster

The development of deep learning has empowered machines with comparable capability of recognizing limited image categories to human beings. However, most existing approaches heavily rely on human-curated training data, which hinders the scalability to large and unlabeled vocabularies in image taggin…

Cited by 48PDFScholar