← Search

Ruimin Hu

27 accepted papers

2026

Enhancing Cross-subject Emotion Recognition via Heterogeneous Distribution Augmentation and Collaborative Learning

ICML 2026poster

Cross-subject emotion recognition aims to improve a model's generalization to previously unseen subjects. Existing methods are mainly built upon domain generalization or data augmentation, but suffer from two major limitations: 1) heavy dependence on modality-specific feature designs—almost exclusiv…

Cited by 0SourceScholar
2026

Revisiting Attention in the Dark for Low-Light Person Re-Identiffcation

AAAI 2026technical

Person re-identification (Re-ID) under extremely low-light conditions suffers from severe image degradation, which significantly impairs the extraction of identity-discriminative features. Existing methods struggle to recover semantic information that is obscured under poor illumination. To better u

Cited by 0SourcePDFScholar
2025

MDFG: Multi-Dimensional Fine-Grained Modeling for Fatigue Detection

AAAI 2025technical

Fatigue is a critical factor contributing to accidents in industries such as safety monitoring and engineering construction. Fatigue exhibits dynamic complexity and non-stationary characteristics, so there are many intermediate states of short-term variation between alert and fatigue. Capturing and…

2025

Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and Evaluations

AAAI 2025technical

Due to the successful development of deep image generation technology, forgery detection plays a more important role in social and economic security. Racial bias has not been explored thoroughly in the deep forgery detection field. In the paper, we first contribute a dedicated dataset called the Fai…

2024

Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model

AAAI 2024technical

Adversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low t…

2024

Hidden Follower Detection: How Is the Gaze-Spacing Pattern Embodied in Frequency Domain?

AAAI 2024technical

Spatiotemporal social behavior analysis is a technique that studies the social behavior patterns of objects and estimates their risks based on their trajectories. In social public scenarios such as train stations, hidden following behavior has become one of the most challenging issues due to its pro…

Cited by 1SourcePDFScholar
2024

Mutuality Attribute Makes Better Video Anomaly Detection

ICASSP 2024accepted

Video anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep featur…

Cited by 0SourceScholar
2024

Robust Heterophilic Graph Learning against Label Noise for Anomaly Detection

IJCAI 2024poster

Given clean labels, Graph Neural Networks (GNNs) have shown promising abilities for graph anomaly detection. However, real-world graphs are inevitably noisy labeled, which drastically degrades the performance of GNNs. To alleviate it, some studies follow the local consistency (a.k.a homophily) assum…

2023

Crowd-Level Abnormal Behavior Detection via Multi-Scale Motion Consistency Learning

AAAI 2023technical

Detecting abnormal crowd motion emerging from complex interactions of individuals is paramount to ensure the safety of crowds. Crowd-level abnormal behaviors (CABs), e.g., counter flow and crowd turbulence, are proven to be the crucial causes of many crowd disasters. In the recent decade, video anom…

Cited by 14SourcePDFScholar
2023

Don't Ignore Alienation and Marginalization: Correlating Fraud Detection

IJCAI 2023poster

The anonymity of online networks makes tackling fraud increasingly costly. Thanks to the superiority of graph representation learning, graph-based fraud detection has made significant progress in recent years. However, upgrading fraudulent strategies produces more advanced and difficult scams. One c…

Cited by 6SourcePDFScholar
2023

Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly Detection

NeurIPS 2023poster

Unsupervised image Anomaly Detection (UAD) aims to learn robust and discriminative representations of normal samples. While separate solutions per class endow expensive computation and limited generalizability, this paper focuses on building a unified framework for multiple classes. Under such a cha…

2022

Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve Refinement

ICASSP 2022accepted

Deep learning networks with deeper layers become a trend for their good performance but lacks the potential for real-time mobile deployment. Another challenge for paired training networks is the limited generalization capacity caused by the sample bias. To overcome these two challenges, we propose a…

Cited by 0SourceScholar
2021

Location Predicts You: Location Prediction via Bi-direction Speculation and Dual-level Association

IJCAI 2021poster

Location prediction is of great importance in location-based applications for the construction of the smart city. To our knowledge, existing models for location prediction focus on the users' preference on POIs from the perspective of the human side. However, modeling users' interests from the histo…

Cited by 0SourcePDFScholar
2019

Cross-view Identical Part Area Alignment for Person Re-identification

ICASSP 2019accepted

Person re-identification aims to associate images captured by non-overlapping cameras. It is a challenging task because images are often in different conditions such as background clutter, illumination variation, viewpoint changes and different camera settings. Viewpoint changes and pose variations…

Cited by 0SourceScholar
2019

Kullback-Leibler Divergence Frequency Warping Scale for Acoustic Scene Classification Using Convolutional Neural Network

ICASSP 2019accepted

Most of current best performing Acoustic Scene Classification (ASC) systems utilize Mel scale spectrograms with Convolutional Neural Networks (CNNs). Mel scale is a common way to suit frequency warping of human ears, with strict decreasing frequency resolution on low to high frequency range. However…

Cited by 0SourceScholar
2019

Long Term Background Reference Based Satellite Video Coding

ICASSP 2019accepted

Video transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellite calls for higher coding efficiency. In…

Cited by 0SourceScholar
2019

Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge

ICASSP 2019accepted

The rapidly increasing surveillance video data has challenged the existing video coding standards. Even though knowledge based video coding scheme proposed for moving objects so far has achieved high efficiency, it does not take full advantages of local information and highly relies on the accuracy…

Cited by 0SourceScholar
2017

A joint learning based Face Super Resolution approach via contextual topological structure

ICASSP 2017accepted

Face Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch bas…

Cited by 0SourceScholar
2017

Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction

ICASSP 2017accepted

NHK has proposed a famous 3D audio system: 22.2 multi-channel system, but its loudspeakers are too many and are troublesome to put in home. Ando and Wang has proposed two simplification methods to reduce its channel number, but only 3D sound field at the central listening point can be recovered well…

Cited by 0SourceScholar
2017

Transferring clothing parsing from fashion dataset to surveillance

ICASSP 2017accepted

In this paper we address the problem of automatic clothing parsing in surveillance video with the information from user-generated tags such as “jeans” and “T-shirt”. Although clothing parsing has achieved great success in fashion clothing, it is quite challenging to parse clothing in practical surve…

Cited by 0SourceScholar
2016

Multiple instance discriminative dictionary learning for action recognition

ICASSP 2016accepted

Action recognition from video is a prominent research area in computer vision, with far-reaching applications. Current state-of-the-art action recognition methods is Fisher Vector (FV) coding model based on spatio-temporal local features. Though high dimensional local features have more representati…

Cited by 0SourceScholar
2015

A down-mixing method for 22.2 multichannel system reproduction

ICASSP 2015accepted

This paper proposes a general multichannel system reproduction method. Firstly, relative to original multichannel system, a general global model is build up by guaranteeing sound pressure and the direction of particle velocity at the receiving point constant, and making the square error of particle…

Cited by 0SourceScholar
2015

Face hallucination via Cauchy regularized sparse representation

ICASSP 2015accepted

In dictionary-learning-based face hallucination, the testing image is represented as a linear combination of the training samples, and how to obtain the optimal coefficients is the primary issue. Sparse representation (SR) has ever been widely used in face hallucination, however, due to the fact tha…

Cited by 0SourceScholar
2015

Super-Resolution Person Re-Identification With Semi-Coupled Low-Rank Discriminant Dictionary Learning

CVPR 2015poster

Person re-identification has been widely studied due to its importance in surveillance and forensics applications. In practice, gallery images are high-resolution (HR) while probe images are usually low-resolution (LR) in the identification scenarios with large variation of illumination, weather or…

Cited by 284SourcePDFScholar