← Search

Ruize Han

11 accepted papers

2026

InfoScan: Information-Efficient Visual Scanning via Resource-Adaptive Walks

ICLR 2026poster

High-resolution visual representation learning remains challenging due to the quadratic complexity of Vision Transformers and the limitations of existing efficient approaches, where fixed scanning patterns in recent Mamba-based models hinder content-adaptive perception. To address these limitations,…

Cited by 0SourceScholar
2026

NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection

CVPR 2026

Despite the remarkable progress in open-vocabulary object detection (OVD), a significant gap remains between the training and testing phases. During training, the RPN and RoI heads often misclassify unlabeled novel-category objects as background, causing some proposals to be prematurely filtered out

Cited by 0SourceScholar
2024

Robust Collaborative Perception without External Localization and Clock Devices

ICRA 2024poster

A consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information exchange among agents. To achieve this spatial-temporal alignment, traditional methods depend on external devices to provide…

Cited by 4SourceScholar
2022

Connecting the Complementary-View Videos: Joint Camera Identification and Subject Association

CVPR 2022poster

We attempt to connect the data from complementary views, i.e., top view from drone-mounted cameras in the air, and side view from wearable cameras on the ground. Collaborative analysis of such complementary-view data can facilitate to build the air-ground cooperative visual system for various kinds…

Cited by 13PDFcodeScholar
2022

Panoramic Human Activity Recognition

ECCV 2022poster

"To obtain a more comprehensive activity understanding for a crowded scene, in this paper, we propose a new problem of panoramic human activity recognition (PAR), which aims to simultaneously achieve the the recognition of individual actions, social group activities, and global activities. This is a…

2022

Self-Supervised Social Relation Representation for Human Group Detection

ECCV 2022poster

"Human group detection, which splits crowd of people into groups, is an important step for video-based human social activity analysis. The core of human group detection is the human social relation representation and division. In this paper, we propose a new two-stage multi-head framework for human…

2020

Key Action and Joint CTC-Attention based Sign Language Recognition

ICASSP 2020accepted

Sign Language Recognition (SLR) translates sign language video into natural language. In practice, sign language video, owning a large number of redundant frames, is necessary to be selected the essential. However, unlike common video that describes actions, sign language video is characterized as c…

Cited by 0SourceScholar