← Search

Yitian Zhao

14 accepted papers

2025

Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?

CVPR 2025poster

Spatial-temporal graph convolutional networks (ST-GCNs) showcase impressive performance in skeleton-based human action recognition (HAR). However, despite the development of numerous models, their recognition performance does not differ significantly after aligning the input settings. With this obse…

2025

DARR: A Dual-Branch Arithmetic Regression Reasoning Framework for Solving Machine Number Reasoning

AAAI 2025technical

Abstract visual reasoning (AVR) is a critical ability of humans, and it has been widely studied, but arithmetic visual reasoning, a unique task in AVR to reason over number sense, is less studied in the literature. To facilitate this research, we construct a Machine Number Reasoning (MNR) dataset to…

2025

DBCR: Exploiting Both Intra-cluster and Extra-cluster Relations for Compositional Reasoning

ICASSP 2025accepted

Most existing models for abstract visual reasoning perform poorly in compositional visual reasoning (CVR), due to complex nature of compositional rules and difficulties in distinguishing tiny rule differences between outliers and normal images. To tackle the challenges, we propose a Dual-Branch Comp…

Cited by 0SourceScholar
2025

DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMs

NeurIPS 2025poster

Abstract Visual Reasoning (AVR) entails discerning latent patterns in visual data and inferring underlying rules. Existing solutions often lack scalability and adaptability, as deep architectures tend to overfit training data, and static neural networks fail to dynamically capture diverse rules. To…

Cited by 0SourcecodeScholar
2024

Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action Recognition

AAAI 2024technical

Graph convolutional networks (GCNs) have attracted great attention and achieved remarkable performance in skeleton-based action recognition. However, most of the previous works are designed to refine skeleton topology without considering the types of different joints and edges, making them infeasibl…

2024

Regression Residual Reasoning with Pseudo-labeled Contrastive Learning for Uncovering Multiple Complex Compositional Relations

IJCAI 2024poster

Abstract Visual Reasoning (AVR) has been widely studied in literature. Our study reveals that AVR models tend to rely on appearance matching rather than a genuine understanding of underlying rules. We hence develop a challenging benchmark, Multiple Complex Compositional Reasoning (MC2R), composed of…

Cited by 4SourcePDFScholar
2024

Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone Imagery

AAAI 2024technical

Object detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of obje…

2022

DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image Classification

CVPR 2022oral

Multiple instance learning (MIL) has been increasingly used in the classification of histopathology whole slide images (WSIs). However, MIL approaches for this specific classification problem still face unique challenges, particularly those related to small sample cohorts. In these, there are limite…

Cited by 407PDFcodeScholar
2022

Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification

ICASSP 2022accepted

Multi-label image classification is a fundamental but challenging task in computer vision. Over the past few decades, solutions exploring relationships between semantic labels have made great progress. However, the underlying spatial-contextual information of labels is under-exploited. To tackle thi…

Cited by 0SourceScholar
2021

Mesh Saliency: An Independent Perceptual Measure or a Derivative of Image Saliency?

CVPR 2021poster

While mesh saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and is well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art mesh saliency methods remain poor at predicting human fixatio…

Cited by 21PDFcodeScholar
2021

Spatial Uncertainty-Aware Semi-Supervised Crowd Counting

ICCV 2021poster

Semi-supervised approaches for crowd counting attract attention, as the fully supervised paradigm is expensive and laborious due to its request for a large number of images of dense crowd scenarios and their annotations. This paper proposes a spatial uncertainty-aware semi-supervised approach via re…

Cited by 123PDFcodeScholar
2020

Regression of Instance Boundary by Aggregated CNN and GCN

ECCV 2020poster

This paper proposes a straightforward, intuitive deep learning approach for (biomedical) image segmentation tasks. Different from the existing dense pixel classification methods, we develop a novel multilevel aggregation network to directly regress the coordinates of the boundary of instances in an…

Cited by 33SourcePDFScholar
2020

Unsupervised Multi-View CNN for Salient View Selection of 3D Objects and Scenes

ECCV 2020poster

We present an unsupervised 3D deep learning framework based on a ubiquitously true proposition named by us view-object consistency as it states that a 3D object and its projected 2D views always belong to the same object class. To validate its effectiveness, we design a multi-view CNN instantiating…

2019

Topology Reconstruction of Tree-Like Structure in Images via Structural Similarity Measure and Dominant Set Clustering

CVPR 2019poster

The reconstruction and analysis of tree-like topological structures in the biomedical images is crucial for biologists and surgeons to understand biomedical conditions and plan surgical procedures. The underlying tree-structure topology reveals how different curvilinear components are anatomically…

Cited by 13PDFScholar