← Search

SHICHAO KAN

9 accepted papers

2026

HSGG: Training-Free Hierarchical Scene Graph Generation with Geometry-Guided Relation Reasoning

ICML 2026poster

Scene Graph Generation (SGG) connects visual perception with structured reasoning, but is limited by scarce annotations and the long-tailed distribution of relational predicates. Training-free methods based on vision-language models (VLMs) reduce supervision requirements, yet often rely on flat grap…

Cited by 0SourceScholar
2025

EPCPE: A Real-time End-to-End Pipeline for RGB-based Category-level 6D Pose Estimation

ICASSP 2025accepted

RGB-based category-level 6D pose estimation methods have faced significant challenges in achieving real-time performance, primarily due to the design of two-stage pipeline. To address this issue, we propose a novel end-to-end pipeline named EPCPE. In detail, we first extract implicit rotation featur…

Cited by 0SourceScholar
2025

How Does the Smoothness Approximation Method Facilitate Generalization for Federated Adversarial Learning?

AAAI 2025technical

Federated Adversarial Learning (FAL) is a robust framework for resisting adversarial attacks on federated learning. Although some FAL studies have developed efficient algorithms, they primarily focus on convergence performance and overlook generalization. Generalization is crucial for evaluating alg…

Cited by 0SourcePDFScholar
2025

Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization for Scene Graph Generation

ICML 2025poster

Scene Graph Generation (SGG) is a fundamental task in visual understanding, aimed at providing more precise local detail comprehension for downstream applications. Existing SGG methods often overlook the diversity of predicate representations and the consistency among similar predicates when dealing…

Cited by 0SourcePDFScholar
2023

Singularformer: Learning to Decompose Self-Attention to Linearize the Complexity of Transformer

IJCAI 2023poster

Transformers achieve excellent performance in a variety of domains since they can capture long-distance dependencies through the self-attention mechanism. However, self-attention is computationally costly due to its quadratic complexity and high memory consumption. In this paper, we propose a novel…

2022

Coded Residual Transform for Generalizable Deep Metric Learning

NeurIPS 2022accept

A fundamental challenge in deep metric learning is the generalization capability of the feature embedding network model since the embedding network learned on training classes need to be evaluated on new test classes. To address this challenge, in this paper, we introduce a new method called coded…

Cited by 4SourcePDFScholar
2021

Relative Order Analysis and Optimization for Unsupervised Deep Metric Learning

CVPR 2021poster

In unsupervised learning of image features without labels, especially on datasets with fine-grained object classes, it is often very difficult to tell if a given image belongs to one specific object class or another, even for human eyes. However, we can reliably tell if image C is more similar to im…

Cited by 14PDFcodeScholar
2020

Unsupervised Deep Metric Learning with Transformed Attention Consistency and Contrastive Clustering Loss

ECCV 2020poster

Existing approaches for unsupervised metric learning focus on exploring self-supervision information within the input image itself. We observe that, when analyzing images, human eyes often compare images against each other instead of examining images individually. In addition, they often pay attenti…

Cited by 26SourcePDFScholar