← Search

Zhihui Wang

22 accepted papers

2026

GOCM: Single-Step Graph Outlier Synthesis via Origin Consistency Model

ICML 2026poster

Supervised Graph Outlier Detection has long been constrained by severe class imbalance, and although recent diffusion-based augmentation methods have improved sample quality, their practical utility is hindered by the high computational costs of multi-step iterative sampling and the stochasticity of…

Cited by 0SourceScholar
2025

2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images

ICASSP 2025accepted

Classifying the status of NSCLC PD-L1 on chest CT is a cost-effective and non-invasive method. The existing multiple instance learning (MIL) methods are not effective for this task, due to the lack of an efficient feature encoder for 3D instances and ignoring the importance of representative instanc…

Cited by 0SourceScholar
2025

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens

ICCV 2025poster

Recently, Vision Large Language Models (VLLMs) with integrated vision encoders have shown promising performance in vision understanding. They encode visual content into sequences of visual tokens, enabling joint processing of visual and textual data. However, understanding videos, especially long vi…

2025

RobAVA: A Large-scale Dataset and Baseline Towards Video based Robotic Arm Action Understanding

ICCV 2025poster

Understanding the behaviors of robotic arms is essential for various robotic applications such as logistics management, precision agriculture, and automated manufacturing. However, the lack of large-scale and diverse datasets significantly hinders progress in video-based robotic arm action understan…

2025

Towards Efficient and Intelligent Laser Weeding: Method and Dataset for Weed Stem Detection

AAAI 2025technical

Weed control is a critical challenge in modern agriculture, as weeds compete with crops for essential nutrient resources, significantly reducing crop yield and quality. Traditional weed control methods, including chemical and mechanical approaches, have real-life limitations such as associated envir…

2024

Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning

ICASSP 2024accepted

Automatic anime sketch colorization aims to generate a color image from a sketch image, which is challenging due to limited structure and semantic understanding, leading to constrained style, and semantic color inconsistency. In this paper, we introduce a sketch to color diffusion model with semanti…

Cited by 0SourceScholar
2023

Decoupling with Entropy-based Equalization for Semi-Supervised Semantic Segmentation

IJCAI 2023poster

Semi-supervised semantic segmentation methods are the main solution to alleviate the problem of high annotation consumption in semantic segmentation. However, the class imbalance problem makes the model favor the head classes with sufficient training samples, resulting in poor performance of the tai…

Cited by 3SourcePDFScholar
2023

Fine-Grained Retrieval Prompt Tuning

AAAI 2023technical

Fine-grained object retrieval aims to learn discriminative representation to retrieve visually similar objects. However, existing top-performing works usually impose pairwise similarities on the semantic embedding spaces or design a localization sub-network to continually fine-tune the entire model…

Cited by 21SourcePDFScholar
2023

Learning to Parameterize Visual Attributes for Open-set Fine-grained Retrieval

NeurIPS 2023poster

Open-set fine-grained retrieval is an emerging challenging task that allows to retrieve unknown categories beyond the training set. The best solution for handling unknown categories is to represent them using a set of visual attributes learnt from known categories, as widely used in zero-shot learn…

Cited by 11SourcePDFScholar
2023

Online Visual SLAM Adaptation against Catastrophic Forgetting with Cycle-Consistent Contrastive Learning

ICRA 2023poster

Visual SLAM (Simultaneous Localisation and Mapping) aims to simultaneously estimate camera poses and depth maps from navigation videos captured. While recent deep learning based methods have achieved great success on this task, they tend to work well on source domain data and suffer from performance…

Cited by 3SourceScholar
2023

Open-Set Fine-Grained Retrieval via Prompting Vision-Language Evaluator

CVPR 2023poster

Open-set fine-grained retrieval is an emerging challenge that requires an extra capability to retrieve unknown subcategories during evaluation. However, current works are rooted in the close-set scenarios, where all the subcategories are pre-defined, and make it hard to capture discriminative knowle…

Cited by 22SourcePDFScholar
2023

Towards Fair and Comprehensive Comparisons for Image-Based 3D Object Detection

ICCV 2023poster

In this work, we build a modular-designed codebase, formulate strong training recipes, design an error diagnosis toolbox, and discuss current methods for image-based 3D object detection. Specifically, different from other highly mature tasks, e.g., 2D object detection, the community of image-based 3…

Cited by 3PDFcodeScholar
2022

Category-Specific Nuance Exploration Network for Fine-Grained Object Retrieval

AAAI 2022technical

Employing additional prior knowledge to model local features as a final fine-grained object representation has become a trend for fine-grained object retrieval (FGOR). A potential limitation of these methods is that they only focus on common parts across the dataset (e.g. head, body or even leg) by…

Cited by 14SourcePDFScholar
2022

MonoDistill: Learning Spatial Features for Monocular 3D Object Detection

ICLR 2022poster

3D object detection is a fundamental and challenging task for 3D scene understanding, and the monocular-based methods can serve as an economical alternative to the stereo-based or LiDAR-based methods. However, accurately locating objects in the 3D space from a single image is extremely difficult due…

2021

Cost Affinity Learning Network for Stereo Matching

ICASSP 2021accepted

Existing stereo matching methods mainly tend to directly aggregate features output from Convolutional Neural Network to obtain more discriminative cost features, but ignore the affinity of each element in the cost feature which also plays a key role in enhancing the cost feature. In this work, we pr…

Cited by 0SourceScholar
2021

Dynamic Position-aware Network for Fine-grained Image Recognition

AAAI 2021technical

Most weakly supervised fine-grained image recognition (WFGIR) approaches predominantly focus on learning the discriminative details which contain the visual variances and position clues. The position clues can be indirectly learnt by utilizing context information of discriminative visual content. Ho…

Cited by 35SourcePDFScholar
2021

Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution

CVPR 2021poster

Existing color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training examples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected…

Cited by 55PDFScholar
2020

Renovating Parsing R-CNN for Accurate Multiple Human Parsing

ECCV 2020poster

Multiple human parsing aims to segment various human parts and associate each part with the corresponding instance simultaneously. This is a very challenging task due to the diverse human appearance, semantic ambiguity of different body parts and clothing, and complex background. Through analysis of…

2020

Weakly Supervised Fine-Grained Image Classification via Guassian Mixture Model Oriented Discriminative Learning

CVPR 2020oral

Existing weakly supervised fine-grained image recognition (WFGIR) methods usually pick out the discriminative regions from the high-level feature maps directly. We discover that due to the operation of stacking local receptive filed, Convolutional Neural Network causes the discriminative region diff…

Cited by 104PDFScholar
2019

Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving

ICCV 2019poster

In this paper, we propose a monocular 3D object detection framework in the domain of autonomous driving. Unlike previous image-based methods which focus on RGB feature extracted from 2D images, our method solves this problem in the reconstructed 3D space in order to exploit 3D contexts explicitly. T…

Cited by 396PDFScholar
2019

Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New Baseline

ICASSP 2019accepted

Online tracking a specific person from a low-altitude unmanned aerial vehicle (UAV) is a very interesting and challenging problem to be solved. However, there exists no large-scale aerial video dataset regarding this online single person tracking (OSPT) task. To promote the study of the OSPT problem…

Cited by 0SourceScholar