← Search

Haokui Zhang

14 accepted papers

2026

Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection

ICML 2026poster

Open-Vocabulary Aerial Detection (OVAD) and Remote Sensing Visual Grounding (RSVG) have emerged as two key paradigms for aerial scene understanding. However, each paradigm suffers from inherent limitations when operating in isolation: OVAD is restricted to coarse category-level semantics, while RSVG…

Cited by 0SourceScholar
2026

UVLM: Benchmarking Video Language Model for Underwater World Understanding

AAAI 2026technical

Recently, video-language models (VidLMs) have gained widespread attention and adoption. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation

Cited by 0SourcePDFScholar
2025

Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy

ICCV 2025poster

A prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived thro…

2025

NN-Former: Rethinking Graph Structure in Neural Architecture Representation

CVPR 2025poster

The growing use of deep learning necessitates efficient network design and deployment, making neural predictors vital for estimating attributes such as accuracy and latency. Recently, Graph Neural Networks (GNNs) and transformers have shown promising performance in representing neural architectures.…

2025

TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

AAAI 2025technical

Currently, inspired by the success of vision-language models (VLMs), an increasing number of researchers are focusing on improving VLMs and have achieved promising results. However, most existing methods concentrate on optimizing the connector and enhancing the language model component, while neglec…

Cited by 4SourcePDFScholar
2024

Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation

NeurIPS 2024poster

A common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of down-projection and up-projection matrices, with the bottleneck…

Cited by 1SourcePDFScholar
2023

NAR-Former V2: Rethinking Transformer for Universal Neural Network Representation Learning

NeurIPS 2023poster

As more deep learning models are being applied in real-world applications, there is a growing need for modeling and learning the representations of neural networks themselves. An effective representation can be used to predict target attributes of networks without the need for actual training and de…

2022

Connecting Compression Spaces with Transformer for Approximate Nearest Neighbor Search

ECCV 2022poster

"We propose a generic feature compression method for Approximate Nearest Neighbor Search (ANNS) problems, which speeds up existing ANNS methods in a plug-and-play manner. Specifically, based on transformer, we propose a new network structure to compress the feature into a low dimensional space, and…

2022

ParC-Net: Position Aware Circular Convolution with Merits from ConvNets and Transformer

ECCV 2022poster

"Recently, vision transformers started to show impressive results which outperform large convolution based models significantly. However, in the area of small models for mobile or resource constrained devices, ConvNet still has its own advantages in both performance and model complexity. We propose…

2020

Memory-Efficient Hierarchical Neural Architecture Search for Image Denoising

CVPR 2020poster

Recently, neural architecture search (NAS) methods have attracted much attention and outperformed manually designed architectures on a few high-level vision tasks. In this paper, we propose HiNAS (Hierarchical NAS), an effort towards employing NAS to automatically design effective neural network arc…

Cited by 89PDFScholar
2020

Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation

IROS 2020poster

Video object segmentation plays a vital role to many robotic tasks, beyond the satisfied accuracy, quickly adapt to the new scenario with very limited annotations and conduct a quick inference are also important. In this paper, we are specifically concerned with the task of fast segmenting all pixel…

Cited by 14SourceScholar
2019

Exploiting Temporal Consistency for Real-Time Video Depth Estimation

ICCV 2019poster

Accuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation p…

Cited by 142PDFScholar