← Search

Zhengqin Xu

18 accepted papers

2026

CP-CLIP: Customized Parameter Generation for Open-vocabulary Semantic Segmentation

AAAI 2026technical

Open-vocabulary semantic segmentation aims to assign pixel-level labels to images based on textual descriptions, even for categories beyond predefined closed sets. While vision-language foundation models like CLIP are widely used for this task, fine-tuning them for pixel-level predictions often comp

Cited by 0SourcePDFScholar
2026

Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation Models

AAAI 2026technical

Vision foundation models (VFMs) have demonstrated remarkable capabilities in learning universal visual representations. However, adapting these models to downstream tasks conventionally requires parameter updates, with even parameter-efficient fine-tuning methods necessitating the modification of th

Cited by 0SourcePDFScholar
2025

Domain Prompt Learning with Quaternion Networks (Extended Abstract)

IJCAI 2025

Foundational vision-language models (VLMs) like CLIP have revolutionized image recognition, but adapting them to specialized domains with limited data remains challenging. We propose Domain Prompt Learning with Quaternion Networks (DPLQ), which leverages domain-specific foundation models and quatern

Cited by 0SourcePDFScholar
2025

FATE: Feature-Adapted Parameter Tuning for Vision-Language Models

AAAI 2025technical

Following the recent popularity of vision language models, several attempts, e.g., parameter-efficient fine-tuning (PEFT), have been made to extend them to different downstream tasks. Previous PEFT works motivate their methods from the view of introducing new parameters for adaptation but still need…

Cited by 0SourcePDFScholar
2025

HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models

NeurIPS 2025oral

Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computational resources (e.g., thousands of GPUs) for training to achieve cross-modal alignment at multi-granularity levels. We arg…

Cited by 0SourceScholar
2025

Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning

ICLR 2025poster

Adapting pre-trained foundation models for various downstream tasks has been prevalent in artificial intelligence. Due to the vast number of tasks and high costs, adjusting all parameters becomes unfeasible. To mitigate this, several fine-tuning techniques have been developed to update the pre-train…

Cited by 1SourcePDFScholar
2025

Robust SAM: On the Adversarial Robustness of Vision Foundation Models

AAAI 2025technical

The Segment Anything Model (SAM) is a widely used vision foundation model with diverse applications, including image segmentation, detection, and tracking. Given SAM's wide applications, understanding its robustness against adversarial attacks is crucial for real-world deployment. However, research…

Cited by 1SourcePDFScholar
2025

Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space

CVPR 2025poster

CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing the text encoder preserves its powerful embeddings, recent studies show that fine-tuning both the text and image encoders jointly significantly enhances segmentation p…

2024

3D-Aware Face Editing via Warping-Guided Latent Direction Learning

CVPR 2024poster

3D facial editing a longstanding task in computer vision with broad applications is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness generalization and efficiency. To overcome these cha…

2024

Domain Prompt Learning with Quaternion Networks

CVPR 2024highlight

Prompt learning has emerged as an effective and data-efficient technique in large Vision-Language Models (VLMs). However when adapting VLMs to specialized domains such as remote sensing and medical imaging domain prompt learning remains underexplored. While large-scale domain-specific foundation mod…

Cited by 13SourcePDFScholar
2024

LERE: Learning-Based Low-Rank Matrix Recovery with Rank Estimation

AAAI 2024technical

A fundamental task in the realms of computer vision, Low-Rank Matrix Recovery (LRMR) focuses on the inherent low-rank structure precise recovery from incomplete data and/or corrupted measurements given that the rank is a known prior or accurately estimated. However, it remains challenging for exist…

2024

Parameter Efficient Fine-tuning via Cross Block Orchestration for Segment Anything Model

CVPR 2024poster

Parameter-efficient fine-tuning (PEFT) is an effective methodology to unleash the potential of large foundation models in novel scenarios with limited training data. In the computer vision community PEFT has shown effectiveness in image classification but little research has studied its ability for…

Cited by 11SourcePDFScholar
2024

SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space Reconstruction

AAAI 2024technical

Segment Anything Model (SAM) has received remarkable attention as it offers a powerful and versatile solution for object segmentation in images. However, fine-tuning SAM for downstream segmentation tasks under different scenarios remains a challenge, as the varied characteristics of different scenar…

Cited by 23SourcePDFScholar
2023

Efficient Robust Principal Component Analysis via Block Krylov Iteration and CUR Decomposition

CVPR 2023poster

Robust principal component analysis (RPCA) is widely studied in computer vision. Recently an adaptive rank estimate based RPCA has achieved top performance in low-level vision tasks without the prior rank, but both the rank estimate and RPCA optimization algorithm involve singular value decompositio…

Cited by 8SourcePDFScholar