← Search

Nan Zhou

10 accepted papers

2026

CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models

CVPR 2026

Multimodal large language models (MLLMs) achieve remarkable progress in cross-modal perception and reasoning, yet a fundamental question remains unresolved: should the vision encoder be fine-tuned or frozen? Despite the success of models such as LLaVA and Qwen-VL, inconsistent design choices and het

Cited by 0SourcecodeScholar
2026

CrossCut: Cross-Patch Aware Interactive Segmentation for Remote Sensing Images

AAAI 2026technical

Interactive segmentation aims to delineate a user-specified target in an image by leveraging positive and negative clicks. While effective on natural images, existing methods often fail in remote sensing scenarios, where satellite imagery is characterized by ultra-high resolution, sparse object dist

Cited by 0SourcePDFScholar
2026

MS-CRL: Multi-Scale Global Path Planning with Progressive Curriculum Reinforcement Learning

ICRA 2026poster

Global path planning provides high-level guidance for autonomous navigation, supplying reference paths for downstream navigation and control modules. Deep Reinforcement Learning (DRL) has shown strong potential in this domain, but existing methods struggle with multi-scale map inputs. This limitatio…

Cited by 0Scholar
2025

ForCenNet: Foreground-Centric Network for Document Image Rectification

ICCV 2025poster

Document image rectification aims to eliminate geometric deformation in photographed documents to facilitate text recognition. However, existing methods often neglect the significance of foreground elements, which provide essential geometric references and layout information for document image corre…

2025

Implicit Modeling for Transferability Estimation of Vision Foundation Models

NeurIPS 2025poster

Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitates deployment and advances the pre-training and fine-tuning paradigm. However, existing methods often struggle to accurat…

Cited by 0SourceScholar
2025

Progressive Parameter Efficient Transfer Learning for Semantic Segmentation

ICLR 2025poster

Parameter Efficient Transfer Learning (PETL) excels in downstream classification fine-tuning with minimal computational overhead, demonstrating its potential within the pre-train and fine-tune paradigm. However, recent PETL methods consistently struggle when fine-tuning for semantic segmentation tas…

2024

Crowd-SAM:SAM as a smart annotator for object detection in crowded scenes

ECCV 2024poster

"Object detection is an important task that finds its application in a wide range of scenarios. Generally, it requires extensive labels for training, which is quite time-consuming, especially in crowded scenes. Recently, Segment Anything Model (SAM) has emerged as a powerful zero-shot segmenter, off…

2023

DR-Tune: Improving Fine-tuning of Pretrained Visual Models by Distribution Regularization with Semantic Calibration

ICCV 2023poster

The visual models pretrained on large-scale benchmarks encode general knowledge and prove effective in building more powerful representations for downstream tasks. Most existing approaches follow the fine-tuning paradigm, either by initializing or regularizing the downstream model based on the pretr…

Cited by 7PDFcodeScholar
2022

Motion Sensitive Contrastive Learning for Self-Supervised Video Representation

ECCV 2022poster

"Contrastive learning has shown great potential in video representation learning. However, existing approaches fail to sufficiently exploit short-term motion dynamics, which are crucial to various down-stream video understanding tasks. In this paper, we propose Motion Sensitive Contrastive Learning…

Cited by 20SourcePDFScholar