← Search

Zelong Sun

6 accepted papers

2026

FineNav: A Versatile Framework Enhancing Ground Robot Navigation in Unstructured Environment

ICRA 2026poster

Autonomous navigation of ground robots in unstructured 3D environments remains a fundamental challenge, as it requires accommodating dynamic obstacles, non-planar ground, and multi-story structures within a unified framework. In this paper, we propose a versatile navigation framework named FineNav. …

Cited by 0Scholar
2026

PortraitRL: Reinforcement Learning for Personalized Portrait Pose Transfer with Multi-Objective Reward Modeling

ICML 2026poster

Portrait pose transfer (PPT) requires generative models to preserve fine-grained identity details while following complex pose and layout modification instructions. Existing methods often struggle with extensive data annotation requirements or employ optimization objectives that are suboptimal for a…

Cited by 0SourceScholar
2026

Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits

CVPR 2026

As social media platforms proliferate, users increasingly demand intuitive ways to create diverse, high-quality portrait collections. In this work, we introduce Portrait Collection Generation (PCG), a novel task that generates coherent portrait collections by editing a reference portrait image throu

Cited by 0SourceScholar
2025

Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval

AAAI 2025technical

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion stra…

Cited by 2SourcePDFScholar
2024

Image Retrieval with Composed Query by Multi-Scale Multi-Modal Fusion

ICASSP 2024accepted

Image retrieval with composed query (IR-CQ) is a challenging task since it aims to retrieve the target image according to a hybrid-modality query which consists of a reference image and a text modifier. Previous approaches mainly focus on designing various multi-modal fusion modules to fuse the hybr…

Cited by 0SourceScholar