← Search

Dongli Xu

12 accepted papers

2026

Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring

ICLR 2026poster

Accurate cell instance segmentation is foundational for digital pathology analysis. Existing methods based on contour detection and distance mapping still face significant challenges in processing complex and dense cellular regions. Graph coloring-based methods provide a new paradigm for this task,…

Cited by 0SourcecodeScholar
2026

PHYSICS-AWARE NOVEL-VIEW ACOUSTIC SYNTHESIS WITH VISION-LANGUAGE PRIORS AND 3D ACOUSTIC ENVIRONMENT MODELING

ICASSP 2026poster

Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on single-view or panoramic inputs improve spatial fidelity but fail t…

Cited by 0SourcePDFScholar
2026

SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model

ICLR 2026poster

Autoregressive (AR) models have emerged as powerful tools for image generation by modeling images as sequences of discrete tokens. While Classifier-Free Guidance (CFG) has been adopted to improve conditional generation, its application in AR models faces two key issues: guidance diminishing, where t…

Cited by 1SourceScholar
2026

SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

CVPR 2026

Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressi

Cited by 0SourcecodeScholar
2025

Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation

ICCV 2025poster

Automatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that naturally follows music remains a challenge, as existing methods lack proper beat alignment and exhibit unnatural motion dynam…

Cited by 0SourcePDFScholar
2025

Band Prompting Aided SAR and Multi-Spectral Data Fusion Framework for Local Climate Zone Classification

ICASSP 2025accepted

Local climate zone (LCZ) classification is of great value for understanding the complex interactions between urban development and local climate. Recent studies have increasingly focused on the fusion of synthetic aperture radar (SAR) and multi-spectral data to improve LCZ classification performance…

Cited by 0SourceScholar
2025

Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization

NeurIPS 2025poster

The generalization ability of deep learning has been extensively studied in supervised settings, yet it remains less explored in unsupervised scenarios. Recently, the Unsupervised Domain Generalization (UDG) task has been proposed to enhance the generalization of models trained with prevalent unsupe…

Cited by 0SourceScholar
2025

Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning

ICCV 2025poster

3D medical image self-supervised learning (mSSL) holds great promise for medical analysis. Effectively supporting broader applications requires considering anatomical structure variations in location, scale, and morphology, which are crucial for capturing meaningful distinctions. However, previous m…

2024

FPN with GMM Based Feature Enhancement Strategy for Object Detection in Remote Sensing Images

ICASSP 2024accepted

In the realm of object detection, the age-old challenge of accommodating large variations in target scales, particularly in the intricate domain of remote sensing imagery, has long perplexed computer vision aficionados. Feature Pyramid Network (FPN) family, a widely-used stalwart, strives to tame th…

Cited by 0SourceScholar
2024

FastDrag: Manipulate Anything in One Step

NeurIPS 2024poster

Drag-based image editing using generative models provides precise control over image contents, enabling users to manipulate anything in an image with a few clicks. However, prevailing methods typically adopt $n$-step iterations for latent semantic optimization to achieve drag-based image editing, wh…