← Search

Weizhan Zhang

13 accepted papers

2026

AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References

CVPR 2026

Identity-preserving video generation offers powerful tools for creative expression, allowing users to customize videos featuring their beloved characters. However, prevailing methods are typically designed and optimized for a single identity reference. This underlying assumption restricts creative f

Cited by 0SourcecodeScholar
2026

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

AAAI 2026technical

Recently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that fine-tune CLIP for segmentation on limited seen categories often lead to overfitting and degrade the pretrained vision-l

Cited by 0SourcePDFScholar
2026

MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation

ICLR 2026poster

Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Existing methods struggle to generalize to multi-humanoid scenarios, which involve diverse humanoid forms, complex interact…

Cited by 0SourcecodeScholar
2026

Recovering Coherent Affective Patterns: Addressing Modality Missing in Multimodal Sentiment Analysis

AAAI 2026technical

Multimodal sentiment analysis (MSA) seeks to decode human emotions by integrating heterogeneous modalities. However, real-world scenarios often involve missing or misaligned data due to sensor failures or transmission errors, leading to disrupted temporal dynamics and degraded cross-modal correlatio

Cited by 0SourcePDFScholar
2026

SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images

CVPR 2026

Effectively grounding complex language to pixels in remote sensing (RS) images is a critical challenge for applications like disaster response and environmental monitoring. Current models can parse simple, single-target commands but fail when presented with complex geospatial scenarios, e.g., segmen

Cited by 0SourcecodeScholar
2026

ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks

CVPR 2026

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a passive perception paradigm, suffering from increased redundancy when obtaining f

Cited by 0SourcecodeScholar
2025

DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial Editability

ICCV 2025poster

Recent advances in text-to-image generation have driven interest in generating personalized human images that depict specific identities from reference images. Although existing methods achieve high-fidelity identity preservation, they are generally limited to single-ID scenarios and offer insuffici…

Cited by 0SourcePDFScholar
2025

EchoShot: Multi-Shot Portrait Video Generation

NeurIPS 2025poster

Video diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge for multiple shots with identity consistency and…

Cited by 0SourcecodeScholar
2025

InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic Perspective

ICML 2025spotlight

The Segment Anything Model (SAM), a vision foundation model, exhibits impressive zero-shot capabilities in general tasks but struggles in specialized domains. Parameter-efficient fine-tuning (PEFT) is a promising approach to unleash the potential of SAM in novel scenarios. However, existing PEFT met…

2025

SpotActor: Training-Free Layout-Controlled Consistent Image Generation

AAAI 2025technical

Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of eac…

Cited by 2SourcePDFScholar
2024

Accelerating Non-Maximum Suppression: A Graph Theory Perspective

NeurIPS 2024poster

Non-maximum suppression (NMS) is an indispensable post-processing step in object detection. With the continuous optimization of network models, NMS has become the ``last mile'' to enhance the efficiency of object detection. This paper systematically analyzes NMS from a graph theory perspective for t…

2024

OneActor: Consistent Subject Generation via Cluster-Conditioned Guidance

NeurIPS 2024poster

Text-to-image diffusion models benefit artists with high-quality image generation. Yet their stochastic nature hinders artists from creating consistent images of the same subject. Existing methods try to tackle this challenge and generate consistent content in various ways. However, they either depe…

2024

SAUI: Scale-Aware Unseen Imagineer for Zero-Shot Object Detection

AAAI 2024technical

Zero-shot object detection (ZSD) aims to localize and classify unseen objects without access to their training annotations. As a prevailing solution to ZSD, generation-based methods synthesize unseen visual features by taking seen features as reference and class semantic embeddings as guideline. Alt…

Cited by 4SourcePDFScholar