← Search

Zhanjie Zhang

21 accepted papers

2026

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

CVPR 2026

Pansharpening aims to generate high-resolution multi-spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. However, most existing methods are evaluated under limited, low-resolution settings, limiting their generalization to real-wo

Cited by 0SourcecodeScholar
2026

InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation

CVPR 2026

E-commerce product poster generation aims to automatically synthesize a single image that effectively conveys product information by presenting a subject, text, and a designed style. Recent diffusion models with fine-grained and efficient controllability have advanced product poster synthesis, yet t

Cited by 0SourceScholar
2026

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

ICML 2026poster

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 7K+ timestamped questions for diagnosing User-centric Continual Spatial intelligence in egocentric video streams. UCS-Bench targets a new problem that emphasizes dynamic spatial reasoning, long-term memory, …

Cited by 0SourceScholar
2026

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

AAAI 2026technical

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural g

Cited by 0SourcePDFScholar
2026

RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers

AAAI 2026technical

The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource

Cited by 0SourcePDFScholar
2026

SBSDM: A Style-aware Bidirectional Stream Diffusion Model for CT-to-PET Synthesis

IJCAI 2026

CT-to-PET synthesis aims to synthesize PET images from the widely available and lower-cost CT scans to address the high cost and additional radiation exposure associated with PET scanning. However, CT-to-PET synthesis faces two key challenges due to the sequential correlation of volumetric imaging:

Cited by 0Scholar
2026

Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interfaces

ICRA 2026poster

Brain–Machine Interfaces (BMIs) provide a direct communication pathway between the brain and external devices, enabling humans to control assistive and robotic technologies, with potential applications in rehabilitation, human motor augmentation, and human-centered robotics. However, due to neural d…

Cited by 0Scholar
2025

DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label Supervision

AAAI 2025technical

Self-supervised stereo matching has drawn attention due to its ability to estimate disparity without needing ground-truth data. However, existing self-supervised stereo matching methods heavily rely on the photo-metric consistency assumption, which is vulnerable to natural disturbances, resulting in…

Cited by 0SourcePDFScholar
2025

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

IJCAI 2025

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations wi

Cited by 0SourcePDFScholar
2025

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

ICCV 2025poster

Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and inter-frame interact…

Cited by 0SourcePDFScholar
2025

Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts

ICCV 2025poster

Recently, learning-based stereo matching networks have advanced significantly.However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets.Leveraging Vision Foundation Models (VFMs)…

2025

WISA: World simulator assistant for physics-aware text-to-video generation

NeurIPS 2025spotlight

Recent advances in text-to-video (T2V) generation, exemplified by models such as Sora and Kling, have demonstrated strong potential for constructing world simulators. However, existing T2V models still struggle to understand abstract physical principles and to generate videos that faithfully obey ph…

Cited by 0SourcecodeScholar
2024

3DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint Videos

CVPR 2024highlight

Constructing photo-realistic Free-Viewpoint Videos (FVVs) of dynamic scenes from multi-view videos remains a challenging endeavor. Despite the remarkable advancements achieved by current neural rendering techniques these methods generally require complete video sequences for offline training and are…

2024

ArtBank: Artistic Style Transfer with Pre-trained Diffusion Model and Implicit Style Prompt Bank

AAAI 2024technical

Artistic style transfer aims to repaint the content image with the learned artistic style. Existing artistic style transfer methods can be divided into two categories: small model-based approaches and pre-trained large-scale model-based approaches. Small model-based approaches can preserve the conte…

2024

Rethinking Diffusion Model for Multi-Contrast MRI Super-Resolution

CVPR 2024poster

Recently diffusion models (DM) have been applied in magnetic resonance imaging (MRI) super-resolution (SR) reconstruction exhibiting impressive performance especially with regard to detailed reconstruction. However the current DM-based SR reconstruction methods still face the following issues: (1) T…

2024

Towards Highly Realistic Artistic Style Transfer via Stable Diffusion with Step-aware and Layer-aware Prompt

IJCAI 2024poster

Artistic style transfer aims to transfer the learned artistic style onto an arbitrary content image, generating artistic stylized images. Existing generative adversarial network-based methods fail to generate highly realistic stylized images and always introduce obvious artifacts and disharmonious p…

2023

Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive Learning

AAAI 2023technical

This paper presents a new adversarial training framework for image inpainting with segmentation confusion adversarial training (SCAT) and contrastive learning. SCAT plays an adversarial game between an inpainting generator and a segmentation network, which provides pixel-level local training signals…

2023

Rethinking Multi-Contrast MRI Super-Resolution: Rectangle-Window Cross-Attention Transformer and Arbitrary-Scale Upsampling

ICCV 2023poster

Recently, several methods have explored the potential of multi-contrast magnetic resonance imaging (MRI) super-resolution (SR) and obtain results superior to single-contrast SR methods. However, existing approaches still have two shortcomings: (1) They can only address fixed integer upsampling scale…

Cited by 22PDFcodeScholar
2023

TeSTNeRF: Text-Driven 3D Style Transfer via Cross-Modal Learning

IJCAI 2023poster

Text-driven 3D style transfer aims at stylizing a scene according to the text and generating arbitrary novel views with consistency. Simply combining image/video style transfer methods and novel view synthesis methods results in flickering when changing viewpoints, while existing 3D style transfer m…

Cited by 16SourcePDFScholar
2023

VGOS: Voxel Grid Optimization for View Synthesis from Sparse Inputs

IJCAI 2023poster

Neural Radiance Fields (NeRF) has shown great success in novel view synthesis due to its state-of-the-art quality and flexibility. However, NeRF requires dense input views (tens to hundreds) and a long training time (hours to days) for a single scene to generate high-fidelity images. Although using…