← Search

Zhenwei Shi

9 accepted papers

2026

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

AAAI 2026technical

Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus lead

Cited by 0SourcePDFScholar
2025

MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting

CVPR 2025poster

Deep learning approaches for marine fog detection and forecasting have outperformed traditional methods, demonstrating significant scientific and practical importance. However, the limited availability of open-source datasets remains a major challenge. Existing datasets, often focused on a single re…

2025

Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes

ICLR 2025poster

Recent advancements in human motion synthesis have focused on specific types of motions, such as human-scene interaction, locomotion or human-human interaction, however, there is a lack of a unified system capable of generating a diverse combination of motion types. In response, we introduce *Sitcom…

2025

Unified Multi-Agent Trajectory Modeling with Masked Trajectory Diffusion

ICCV 2025poster

Understanding movements in multi-agent scenarios is a fundamental problem in intelligent systems. Previous research assumes complete and synchronized observations. However, real-world partial observation caused by occlusions leads to inevitable model failure, which demands a unified framework for co…

2023

Zero-Shot Text-to-Parameter Translation for Game Character Auto-Creation

CVPR 2023poster

Recent popular Role-Playing Games (RPGs) saw the great success of character auto-creation systems. The bone-driven face model controlled by continuous parameters (like the position of bones) and discrete parameters (like the hairstyles) makes it possible for users to personalize and customize in-gam…

2020

Deep Adversarial Decomposition: A Unified Framework for Separating Superimposed Images

CVPR 2020poster

Separating individual image layers from a single mixed image has long been an important but challenging task. We propose a unified framework named "deep adversarial decomposition" for single superimposed image separation. Our method deals with both linear and non-linear mixtures under an adversarial…

Cited by 86PDFScholar
2019

Face-to-Parameter Translation for Game Character Auto-Creation

ICCV 2019poster

Character customization system is an important component in Role-Playing Games (RPGs), where players are allowed to edit the facial appearance of their in-game characters with their own preferences rather than using default templates. This paper proposes a method for automatically creating in-game c…

Cited by 75PDFScholar
2019

Generative Adversarial Training for Weakly Supervised Cloud Matting

ICCV 2019poster

The detection and removal of cloud in remote sensing images are essential for earth observation applications. Most previous methods consider cloud detection as a pixel-wise semantic segmentation process (cloud v.s. background), which inevitably leads to a category-ambiguity problem when dealing with…

Cited by 42PDFScholar