← Search

Zihang Lin

5 accepted papers

2025

Style Nursing with Spatial and Semantic Guidance for Zero-Shot Traffic Scene Style Transfer

AAAI 2025technical

Recent advances in text-to-image diffusion models have shown an outstanding ability in zero-shot style transfer. However, existing methods often struggle to balance preserving the semantic content of the input image and faithfully transferring the target style in line with the edit prompt. Especiall…

Cited by 0SourcePDFScholar
2023

Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding

CVPR 2023poster

Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the la…

2023

Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding

CVPR 2023poster

Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex sem…

Cited by 24SourcePDFScholar
2021

Action-guided 3D Human Motion Prediction

NeurIPS 2021poster

The ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human mo…

Cited by 10SourcePDFScholar
2021

Predictive Feature Learning for Future Segmentation Prediction

ICCV 2021poster

Future segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation models. However, these segmentation features are learned to be local discriminative…

Cited by 20PDFScholar