← Search

Yuang Zhang

9 accepted papers

2026

Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping

ICML 2026poster

Precise camera pose control is critical for video diffusion, yet maintaining geometric consistency remains a challenge. Existing methods that directly inject numerical camera parameters into the diffusion backbone often fail to bridge the gap between abstract coordinates and visual content, leading …

Cited by 1SourceScholar
2025

MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

ICML 2025poster

In recent years, while generative AI has advanced significantly in image generation, video generation continues to face challenges in controllability, length, and detail quality, which hinder its application. We present MimicMotion, a framework for generating high-quality human videos of arbitrary l…

2024

SmartCooper: Vehicular Collaborative Perception with Adaptive Fusion and Judger Mechanism

ICRA 2024poster

In recent years, autonomous driving has garnered significant attention due to its potential for improving road safety through collaborative perception among connected and autonomous vehicles (CAVs). However, time-varying channel variations in vehicular transmission environments demand dynamic alloca…

Cited by 5SourceScholar
2023

MOTRv2: Bootstrapping End-to-End Multi-Object Tracking by Pretrained Object Detectors

CVPR 2023poster

In this paper, we propose MOTRv2, a simple yet effective pipeline to bootstrap end-to-end multi-object tracking with a pretrained object detector. Existing end-to-end methods, e.g. MOTR and TrackFormer are inferior to their tracking-by-detection counterparts mainly due to their poor detection perfor…

2023

OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation

ICCV 2023poster

Referring video object segmentation (RVOS) aims at segmenting an object in a video following human instruction. Current state-of-the-art methods fall into an offline pattern, in which each clip independently interacts with text embedding for cross-modal understanding. They usually present that the o…

Cited by 55PDFcodeScholar
2022

MOTR: End-to-End Multiple-Object Tracking with TRansformer

ECCV 2022poster

"Temporal modeling of objects is a key challenge in multiple-object tracking (MOT). Existing methods track by associating detections through motion-based and appearance-based similarity heuristics. The post-processing nature of association prevents end-to-end exploitation of temporal variations in v…

2022

Progressive End-to-End Object Detection in Crowded Scenes

CVPR 2022poster

In this paper, we propose a new query-based detection framework for crowd detection. Previous query-based detectors suffer from two drawbacks: first, multiple predictions will be inferred for a single object, typically in crowded scenes; second, the performance saturates as the depth of the decoding…

Cited by 85PDFcodeScholar