← Search

Zhipeng Fan

10 accepted papers

2026

UniT: Unified Multimodal Chain-of-Thought Test-time Scaling

CVPR 2026

Unified models can handle both multimodal understanding and generation within a single architecture, yet they typically operate in a single pass without iteratively refining their outputs. Many multimodal tasks, especially those involving complex spatial compositions, multiple interacting objects, o

Cited by 0SourceScholar
2024

MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers

CVPR 2024poster

Recent advances in generative AI have significantly enhanced image and video editing particularly in the context of text prompt control. State-of-the-art approaches predominantly rely on diffusion models to accomplish these tasks. However the computational demands of diffusion-based methods are subs…

Cited by 4SourcePDFScholar
2024

SM3: Self-supervised Multi-task Modeling with Multi-view 2D Images for Articulated Objects

ICRA 2024poster

Reconstructing real-world objects and estimating their movable joint structures are pivotal technologies within the field of robotics. Previous research has predominantly focused on supervised approaches, relying on annotated datasets to model articulated objects within limited categories. However,…

Cited by 1SourceScholar
2023

DiffPose: Toward More Reliable 3D Pose Estimation

CVPR 2023poster

Monocular 3D human pose estimation is quite challenging due to the inherent ambiguity and occlusion, which often lead to high uncertainty and indeterminacy. On the other hand, diffusion models have recently emerged as an effective tool for generating high-quality images from noise. Inspired by their…

2023

System-Status-Aware Adaptive Network for Online Streaming Video Understanding

CVPR 2023poster

Recent years have witnessed great progress in deep neural networks for real-time applications. However, most existing works do not explicitly consider the general case where the device's state and the available resources fluctuate over time, and none of them investigate or address the impact of vary…

2022

GradAuto: Energy-Oriented Attack on Dynamic Neural Networks

ECCV 2022poster

"Dynamic neural networks could adapt their structures or parameters based on different inputs. By reducing the computation redundancy for certain samples, it can greatly improve the computational efficiency without compromising the accuracy. In this paper, we investigate the robustness of dynamic ne…

2022

REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in Videos

AAAI 2022technical

Existing approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose va…

Cited by 13SourcePDFScholar