← Search

Kai Su

4 accepted papers

2026

BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration

ICLR 2026poster

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in subject-consistent video generation due to an inherent difficulty in pa…

Cited by 0SourceScholar
2022

QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query

NeurIPS 2022accept

We propose a sparse end-to-end multi-person pose regression framework, termed QueryPose, which can directly predict multi-person keypoint sequences from the input image. The existing end-to-end methods rely on dense representations to preserve the spatial detail and structure for precise keypoint lo…

2021

Weakly Supervised Person Search With Region Siamese Networks

ICCV 2021poster

Supervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to collect, especially for person identities. A natural question is whether a good person search model can be trained without th…

Cited by 30PDFScholar
2019

Multi-Person Pose Estimation With Enhanced Channel-Wise and Spatial Information

CVPR 2019poster

Multi-person pose estimation is an important but challenging problem in computer vision. Although current approaches have achieved significant progress by fusing the multi-scale feature maps, they pay little attention to enhancing the channel-wise and spatial information of the feature maps. In this…

Cited by 189PDFScholar