← Search

Dongdong Yu

14 accepted papers

2025

LaVin-DiT: Large Vision Diffusion Transformer

CVPR 2025poster

This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative framework. Unlike existing large vision models directly adapted from natural language processing architectures, which rely…

2025

TrackGo: A Flexible and Efficient Method for Controllable Video Generation

AAAI 2025technical

Recent years have seen substantial progress in diffusion-based controllable video generation. However, achieving precise control in complex scenarios, including fine-grained object parts, sophisticated motion trajectories, and coherent background movement, remains a challenge. In this paper, we in…

Cited by 11SourcePDFScholar
2023

Box-Level Active Detection

CVPR 2023highlight

Active learning selects informative samples for annotation within budget, which has proven efficient recently on object detection. However, the widely used active detection benchmarks conduct image-level evaluation, which is unrealistic in human workload estimation and biased towards crowded images.…

2022

AdaptivePose: Human Parts as Adaptive Points

AAAI 2022technical

Multi-person pose estimation methods generally follow top-down and bottom-up paradigms, both of which can be considered as two-stage approaches thus leading to the high computation cost and low efficiency. Towards a compact and efficient pipeline for multi-person pose estimation task, in this paper,…

Cited by 26SourcePDFScholar
2022

ByteTrack: Multi-Object Tracking by Associating Every Detection Box

ECCV 2022poster

"Multi-object tracking (MOT) aims at estimating bounding boxes and identities of objects in videos. Most methods obtain identities by associating detection boxes whose scores are higher than a threshold. The objects with low detection scores, e.g. occluded objects, are simply thrown away, which brin…

2022

Learning Quality-Aware Representation for Multi-Person Pose Regression

AAAI 2022technical

Off-the-shelf single-stage multi-person pose regression methods generally leverage the instance score (i.e., confidence of the instance localization) to indicate the pose quality for selecting the pose candidates. We consider that there are two gaps involved in existing paradigm: 1) The instance sco…

Cited by 17SourcePDFScholar
2022

QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query

NeurIPS 2022accept

We propose a sparse end-to-end multi-person pose regression framework, termed QueryPose, which can directly predict multi-person keypoint sequences from the input image. The existing end-to-end methods rely on dense representations to preserve the spatial detail and structure for precise keypoint lo…

2021

F2Net: Learning to Focus on the Foreground for Unsupervised Video Object Segmentation

AAAI 2021technical

Although deep learning based methods have achieved great progress in unsupervised video object segmentation, difficult scenarios (e.g., visual similarity, occlusions, and appearance changing) are still no well-handled. To alleviate these issues, we propose a novel Focus on Foreground Network (F2Net…

Cited by 51SourcePDFScholar
2021

Mining Contextual Information Beyond Image for Semantic Segmentation

ICCV 2021poster

This paper studies the context aggregation problem in semantic image segmentation. The existing researches focus on improving the pixel representations by aggregating the contextual information within individual images. Though impressive, these methods neglect the significance of the representations…

Cited by 105PDFcodeScholar
2021

Weakly Supervised Person Search With Region Siamese Networks

ICCV 2021poster

Supervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to collect, especially for person identities. A natural question is whether a good person search model can be trained without th…

Cited by 30PDFScholar
2019

Multi-Person Pose Estimation With Enhanced Channel-Wise and Spatial Information

CVPR 2019poster

Multi-person pose estimation is an important but challenging problem in computer vision. Although current approaches have achieved significant progress by fusing the multi-scale feature maps, they pay little attention to enhancing the channel-wise and spatial information of the feature maps. In this…

Cited by 189PDFScholar