← Search

Dongkai Wang

10 accepted papers

2026

PMDformer: Patch-Mean Decoupling Transformer for Long-term Forecasting

ICLR 2026poster

Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-based strategies to capture long-range dependencies, but accurately modeling shape similarities across patches and variables…

Cited by 0SourceScholar
2026

Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition

ICLR 2026poster

Zero-shot Human-object interaction (HOI) detection aims to locate humans and objects in images and recognize their interactions. While advances in open-vocabulary object detection provide promising solutions for object localization, interaction recognition (IR) remains challenging due to the combina…

Cited by 0SourcecodeScholar
2025

Generalizable Object Keypoint Localization from Generative Priors

CVPR 2025poster

Generalizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data cannot provide generalizable shape and semantic cues, leading to inferior performan…

Cited by 0SourcePDFScholar
2025

InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions

NeurIPS 2025spotlight

In multimodal representation learning, synergistic interactions between modalities not only provide complementary information but also create unique outcomes through specific interaction patterns that no single modality could achieve alone. Existing methods may struggle to effectively capture the fu…

Cited by 0SourcecodeScholar
2024

LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model

CVPR 2024highlight

The capacity of existing human keypoint localization models is limited by keypoint priors provided by the training data. To alleviate this restriction and pursue more general model this work studies keypoint localization from a different perspective by reasoning locations based on keypiont clues in…

2021

Robust Pose Estimation in Crowded Scenes with Direct Pose-Level Inference

NeurIPS 2021poster

Multi-person pose estimation in crowded scenes is challenging because overlapping and occlusions make it difficult to detect person bounding boxes and infer pose cues from individual keypoints. To address those issues, this paper proposes a direct pose-level inference strategy that is free of boundi…