← Search

Wen-Li Wei

6 accepted papers

2025

Birds of a Feather: Learning to Retrieve Dance Poses From Music Via Ground-Truth Annotation Lifting

ICASSP 2025accepted

Learning to retrieve dance poses from music, a cross-modal retrieval task, has gained prominence in assisting choreographers in creating dances that harmonize with music. The recent predominant approach is to map music into 3D pose and shape space, and then match it with dance poses [1]. However, we…

Cited by 0SourceScholar
2024

Music-to-Dance Poses: Learning to Retrieve Dance Poses from Music

ICASSP 2024accepted

Choreography is an artful blend of technique and creativity, requiring the meticulous design of movement sequences in harmony with music. To support choreographers in this intricate task, this work proposes a "music-to-dance pose retrieval" system that uses music snippets to retrieve dance poses, pr…

Cited by 4SourceScholar
2022

Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation From Monocular Video

CVPR 2022poster

Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human mot…

Cited by 108PDFcodeScholar
2021

Positions, Channels, and Layers: Fully Generalized Non-Local Network for Singer Identification

AAAI 2021technical

Recently, a non-local (NL) operation has been designed as the central building block for deep-net models to capture long-range dependencies (Wang et al. 2018). Despite its excellent performance, it does not consider the interaction between positions across channels and layers, which is crucial in fi…

2017

Deep-net fusion to classify shots in concert videos

ICASSP 2017accepted

Varying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. W…

Cited by 0SourceScholar
2016

DEMV-matchmaker: Emotional temporal course representation and deep similarity matching for automatic music video generation

ICASSP 2016accepted

This paper presents a deep similarity matching-based emotion-oriented music video (MV) generation system, called DEMV-matchmaker, which utilizes an emotion-oriented deep similarity matching (EDSM) metric as a bridge to connect music and video. Specifically, we adopt an emotional temporal course mode…

Cited by 12SourceScholar