← Search

Mengyuan Liu

39 accepted papers

2026

MP1: MeanFlow Tames Policy Learning in 1-step for Robotic Manipulation

AAAI 2026technical

In robot manipulation, robot learning has become a prevailing approach. However, generative models within this field face a fundamental trade-off between the slow, iterative sampling of diffusion models and the architectural constraints of faster Flow-based methods, which often rely on explicit cons

Cited by 0SourcePDFScholar
2026

Masked Clustering Prediction for Unsupervised Point Cloud Pre-training

AAAI 2026technical

Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We pr

Cited by 0SourcePDFScholar
2026

Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image Restoration

ICLR 2026poster

All-in-one image restoration (IR) aims to recover high-quality images from diverse degradations, which in real-world settings are often mixed and unknown. Unlike single-task IR, this problem requires a model to approximate a family of heterogeneous inverse functions, making it fundamentally more cha…

Cited by 0SourceScholar
2026

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation

CVPR 2026

Human motion analysis tasks, such as temporal 3D pose estimation, motion prediction, and motion in-betweening, play an essential role in computer vision. However, current paradigms suffer from severe fragmentation. First, the field is split between "perception" models that understand motion from vid

Cited by 0SourcecodeScholar
2026

Universal Skeleton Understanding via Differentiable Rendering and MLLMs

ICML 2026poster

Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet remain confined to their native modalities and cannot directly process structured, non-visual data such as human skeletons. Existing methods either compress skeleton dynamics into lossy feature vectors for text al…

Cited by 0SourceScholar
2025

Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment

AAAI 2025technical

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it…

2025

Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking

CVPR 2025poster

Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New method…

2025

Recognizing Actions from Robotic View for Natural Human-Robot Interaction

ICCV 2025poster

Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationary. This setup is more flexible and practical than conventional human action recognition tasks. However, existing benchma…

2025

SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose Estimation

AAAI 2025technical

Recently, transformer-based methods have been introduced to estimate 3D human pose from multiple views by aggregating the spatial-temporal information of human joints to achieve the lifting of 2D to 3D. However, previous approaches cannot model the inter-frame correspondence of each view's joint ind…

2025

TCNet: A Temporally Consistent Network for Self-supervised Monocular Depth Estimation

IROS 2025

Despite significant advances in self-supervised monocular depth estimation methods, achieving temporally consistent and accurate depth maps from frame sequences remains a formidable challenge. Existing approaches often estimate depth maps for individual frames in isolation, neglecting the rich geome

Cited by 0SourceScholar
2025

TCPFormer: Learning Temporal Correlation with Implicit Pose Proxy for 3D Human Pose Estimation

AAAI 2025technical

Recent multi-frame lifting methods have dominated the 3D human pose estimation. However, previous methods ignore the intricate dependence within the 2D pose sequence and learn single temporal correlation. To alleviate this limitation, we propose TCPFormer, which leverages an implicit pose proxy as a…

2024

CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition

NeurIPS 2024poster

Skeleton-based multi-entity action recognition is a challenging task aiming to identify interactive actions or group activities involving multiple diverse entities. Existing models for individuals often fall short in this task due to the inherent distribution discrepancies among entity skeletons, le…

2024

Denoising Diffusion Probabilistic Models for Action-Conditioned 3D Motion Generation

ICASSP 2024accepted

Diffusion-based generative models have proven to be highly effective in various domains of synthesis. In this work, we propose a conditional paradigm utilizing the denoising diffusion probabilistic model (DDPM) to address the challenge of realistic and diverse action-conditioned 3D skeleton-based mo…

Cited by 0SourceScholar
2024

Diffusion-Based Pose Refinement and Multi-Hypothesis Generation for 3D Human Pose Estimation

ICASSP 2024accepted

Previous probabilistic models for 3D Human Pose Estimation (3DHPE) aimed to enhance pose accuracy by generating multiple hypotheses. However, most of the hypotheses generated deviate substantially from the true pose. Compared to deterministic models, the excessive uncertainty in probabilistic models…

Cited by 0SourceScholar
2024

Expressive Forecasting of 3D Whole-Body Human Motions

AAAI 2024technical

Human motion forecasting, with the goal of estimating future human behavior over a period of time, is a fundamental task in many real-world applications. However, existing works typically concentrate on foretelling the major joints of the human body without considering the delicate movements of the…

2024

GCNext: Towards the Unity of Graph Convolutions for Human Motion Prediction

AAAI 2024technical

The past few years has witnessed the dominance of Graph Convolutional Networks (GCNs) over human motion prediction. Various styles of graph convolutions have been proposed, with each one meticulously designed and incorporated into a carefully-crafted network architecture. This paper breaks the limit…

2024

Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation

CVPR 2024highlight

Transformers have been successfully applied in the field of video-based 3D human pose estimation. However the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper we present a plug-and-play pruning-and-recovering framew…

2024

Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking

ICML 2024poster

Harnessing transformer-based models, visual tracking has made substantial strides. However, the sluggish performance of current trackers limits their practicality on devices with constrained computational capabilities, especially for real-time unmanned aerial vehicle (UAV) tracking. Addressing this…

2024

Sharing Key Semantics in Transformer Makes Efficient Image Restoration

NeurIPS 2024poster

Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergence of Vision Transformers (ViTs) has further propelled these advancements. When computing, the self-attention mechanism,…

2024

VG4D: Vision-Language Model Goes 4D Video Recognition

ICRA 2024poster

Understanding the real world through point cloud video is a crucial aspect of robotics and autonomous driving systems. However, prevailing methods for 4D point cloud recognition have limitations due to sensor resolution, which leads to a lack of detailed information. Recent advances have shown that…

Cited by 9SourcecodeScholar
2023

A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation

NeurIPS 2023poster

The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation, intractable computation and the non-causal problem. This can be…

2023

Body Prior Guided Graph Convolutional Neural Network for Skeleton-Based Action Recognition

ICASSP 2023accepted

Graph Convolutional Network (GCN) has achieved high success in the skeleton-based human action recognition task by modeling the human skeleton as a graph. However, it remains a problem for GCN-based methods to learn distinctive action features from a limited number of training samples. Via taking fu…

Cited by 0SourceScholar
2023

Explore In-Context Learning for 3D Point Cloud Understanding

NeurIPS 2023spotlight

With the rise of large-scale models trained on broad data, in-context learning has become a new learning paradigm that has demonstrated significant potential in natural language processing and computer vision tasks. Meanwhile, in-context learning is still largely unexplored in the 3D point cloud dom…

2023

Interactive Spatiotemporal Token Attention Network for Skeleton-Based General Interactive Action Recognition

IROS 2023poster

Recognizing interactive action plays an important role in human-robot interaction and collaboration. Previous methods use late fusion and co-attention mechanism to capture interactive relations, which have limited learning capability or inefficiency to adapt to more interacting entities. With assump…

Cited by 26SourcecodeScholar
2023

Multi-Stream Facial Adaptive Network for Expression Recognition from a Single Image

ICASSP 2023accepted

Facial expression recognition from a single image has potential applications in fields including human-computer interaction and medical diagnosis. Most recent methods use deep neural networks to directly learn from a roughly cropped facial image which is usually detected from a whole image by face d…

Cited by 0SourceScholar
2023

Novel Motion Patterns Matter for Practical Skeleton-Based Action Recognition

AAAI 2023technical

Most skeleton-based action recognition methods assume that the same type of action samples in the training set and the test set share similar motion patterns. However, action samples in real scenarios usually contain novel motion patterns which are not involved in the training set. As it is laboriou…

Cited by 32SourcePDFScholar
2023

Part Aware Contrastive Learning for Self-Supervised Action Recognition

IJCAI 2023poster

In recent years, remarkable results have been achieved in self-supervised action recognition using skeleton sequences with contrastive learning. It has been observed that the semantic distinction of human action features is often represented by local body parts, such as legs or hands, which are adva…

2023

PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation

CVPR 2023poster

Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relations of human joints in each video frame and human dynamics across frames with cascaded transformer layers and has achieved…

2023

RenderIH: A Large-Scale Synthetic Dataset for 3D Interacting Hand Pose Estimation

ICCV 2023poster

The current interacting hand (IH) datasets are relatively simplistic in terms of background and texture, with hand joints being annotated by a machine annotator, which may result in inaccuracies, and the diversity of pose distribution is limited. However, the variability of background, pose distribu…

Cited by 19PDFcodeScholar
2022

Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action Recognition

AAAI 2022technical

In recent years, self-supervised representation learning for skeleton-based action recognition has been developed with the advance of contrastive learning methods. The existing contrastive learning methods use normal augmentations to construct similar positive samples, which limits the ability to ex…

2020

GFNet: A Lightweight Group Frame Network for Efficient Human Action Recognition

ICASSP 2020accepted

Human action recognition aims at assigning an action label to a well-segmented video. Recent work using two-stream or 3D convolutional neural networks achieves high recognition rates at the cost of huge computation complexity, memory footprint, and parameters. In this paper, we propose a lightweight…

Cited by 0SourceScholar
2018

Deformable Pose Traversal Convolution for 3D Action and Gesture Recognition

ECCV 2018poster

The representation of 3D pose plays a critical role for 3D body action and hand gesture recognition. Rather than directly representing the 3D pose using its joint locations, in this paper, we propose Deformable Pose Traversal Convolution which applies one-dimensional convolution to traverse the 3D p…

Cited by 87SourcePDFScholar
2018

Learning Explicit Shape and Motion Evolution Maps for Skeleton-Based Human Action Recognition

ICASSP 2018accepted

Human action recognition based on skeleton sequences has wide applications in human-computer interaction and intelligent surveillance. Although previous methods have successfully applied Long Short-Term Memory(LSTM) networks to model shape evolution of human actions, it still remains a problem to ef…

Cited by 0SourceScholar
2017

Human action recognition using Adaptive Hierarchical Depth Motion Maps and Gabor filter

ICASSP 2017accepted

Depth motion maps (DMMs) have shown effectiveness in human action recognition, however, they lose the temporal information and suffer from intra-class variations caused by action speed variations. To address these challenges, we propose a novel method for human action recognition. Firstly, Adaptive…

Cited by 0SourceScholar