← Search

Yingda Lyu

10 accepted papers

2026

Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose Estimation

AAAI 2026technical

Video-based human pose estimation has vast applications such as action recognition, sports analytics, and crime detection. However, this task is challenging as it involves interpreting both spatial context and temporal dynamics to accurately localize human anatomical keypoints in video sequences. Cu

Cited by 0SourcePDFScholar
2026

Causality-Aligned Semantic Recovery for Incomplete Cross-Modal Retrieval

AAAI 2026technical

Incomplete cross-modal retrieval (ICMR) requires models to recover missing modalities and robustly align heterogeneous ones for effective retrieval. Existing methods, however, fall short in both aspects. They often rely on limited semantic cues, such as single samples or coarse category prototypes,

Cited by 0SourcePDFScholar
2026

Diffusion-Based Native Adversarial Synthesis for Enhanced Medical Segmentation Generalization

CVPR 2026

Diffusion models (DMs) can generate anatomically realistic medical images, offering a compelling route to improving generalization through synthetic augmentation. Yet high visual realism does not necessarily translate into improved downstream utility. This work addresses two key questions in diffusi

Cited by 0SourceScholar
2026

Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in Videos

AAAI 2026technical

Video-based human pose estimation aims to localize keypoints across frames, enabling robust analysis of human motion in applications such as sports, surveillance, and healthcare. However, existing methods rely solely on visual cues, limiting their robustness in complex scenes involving occlusion, mo

Cited by 0SourcePDFScholar
2026

VGD: Value-Guided Diffusion Toward High-Utility Medical Image Segmentation

AAAI 2026technical

Progress in medical image segmentation is fundamentally constrained by the scarcity of annotated data. While diffusion models offer a promising solution by generating high-fidelity image–mask pairs, their utility for downstream tasks remains underexplored. A key bottleneck lies in the misalignment

Cited by 0SourcePDFScholar
2025

Causal-Inspired Multitask Learning for Video-Based Human Pose Estimation

AAAI 2025technical

Video-based human pose estimation has long been a fundamental yet challenging problem in computer vision. Previous studies focus on spatio-temporal modeling through the enhancement of architecture design and optimization strategies. However, they overlook the causal relationships in the joints, lead…

Cited by 1SourcePDFScholar
2025

Enhancing Semantic Clarity: Discriminative and Fine-grained Information Mining for Remote Sensing Image-Text Retrieval

IJCAI 2025

Remote sensing image-text retrieval is a fundamental task in remote sensing multimodal analysis, promoting the alignment of visual and language representations. The mainstream approaches commonly focus on capturing shared semantic representations between visual and textual modalities. However, the i

Cited by 0SourcePDFScholar
2025

Skeleton-based Action Recognition with Non-linear Dependency Modeling and Hilbert-Schmidt Independence Criterion

AAAI 2025technical

Human skeleton-based action recognition has long been an indispensable aspect of artificial intelligence. Current state-of-the-art methods tend to consider only the dependencies between connected skeletal joints, limiting their ability to capture non-linear dependencies between physically distant jo…

2024

Rethinking Human Motion Prediction with Symplectic Integral

CVPR 2024poster

Long-term and accurate forecasting is the long-standing pursuit of the human motion prediction task. Existing methods typically suffer from dramatic degradation in prediction accuracy with the increasing prediction horizon. It comes down to two reasons:1? Insufficient numerical stability.Unforeseen…

Cited by 2SourcePDFScholar
2023

Action Recognition with Multi-stream Motion Modeling and Mutual Information Maximization

IJCAI 2023poster

Action recognition has long been a fundamental and intriguing problem in artificial intelligence. The task is challenging due to the high dimensionality nature of an action, as well as the subtle motion details to be considered. Current state-of-the-art approaches typically learn from articulated mo…