← Search

Lipeng Ke

5 accepted papers

2026

ConsMSA: Semantic Distribution Consistency Learning for Multimodal Sentiment Analysis

ICML 2026poster

Multimodal sentiment analysis (MSA) aims to predict human sentiments by integrating signals from different modalities such as text, video, and audio. However, raw multimodal sequences often suffer from semantic inconsistencies--exhibiting redundancy or conflicts within and across modalities--which h…

Cited by 0SourceScholar
2026

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

ICLR 2026poster

We propose a method for audio-visual knowledge distillation. Existing methods typically distill from the latent embeddings or outputs. The former requires matching feature dimensions, if not the same architecture, between teacher and student models while the latter supports any teacher-student pairi…

Cited by 0SourceScholar
2022

Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action Recognition

AAAI 2022technical

Graph Convolutional Networks (GCNs) have been widely used to model the high-order dynamic dependencies for skeleton-based action recognition. Most existing approaches do not explicitly embed the high-order spatio-temporal importance to joints’ spatial connection topology and intensity, and they do n…

Cited by 63SourcePDFScholar
2018

Multi-Scale Structure-Aware Network for Human Pose Estimation

ECCV 2018poster

We develop a robust multi-scale structure-aware neural network for human pose estimation. This method improves the recent deep conv-deconv hourglass models with four key improvements: (1) multi-scale supervision to strengthen contextual feature learning in matching body keypoints by combining featur…

Cited by 382SourcePDFScholar