← Search

Zhidong Deng

20 accepted papers

2026

From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper

ICLR 2026poster

In multi-view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a specific moment. While providing accurate spatial information, this traditional approach often overlooks the rich temporal dependencies between adjacent…

Cited by 0SourcecodeScholar
2025

Exploring Timeline Control for Facial Motion Generation

CVPR 2025poster

This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arran…

Cited by 0SourcePDFScholar
2025

LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models

ICASSP 2025accepted

Recent advances in large vision-language models (LVLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of LVLMs. In this pa…

Cited by 0SourceScholar
2025

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

ICCV 2025poster

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstru…

Cited by 0SourcePDFScholar
2025

PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection

ICLR 2025poster

Single point supervised oriented object detection has gained attention and made initial progress within the community. Diverse from those approaches relying on one-shot samples or powerful pretrained models (e.g. SAM), PointOBB has shown promise due to its prior-free feature. In this paper, we propo…

Cited by 22SourcePDFScholar
2024

A Novel Contrastive Diffusion Graph Convolutional Network for Few-Shot Skeleton-Based Action Recognition

ICASSP 2024accepted

Existing skeleton spatial-temporal models tend to deteriorate the positional distinguishability of skeleton joints and lead to inaccurate spatial matching and poor interpretability. This paper proposes a novel contrastive diffusion graph convolutional network (CD-GCN) for few-shot action recognition…

Cited by 0SourceScholar
2024

Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition

ICRA 2024poster

This paper presents Action-SGFA, a novel action feature alignment approach to learn unified joint embeddings across four action modalities incorporating scene graph (SG) comprehension. A new training paradigm for Action-SGFA is also devised to improve pre-trained VL models using datasets with SG ann…

Cited by 3SourceScholar
2024

Open-Vocabulary Skeleton Action Recognition with Diffusion Graph Convolutional Network and Pre-Trained Vision-Language Models

ICASSP 2024accepted

This study explores unsupervised open-vocabulary skeleton action recognition, aiming at addressing inaccurate spatial matching and poor interpretability of existing GCN models. We present Skeleton-DGCFA, an approach to make feature alignment (FA) of skeleton with image modalities based on a large pr…

Cited by 0SourceScholar
2023

3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment

ICCV 2023poster

3D vision-language grounding (3D-VL) is an emerging field that aims to connect the 3D physical world with natural language, which is crucial for achieving embodied intelligence. Current 3D-VL models rely heavily on sophisticated modules, auxiliary losses, and optimization tricks, which calls for a s…

Cited by 129PDFScholar
2023

Cross-Modality Time-Variant Relation Learning for Generating Dynamic Scene Graphs

ICRA 2023poster

Dynamic scene graphs generated from video clips could help enhance the semantic visual understanding in a wide range of challenging tasks such as environmental perception, autonomous navigation, and task planning of self-driving vehicles and mobile robots. In the process of temporal and spatial mode…

Cited by 9SourcecodeScholar
2023

Learnable Flow Model Conditioned on Graph Representation Memory for Anomaly Detection

ICASSP 2023accepted

Anomaly detection could be applied in a wide range of fields from industrial scene to medical imaging analysis. Although invertible flow models are developed to accomplish unsupervised anomaly detection, they are usually hard to train and have limited capabilities of accurately modeling the distribu…

Cited by 0SourceScholar
2023

StyleTalk: One-Shot Talking Head Generation with Controllable Speaking Styles

AAAI 2023technical

Different people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they still cannot generate diverse speaking styles in the final talking head videos. To t…

2022

DuMLP-Pin: A Dual-MLP-Dot-Product Permutation-Invariant Network for Set Feature Extraction

AAAI 2022technical

Existing permutation-invariant methods can be divided into two categories according to the aggregation scope, i.e. global aggregation and local one. Although the global aggregation methods, e. g., PointNet and Deep Sets, get involved in simpler structures, their performance is poorer than the local…

2021

Feature Enhanced Projection Network for Zero-shot Semantic Segmentation

ICRA 2021poster

In environmental perception of autonomous driving, zero-shot semantic segmentation that can make prediction of new categories without using any labeled training samples is considered as a challenging task. One key step in this task is to transfer knowledge across categories via auxiliary semantic wo…

Cited by 4SourceScholar
2019

DrivingStereo: A Large-Scale Dataset for Stereo Matching in Autonomous Driving Scenarios

CVPR 2019poster

Great progress has been made on estimating disparity maps from stereo images. However, with the limited stereo data available in the existing datasets and unstable ranging precision of current stereo methods, industry-level stereo matching in autonomous driving remains challenging. In this paper, we…

Cited by 248PDFcodeScholar
2018

SegStereo: Exploiting Semantic Information for Disparity Estimation

ECCV 2018poster

Disparity estimation for binocular stereo images finds a wide range of applications. Traditional algorithms may fail on featureless regions, which could be handled by high-level clues such as semantic segments. In this paper, we suggest that appropriate incorporation of semantic cues can greatly rec…

Cited by 429SourcePDFScholar