← Search

Xiaoming Deng

18 accepted papers

2026

BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model

RA-L 2026

Vision-Language-Action (VLA) models have achieved significant breakthroughs by leveraging Large Vision Language Models (VLMs) to jointly interpret instructions and visual inputs. However, the substantial increase in visual tokens, particularly from multi-view inputs, poses serious challenges to real

Cited by 0SourceScholar
2026

Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagrams

CVPR 2026

Diagrams are widely used in daily life. However, offline diagrams typically exist in the form of images, lacking structured data representation, which significantly limits their reusability and editability. Current research mainly focuses on supporting basic query tasks for online diagrams and does

Cited by 0SourceScholar
2026

SketchRevive: Fine-Grained Pixel-to-Vector Sketch Completion with Diffusion-Prior-Guided Multimodal LLMs

CVPR 2026

Transforming sparse, partial pixel sketches from diverse media into complete, editable vector drawings is essential yet underexplored in digital creation. Prior methods either generate from scratch or inpaint local gaps without predicting global structure, leading to coarse contours and limited deta

Cited by 0SourceScholar
2026

TCATSEG: A TOOTH CENTER-WISE ATTENTION NETWORK FOR 3D DENTAL MODEL SEMANTIC SEGMENTATION

ICASSP 2026poster

Accurate semantic segmentation of 3D dental models is essential for digital dentistry applications such as orthodontics and dental implants. However, due to complex tooth arrangements and similarities in shape among adjacent teeth, existing methods struggle with accurate segmentation, because they o…

Cited by 0SourcePDFScholar
2025

DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model

AAAI 2025technical

Generating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of involvement of each hand given the complex shapes of objects in d…

Cited by 1SourcePDFScholar
2025

HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation

AAAI 2025technical

Understanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it is challenging to collect and annotate a high-quality, large-scale dataset, whi…

Cited by 0SourcePDFScholar
2025

Universal Features Guided Zero-Shot Category-Level Object Pose Estimation

AAAI 2025technical

Object pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to estab…

Cited by 0SourcePDFScholar
2024

Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction

CVPR 2024poster

Reliable hand mesh reconstruction (HMR) from commonly-used color and depth sensors is challenging especially under scenarios with varied illuminations and fast motions. Event camera is a highly promising alternative for its high dynamic range and dense temporal resolution properties but it lacks key…

Cited by 8SourcePDFScholar
2024

SceneDiff: Generative Scene-Level Image Retrieval with Text and Sketch Using Diffusion Models

IJCAI 2024poster

Jointly using text and sketch for scene-level image retrieval utilizes the complementary between text and sketch to describe the fine-grained scene content and retrieve the target image, which plays a pivotal role in accurate image retrieval. Existing methods directly fuse the features of sketch and…

Cited by 0SourcePDFScholar
2024

SpaceGTN: A Time-Agnostic Graph Transformer Network for Handwritten Diagram Recognition and Segmentation

AAAI 2024technical

Online handwriting recognition is pivotal in domains like note-taking, education, healthcare, and office tasks. Existing diagram recognition algorithms mainly rely on the temporal information of strokes, resulting in a decline in recognition performance when dealing with notes that have been modifie…

Cited by 2SourcePDFScholar
2023

Novel-View Synthesis and Pose Estimation for Hand-Object Interaction from Sparse Views

ICCV 2023poster

Hand-object interaction understanding and the barely addressed novel view synthesis are highly desired in the immersive communication, whereas it is challenging due to the high deformation of hand and heavy occlusions between hand and object. In this paper, we propose a neural rendering and pose est…

Cited by 16PDFcodeScholar
2023

Self-supervised Learning of Implicit Shape Representation with Dense Correspondence for Deformable Objects

ICCV 2023poster

Learning 3D shape representation with dense correspondence for deformable objects is a fundamental problem in computer vision. Existing approaches often need additional annotations of specific semantic domain, e.g., skeleton pose for human body or animals, which require extra annotation effort and s…

Cited by 9PDFScholar
2022

Efficient Virtual View Selection for 3D Hand Pose Estimation

AAAI 2022technical

3D hand pose estimation from single depth is a fundamental problem in computer vision, and has wide applications. However, the existing methods still can not achieve satisfactory hand pose estimation results due to view variation and occlusion of human hand. In this paper, we propose a new virtual v…

2021

Interacting Two-Hand 3D Pose and Shape Reconstruction From Single Color Image

ICCV 2021poster

In this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solu…

Cited by 111PDFcodeScholar
2021

Sequential 3D Human Pose Estimation Using Adaptive Point Cloud Sampling Strategy

IJCAI 2021poster

3D human pose estimation is a fundamental problem in artificial intelligence, and it has wide applications in AR/VR, HCI and robotics. However, human pose estimation from point clouds still suffers from noisy points and estimated jittery artifacts because of handcrafted-based point cloud sampling an…

2020

SceneSketcher: Fine-Grained Image Retrieval with Scene Sketches

ECCV 2020poster

Sketch-based image retrieval (SBIR) has been a popular research topic in recent years. Existing works concentrate on mapping the visual information of sketches and images to a semantic space at the object level. In this paper, for the first time, we study the fine-grained scene-level SBIR problem wh…

Cited by 45SourcePDFScholar
2019

Cascaded Point Network for 3D Hand Pose Estimation

ICASSP 2019accepted

Recent PointNet-family hand pose methods have the advantages of high pose estimation performance and small model size, and it is a key problem to get effective sample points for PointNet-family methods. In this paper, we propose a two-stage coarse to fine hand pose estimation method, which belongs t…

Cited by 0SourceScholar
2019

SketchGAN: Joint Sketch Completion and Recognition With Generative Adversarial Network

CVPR 2019poster

Hand-drawn sketch recognition is a fundamental problem in computer vision, widely used in sketch-based image and video retrieval, editing, and reorganization. Previous methods often assume that a complete sketch is used as input; however, hand-drawn sketches in common application scenarios are often…

Cited by 71PDFScholar