← Search

Hongan Wang

19 accepted papers

2026

BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model

RA-L 2026

Vision-Language-Action (VLA) models have achieved significant breakthroughs by leveraging Large Vision Language Models (VLMs) to jointly interpret instructions and visual inputs. However, the substantial increase in visual tokens, particularly from multi-view inputs, poses serious challenges to real

Cited by 0SourceScholar
2026

Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagrams

CVPR 2026

Diagrams are widely used in daily life. However, offline diagrams typically exist in the form of images, lacking structured data representation, which significantly limits their reusability and editability. Current research mainly focuses on supporting basic query tasks for online diagrams and does

Cited by 0SourceScholar
2026

TCATSEG: A TOOTH CENTER-WISE ATTENTION NETWORK FOR 3D DENTAL MODEL SEMANTIC SEGMENTATION

ICASSP 2026poster

Accurate semantic segmentation of 3D dental models is essential for digital dentistry applications such as orthodontics and dental implants. However, due to complex tooth arrangements and similarities in shape among adjacent teeth, existing methods struggle with accurate segmentation, because they o…

Cited by 0SourcePDFScholar
2025

DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model

AAAI 2025technical

Generating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of involvement of each hand given the complex shapes of objects in d…

Cited by 1SourcePDFScholar
2025

HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation

AAAI 2025technical

Understanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it is challenging to collect and annotate a high-quality, large-scale dataset, whi…

Cited by 0SourcePDFScholar
2025

Representation Learning with Mutual Influence of Modalities for Node Classification in Multi-Modal Heterogeneous Networks

IJCAI 2025

Nowadays, numerous online platforms can be described as multi-modal heterogeneous networks (MMHNs), such as Douban's movie networks and Amazon's product review networks. Accurately categorizing nodes within these networks is crucial for analyzing the corresponding entities, which requires effective

2025

Universal Features Guided Zero-Shot Category-Level Object Pose Estimation

AAAI 2025technical

Object pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to estab…

Cited by 0SourcePDFScholar
2024

SceneDiff: Generative Scene-Level Image Retrieval with Text and Sketch Using Diffusion Models

IJCAI 2024poster

Jointly using text and sketch for scene-level image retrieval utilizes the complementary between text and sketch to describe the fine-grained scene content and retrieve the target image, which plays a pivotal role in accurate image retrieval. Existing methods directly fuse the features of sketch and…

Cited by 0SourcePDFScholar
2024

SpaceGTN: A Time-Agnostic Graph Transformer Network for Handwritten Diagram Recognition and Segmentation

AAAI 2024technical

Online handwriting recognition is pivotal in domains like note-taking, education, healthcare, and office tasks. Existing diagram recognition algorithms mainly rely on the temporal information of strokes, resulting in a decline in recognition performance when dealing with notes that have been modifie…

Cited by 2SourcePDFScholar
2023

Novel-View Synthesis and Pose Estimation for Hand-Object Interaction from Sparse Views

ICCV 2023poster

Hand-object interaction understanding and the barely addressed novel view synthesis are highly desired in the immersive communication, whereas it is challenging due to the high deformation of hand and heavy occlusions between hand and object. In this paper, we propose a neural rendering and pose est…

Cited by 16PDFcodeScholar
2023

Self-supervised Learning of Implicit Shape Representation with Dense Correspondence for Deformable Objects

ICCV 2023poster

Learning 3D shape representation with dense correspondence for deformable objects is a fundamental problem in computer vision. Existing approaches often need additional annotations of specific semantic domain, e.g., skeleton pose for human body or animals, which require extra annotation effort and s…

Cited by 9PDFScholar
2023

Unsupervised Model-Based Speaker Adaptation of End-To-End Lattice-Free MMI Model for Speech Recognition

ICASSP 2023accepted

Modeling the speaker variability is a key challenge for automatic speech recognition (ASR) systems. In this paper, the learning hidden unit contributions (LHUC) based adaptation techniques with compact speaker dependent (SD) parameters are used to facilitate both speaker adaptive training (SAT) and…

Cited by 0SourceScholar
2022

Efficient Virtual View Selection for 3D Hand Pose Estimation

AAAI 2022technical

3D hand pose estimation from single depth is a fundamental problem in computer vision, and has wide applications. However, the existing methods still can not achieve satisfactory hand pose estimation results due to view variation and occlusion of human hand. In this paper, we propose a new virtual v…

2021

Interacting Two-Hand 3D Pose and Shape Reconstruction From Single Color Image

ICCV 2021poster

In this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solu…

Cited by 111PDFcodeScholar
2020

Dataless Short Text Classification Based on Biterm Topic Model and Word Embeddings

IJCAI 2020poster

Dataless text classification has attracted increasing attentions recently. It only needs very few seed words of each category to classify documents, which is much cheaper than supervised text classification that requires massive labeling efforts. However, most of existing models pay attention to lon…

2020

SceneSketcher: Fine-Grained Image Retrieval with Scene Sketches

ECCV 2020poster

Sketch-based image retrieval (SBIR) has been a popular research topic in recent years. Existing works concentrate on mapping the visual information of sketches and images to a semantic space at the object level. In this paper, for the first time, we study the fine-grained scene-level SBIR problem wh…

Cited by 45SourcePDFScholar
2019

Cascaded Point Network for 3D Hand Pose Estimation

ICASSP 2019accepted

Recent PointNet-family hand pose methods have the advantages of high pose estimation performance and small model size, and it is a key problem to get effective sample points for PointNet-family methods. In this paper, we propose a two-stage coarse to fine hand pose estimation method, which belongs t…

Cited by 0SourceScholar
2019

SketchGAN: Joint Sketch Completion and Recognition With Generative Adversarial Network

CVPR 2019poster

Hand-drawn sketch recognition is a fundamental problem in computer vision, widely used in sketch-based image and video retrieval, editing, and reorganization. Previous methods often assume that a complete sketch is used as input; however, hand-drawn sketches in common application scenarios are often…

Cited by 71PDFScholar