← Search

Hao Xu

75 accepted papers

2026

AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora

AAAI 2026technical

Comprehension of ancient texts plays an important role in archaeology and understanding of Chinese history and civilization. The rapid development of large language models needs benchmarks that can evaluate their comprehension of ancient characters. Existing Chinese benchmarks are mostly targeted at

Cited by 0SourcePDFScholar
2026

Diff-VIO: A Diffusion Model-Based Pose Optimizer for Visual Inertial Odometry

ICRA 2026poster

Visual inertial odometry (VIO) serves as a cornerstone of environmental perception and spatial localization, with broad applications in autonomous driving, robotic navigation, and embodied intelligence. Although recent deep learning based VIO methods have achieved impressive accuracy and computation…

Cited by 0Scholar
2026

EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

ICLR 2026poster

Robust 3D hand reconstruction is challenging in egocentric vision due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior works attempt to mitigate the challenges by scaling up training data or incorporating auxiliary cues, often falling short of effectively handling unse…

Cited by 0SourcecodeScholar
2026

GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks

ICLR 2026poster

This paper introduces GraphOmni, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs on graph-theoretic tasks articulated in natural language. GraphOmni spans diverse graph types, serialization formats, and prompting schemes, substantially extending upon prior efforts i…

Cited by 0SourcecodeScholar
2026

IMH-MOT: Interactive Multi-Hierarchical Image and Point Cloud Fusion for Multi-Object Tracking

ICRA 2026poster

Multi-object tracking (MOT) plays a critical role in applications such as autonomous driving and surveillance. Camera-based approaches offer rich texture features for object association, while LiDAR-based methods provide accurate geometric information for spatial reasoning. Although each modality ad…

Cited by 0SourceScholar
2026

InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

AAAI 2026technical

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical texts. First, the scarcity of historical language samples r

Cited by 0SourcePDFScholar
2026

LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency

CVPR 2026

This paper introduces LaS-Comp, a zero-shot and category-agnostic approach that leverages the rich geometric priors of 3D foundation models to enable 3D shape completion across diverse types of partial observations. Our contributions are threefold: First, LaS-Comp harnesses these powerful generative

Cited by 0SourcecodeScholar
2026

PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation

CVPR 2026

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over talking face, such as speaking style and emotional expression, resulting in uniform facial motion. In this paper, we focus on impro

Cited by 0SourcecodeScholar
2026

Step-Level Sparse Autoencoder for Reasoning Process Interpretation

ICML 2026poster

Large Language Models (LLMs) have achieved strong complex reasoning capabilities through Chain-of-Thought (CoT) reasoning. However, their reasoning patterns remain too complicated to analyze. While Sparse Autoencoders (SAEs) have emerged as a powerful tool for interpretability, existing approaches p…

Cited by 0SourceScholar
2025

AdsQA: Towards Advertisement Video Understanding

ICCV 2025poster

Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpose models to continuously evolve via learning deeper expertise. Now is thus the time further to extend the diversity of…

2025

Annealing Distillation Algorithm for Transferring Unsupervised Clustering Knowledge to Supervised Student Models

ICASSP 2025accepted

In knowledge distillation, the performance of teacher models often serves as an upper limit for student models. For a long time, deeper and more accurate supervised learning algorithms have been the first choice for teacher models in image classification tasks where unsupervised models typically und…

Cited by 0SourceScholar
2025

Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents

EMNLP 2025

Large Language Models (LLMs) based agent systems have made great strides in real-world applications beyond traditional NLP tasks. This paper proposes a new LLM-based Multi-Agent System (LLM-MAS) benchmark, Collab-Overcooked, built on the popular Overcooked-AI game with more applicable and challengin

2025

DGJA: Dependency Graph-enhanced Joint Attention Structure for Multimodal Sarcasm Detection

ICASSP 2025accepted

Multimodal sarcasm detection (MSD) leverages multimodal data, including both images and text, to detect whether the input content contains sarcastic information. Despite recent advances, existing MSD approaches often overlook the imbalance in sarcastic content between text and image modalities, wher…

Cited by 0SourceScholar
2025

Dual-Pyramid Attention Collaborative Network for Oracle Bone Inscription Classification

ICASSP 2025accepted

Recent advances in oracle bone inscriptions (OBI) classification have explored various strategies such as zero-shot learning, augmentation, and complex convolution architectures. These strategies ultimately represent samples as global feature vectors in various forms but often fail to effectively ca…

Cited by 0SourceScholar
2025

Embodied Representation Alignment with Mirror Neurons

ICCV 2025poster

Mirror neurons are a class of neurons that activate both when an individual observes an action and when they perform the same action. This mechanism reveals a fundamental interplay between action understanding and embodied execution, suggesting that these two abilities are inherently connected. None…

Cited by 0SourcePDFScholar
2025

How Do Position Encodings Affect Length Generalization? Case Studies On In-Context Function Learning

AAAI 2025technical

The capability of In-Context Learning (ICL) is crucial for large language models to generalize across a wide range of tasks. By utilizing prompts, these models can accurately predict outcomes for previously unseen tasks without necessitating retraining. However, this generalization ability does not…

2025

HybridReg: Robust 3D Point Cloud Registration with Hybrid Motions

AAAI 2025technical

Scene-level point cloud registration is very challenging when considering dynamic foregrounds. Existing indoor datasets mostly assume rigid motions, so the trained models cannot robustly handle scenes with non-rigid motions. On the other hand, non-rigid datasets are mainly object-level, so the train…

2025

IMH-MOT: Interactive Multi-Hierarchical Image and Point Cloud Fusion for Multi-Object Tracking

RA-L 2025

Multi-object tracking (MOT) plays a critical role in applications such as autonomous driving and surveillance. Camera-based approaches offer rich texture features for object association, while LiDAR-based methods provide accurate geometric information for spatial reasoning. Although each modality ad

Cited by 0SourceScholar
2025

Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index

EMNLP 2025

Language models are trained mainly on massive text data from the Internet, and it becomes increasingly important to understand this data source. Exact-match search engines enable searching in large text corpora – counting string appearances and retrieving the enclosing documents – yet the high stora

2025

LOG: A Local-to-Global Optimization Approach for Retrieval-based Explainable Multi-Hop Question Answering

COLING 2025main

Multi-hop question answering (MHQA) aims to utilize multi-source intensive documents retrieved to derive the answer. However, it is very challenging to model the importance of knowledge retrieved. Previous approaches primarily emphasize single-step and multi-step iterative decomposition or retrieval…

2025

Learning to Factorize Spatio-Temporal Foundation Models

NeurIPS 2025spotlight

Spatio-Temporal Foundation Models (STFMs) promise zero/few-shot generalization across various datasets, yet joint spatio-temporal pretraining is computationally prohibitive and struggles with domain-specific spatial correlations. To this end, we introduce FactoST, a factorized STFM that decouples un…

Cited by 0SourceScholar
2025

MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable Recommendation

ACL 2025long

Explainable Recommendation task is designed to receive a pair of user and item and output explanations to justify why an item is recommended to a user. Many models approach review generation as a proxy for explainable recommendations. While these models can produce fluent and grammatically correct s…

2025

Mat-Instructions: A Large-Scale Inorganic Material Instruction Dataset for Large Language Models

IJCAI 2025

Recent advancements in large language models (LLMs) have revolutionized research discovery across various scientific disciplines, including materials science. The discovery of novel materials, particularly crystal materials, is essential for achieving sustainable development goals (SDGs), as they dr

2025

MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds

NeurIPS 2025poster

Reconstructing 3D objects into editable programs is pivotal for applications like reverse engineering and shape editing. However, existing methods often rely on limited domain-specific languages (DSLs) and small-scale datasets, restricting their ability to model complex geometries and structures. To…

Cited by 0SourceScholar
2025

MoME: Mixture of Multi-Domain Experts for Multivariate Long-Term Series Forecasting

ICASSP 2025accepted

Time series forecasting is always important, with multivariate long-term series forecasting being its most challenging task. Here, the existing methods typically learn only in a single domain and focus on optimizing model structures, leading to incomplete information mining and imprecise predictions…

Cited by 0SourceScholar
2025

Multirate Neural Image Compression with Adaptive Lattice Vector Quantization

CVPR 2025highlight

Recent research has explored integrating lattice vector quantization (LVQ) into learned image compression models. Due to its more efficient Voronoi covering of vector space than scalar quantization (SQ), LVQ achieves better rate-distortion (R-D) performance than SQ, while still retaining the low com…

2025

RAOCSL: A BERT-Based Strategy for Identifying Learner Confusion under Class Imbalance

ICASSP 2025accepted

Understanding and identifying the nature of learner confusion is important for online learning platforms. In this study, we address this problem by analyzing forum posts from large-scale online courses. However, due to the large volume of comments and frequent interactions, confusion posts are often…

Cited by 0SourceScholar
2025

Sensitivity-LoRA : Low-Load Sensitivity-Based Fine-Tuning for Large Language Models

EMNLP 2025

Large Language Models (LLMs) have transformed both everyday life and scientific research. However, adapting LLMs from general-purpose models to specialized tasks remains challenging, particularly in resource-constrained environments. Low-Rank Adaptation (LoRA), a prominent method within Parameter-Ef

Cited by 0SourcePDFScholar
2025

Simulating Dual-Pixel Images From Ray Tracing For Depth Estimation

ICCV 2025poster

Many studies utilize dual-pixel (DP) sensor phase information for various applications, such as depth estimation and deblurring. However, since DP image features are entirely determined by the camera hardware, DP-depth paired datasets are very scarce, especially when performing depth estimation on c…

2025

UN-DETR: Promoting Objectness Learning via Joint Supervision for Unknown Object Detection

AAAI 2025technical

Unknown Object Detection (UOD) aims to identify objects of unseen categories, differing from the traditional detection paradigm limited by the closed-world assumption. A key component of UOD is learning a generalized representation, i.e. objectness for both known and unknown categories to distinguis…

2025

UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation

CVPR 2025poster

Estimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly handle both scenarios and their performance degrades when appli…

2024

An Effective Span-based Multimodal Named Entity Recognition with Consistent Cross-Modal Alignment

COLING 2024main

With the increasing availability of multimodal content on social media, consisting primarily of text and images, multimodal named entity recognition (MNER) has gained a wide-spread attention. A fundamental challenge of MNER lies in effectively aligning different modalities. However, the majority of…

Cited by 0SourcePDFScholar
2024

Ancient Chinese Glyph Identification Powered by Radical Semantics

ACL 2024findings

The ancestor of Chinese character – the ancient characters from about 1300 BC to 200 BC are not fixed in their writing glyphs. At the same or different points in time, one character can possess multiple glyphs that are different in shapes or radicals. Nearly half of ancient glyphs have not been deci…

2024

Build a 50+ Hours Chinese Mandarin Corpus for Children's Speech Recognition

ICASSP 2024accepted

Children’s speech recognition plays an important role in the education research of children. The usual automatic speech recognition (ASR) systems are not satisfactory in terms of speech recognition for children, mainly due to the lack of child speech corpus. In recent years, there have been a large…

Cited by 0SourceScholar
2024

Deciphering Compatibility Relationships with Textual Descriptions via Extraction and Explanation

AAAI 2024technical

Understanding and accurately explaining compatibility relationships between fashion items is a challenging problem in the burgeoning domain of AI-driven outfit recommendations. Present models, while making strides in this area, still occasionally fall short, offering explanations that can be element…

2024

EmoPrompt-ECPE: Emotion Knowledge-aware Prompt-tuning for Emotion-Cause Pair Extraction

COLING 2024main

Emotion-cause pair extraction (ECPE) main focus is on extracting all potential emotion clauses and corresponding cause clauses from unannotated documents. Existing methods achieve promising results with the help of fine-tuning and prompt paradigms, but they present three downsides. First, most appro…

2024

HHD-GP: Incorporating Helmholtz-Hodge Decomposition into Gaussian Processes for Learning Dynamical Systems

NeurIPS 2024poster

Machine learning models provide alternatives for efficiently recognizing complex patterns from data, but the main concern in applying them to modeling physical systems stems from their physics-agnostic design, leading to learning methods that lack interpretability, robustness, and data efficiency. T…

Cited by 0SourcePDFScholar
2024

HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object Interactions

CVPR 2024poster

Reconstructing 3D hand mesh robustly from a single image is very challenging due to the lack of diversity in existing real-world datasets. While data synthesis helps relieve the issue the syn-to-real gap still hinders its usage. In this work we present HandBooster a new approach to uplift the data d…

2024

Hierarchical Topic Modeling via Contrastive Learning and Hyperbolic Embedding

COLING 2024main

Hierarchical topic modeling, which can mine implicit semantics in the corpus and automatically construct topic hierarchical relationships, has received considerable attention recently. However, the current hierarchical topic models are mainly based on Euclidean space, which cannot well retain the im…

2024

OmniNxt: A Fully Open-source and Compact Aerial Robot with Omnidirectional Visual Perception

IROS 2024poster

Adopting omnidirectional Field of View (FoV) cameras in aerial robots vastly improves perception ability, significantly advancing aerial robotics’s capabilities in inspection, reconstruction, and rescue tasks. However, such sensors also elevate system complexity, e.g., hardware design, and correspon…

Cited by 9SourcecodeScholar
2024

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

ECCV 2024poster

"In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel method, namely Diff2Scene, which leverages frozen representations from text-image generative models, along with salien…

Cited by 4SourcePDFScholar
2024

POP-CEE: Position-oriented Prompt-tuning Model for Causal Emotion Entailment

ACL 2024findings

The objective of the Causal Emotion Entailment (CEE) task is to identify the causes of the target emotional utterances in a given conversation. Most existing studies have focused on a fine-tuning paradigm based on a pretrained model, e.g., the BERT model. However, there are gaps between the pretrain…

2024

PointRegGPT: Boosting 3D Point Cloud Registration using Generative Point-Cloud Pairs for Training

ECCV 2024poster

"Data plays a crucial role in training learning-based methods for 3D point cloud registration. However, the real-world dataset is expensive to build, while rendering-based synthetic data suffers from domain gaps. In this work, we present , boosting 3D Point cloud Registration using Generative Point-…

2024

SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth Estimation

AAAI 2024technical

Recently, self-supervised monocular depth estimation has gained popularity with numerous applications in autonomous driving and robotics. However, existing solutions primarily seek to estimate depth from immediate visual features, and struggle to recover fine-grained scene details. In this paper, we…

2024

Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness

EMNLP 2024finding

Recently, there has been significant interest in replacing the reward model in Reinforcement Learning with Human Feedback (RLHF) methods for Large Language Models (LLMs), such as Direct Preference Optimization (DPO) and its variants. These approaches commonly use a binary cross-entropy mechanism on…

2024

SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation

AAAI 2024technical

Estimating 3D hand mesh from RGB images is a longstanding track, in which occlusion is one of the most challenging problems. Existing attempts towards this task often fail when the occlusion dominates the image space. In this paper, we propose SiMA-Hand, aiming to boost the mesh reconstruction perfo…

2024

TACIT: A Target-Agnostic Feature Disentanglement Framework for Cross-Domain Text Classification

AAAI 2024technical

Cross-domain text classification aims to transfer models from label-rich source domains to label-poor target domains, giving it a wide range of practical applications. Many approaches promote cross-domain generalization by capturing domaininvariant features. However, these methods rely on unlabeled…

2023

D-IF: Uncertainty-aware Human Digitization via Implicit Distribution Field

ICCV 2023poster

Realistic virtual humans play a crucial role in numerous industries, such as metaverse, intelligent healthcare, and self-driving simulation. But creating them on a large scale with high levels of realism remains a challenge. The utilization of deep implicit function sparks a new era of image-based 3…

Cited by 40PDFcodeScholar
2023

Differentiable Meta Multigraph Search with Partial Message Propagation on Heterogeneous Information Networks

AAAI 2023technical

Heterogeneous information networks (HINs) are widely employed for describing real-world data with intricate entities and relationships. To automatically utilize their semantic information, graph neural architecture search has recently been developed for various tasks of HINs. Existing works, on the…

2023

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

ICCV 2023poster

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content. To address this issue, this paper proposes an end-to-end neur…

Cited by 117PDFcodeScholar
2023

Enhancing Ontology Translation Through Cross-Lingual Agreement

ICASSP 2023accepted

Ontology serves as the foundation for the underlying representation of knowledge. In order to achieve the sharing of knowledge across languages, ontologies that are typically only represented in English must be translated into different languages. Building a domain-specific translation system is nec…

Cited by 0SourceScholar
2023

FCC: Feature Clusters Compression for Long-Tailed Visual Recognition

CVPR 2023poster

Deep Neural Networks (DNNs) are rather restrictive in long-tailed data, since they commonly exhibit an under-representation for minority classes. Various remedies have been proposed to tackle this problem from different perspectives, but they ignore the impact of the density of Backbone Features (BF…

2023

H2ONet: Hand-Occlusion-and-Orientation-Aware Network for Real-Time 3D Hand Mesh Reconstruction

CVPR 2023poster

Real-time 3D hand mesh reconstruction is challenging, especially when the hand is holding some object. Beyond the previous methods, we design H2ONet to fully exploit non-occluded information from multiple frames to boost the reconstruction quality. First, we decouple hand mesh reconstruction into tw…

2023

Planning Assembly Sequence with Graph Transformer

ICRA 2023poster

Assembly Sequence Planning (ASP) is the essential process for modern manufacturing, proven to be NP-complete thus its effective and efficient solution has been a challenge for researchers in the field. In this paper, we present a graph-transformer based framework for the ASP problem which is trained…

Cited by 23SourcecodeScholar
2023

RZCR: Zero-shot Character Recognition via Radical-based Reasoning

IJCAI 2023poster

The long-tail effect is a common issue that limits the performance of deep learning models on real-world datasets. Character image datasets are also affected by such unbalanced data distribution due to differences in character usage frequency. Thus, current character recognition methods are limited…

Cited by 14SourcePDFScholar
2023

SIRA-PCR: Sim-to-Real Adaptation for 3D Point Cloud Registration

ICCV 2023poster

Point cloud registration is essential for many applications. However, existing real datasets require extremely tedious and costly annotations, yet may not provide accurate camera poses. For the synthetic datasets, they are mainly object-level, so the trained models may not generalize well to real sc…

Cited by 22PDFcodeScholar
2023

Speaker Diaphragm Excursion Prediction: Deep Attention and Online Adaptation

ICASSP 2023accepted

Speaker protection algorithm is to leverage the playback signal properties to prevent over excursion while maintaining maximum loudness, especially for the mobile phone with tiny loudspeakers. This paper proposes efficient DL solutions to accurately model and predict the nonlinear excursion, which i…

Cited by 0SourceScholar
2022

A Simple Contrastive Learning Framework for Interactive Argument Pair Identification via Argument-Context Extraction

EMNLP 2022main

Interactive argument pair identification is an emerging research task for argument mining, aiming to identify whether two arguments are interactively related. It is pointed out that the context of the argument is essential to improve identification performance. However, current context-based methods…

2022

Ensemble Multi-Relational Graph Neural Networks

IJCAI 2022poster

It is well established that graph neural networks (GNNs) can be interpreted and designed from the perspective of optimization objective. With this clear optimization objective, the deduced GNNs architecture has sound theoretical foundation, which is able to flexibly remedy the weakness of GNNs. Howe…

2022

FINet: Dual Branches Feature Interaction for Partial-to-Partial Point Cloud Registration

AAAI 2022technical

Data association is important in the point cloud registration. In this work, we propose to solve the partial-to-partial registration from a new perspective, by introducing multi-level feature interactions between the source and the reference clouds at the feature extraction stage, such that the regi…

2022

Federated Multi-Task Attention for Cross-Individual Human Activity Recognition

IJCAI 2022poster

Federated Learning (FL) is an emerging privacy-aware machine learning technique that applies successfully to the collaborative learning of global models for Human Activity Recognition (HAR). As of now, the applications of FL for HAR assume that the data associated with diverse individuals follow the…

2022

Improving Few-Shot Text-to-SQL with Meta Self-Training via Column Specificity

IJCAI 2022poster

The few-shot problem is an urgent challenge for single-table text-to-SQL. Existing methods ignore the potential value of unlabeled data, and merely rely on a coarse-grained Meta-Learning (ML) algorithm that neglects the differences of column contributions to the optimization object. This paper propo…

2022

ZiNet: Linking Chinese Characters Spanning Three Thousand Years

ACL 2022findings

Modern Chinese characters evolved from 3,000 years ago. Up to now, tens of thousands of glyphs of ancient characters have been discovered, which must be deciphered by experts to interpret unearthed documents. Experts usually need to compare each ancient character to be examined with similar known on…

2021

Design and Analysis of a Bi-directional Transformable Wheel Robot Trimode

IROS 2021poster

This article presents a novel transformable wheeled robot with three motion modes. Based on the four-bar mechanism design, the transformable wheel can transform into CW (clockwise) legged wheel mode and CCW (counterclock-wise) legged wheel mode from the circular wheel. Both legged wheel modes achiev…

Cited by 15SourceScholar
2021

OMNet: Learning Overlapping Mask for Partial-to-Partial Point Cloud Registration

ICCV 2021poster

Point cloud registration is a key task in many computational fields. Previous correspondence matching based methods require the inputs to have distinctive geometric structures to fit a 3D rigid transformation according to point-wise sparse feature matches. However, the accuracy of transformation hea…

Cited by 207PDFcodeScholar
2020

Decentralized Visual-Inertial-UWB Fusion for Relative State Estimation of Aerial Swarm

ICRA 2020poster

The collaboration of unmanned aerial vehicles (UAVs) has become a popular research topic for its practicability in multiple scenarios. The collaboration of multiple UAVs, which is also known as aerial swarm is a highly complex system, which still lacks a state-of-art decentralized relative state est…

Cited by 161SourceScholar
2020

Distributed Consensus Control of Multiple UAVs in a Constrained Environment

ICRA 2020poster

In this paper, we investigate the consensus problem of multiple unmanned aerial vehicles (UAVs) in the presence of environmental constraints under a general communication topology containing a directed spanning tree. First, based on a position transformation function, we propose a novel dynamic refe…

Cited by 15SourceScholar
2018

The Deformable Quad-Rotor Enabled and Wasp-Pedal-Carrying Inspired Aerial Gripper

IROS 2018poster

The paper presents the development of a novel deformable quad-rotor enabled aerial gripper. The mechanism of our deformable quad-rotor is based on simultaneous expansion or contraction of the quad-rotor body, which is generated by controlling a rigid elements based morphing structure (REMS). Such de…

Cited by 25SourceScholar