← Search

Hao Xiong

25 accepted papers

2026

AutoMat: Physics-Guided Agentic Reasoning for Solving Ill-Posed Inverse Microscopy Problems

ICML 2026poster

Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain similar contrast, and purely feed-forward models cannot verify physical validity. We present **AutoMat**, a failure-aware agentic *controller* that performs …

Cited by 0SourceScholar
2026

CoEvol-NO: State and Coordinate Co-Evolution with an Error-Driven Predictor-Corrector Paradigm for Neural Operator Transformer

ICML 2026oral

Despite the fast progress in neural operator learning, long-sequence modeling still is a standing challenge whereby latent states have been introduced with techniques well derived. Diverging from existing methods that treat latent states as transient variables or decoupled representations, CoEvol-NO…

Cited by 0SourceScholar
2026

DART: Navigating Last-Mile Heterogeneity in Instant Delivery via Distribution-Adaptive Splines

IJCAI 2026

On-demand delivery platforms rely on Travel Time Estimation (TTE) to balance courier earnings and overdue risks. In collaboration with one of China's largest platforms, we address a critical "Fairness Gap" in TTE: current systems fail to capture complex delivery patterns in GNSS-denied environments,

Cited by 0Scholar
2026

SAGA: Source Attribution of Generative AI Videos

CVPR 2026

The proliferation of generative AI has led to hyper-realistic synthetic videos, escalating misuse risks and outstripping binary real/fake detectors. We introduce \texttt SAGA (\underline S ource \underline A ttribution of \underline G enerative \underline A I videos), the first comprehensive framewo

Cited by 0SourcecodeScholar
2026

SCE-Depth: A Spherical Compound Eye Framework for Wide FOV Depth Estimation

CVPR 2026

Accurate depth estimation in wide field is highly desired in applications of autonomous driving, robot vision and drone controls. Biological compound eyes inspire wide Field of View (FOV) depth estimation, yet their artificial implementations face the challenge of modality misalignment. Specifically

Cited by 0SourcecodeScholar
2025

A Conditional KAN Diffusion Network for Human Activity Recognition with Missing Sensor Signal Series

ICASSP 2025accepted

Human Activity Recognition (HAR) is crucial for applications like urban traffic management and health monitoring but faces challenges in handling complex patterns and missing sensor data. In this work, we propose a conditional Kolmogorov-Arnold network diffusion (CKAD) framework for HAR, which separ…

Cited by 0SourceScholar
2025

CSS: Overcoming Pose and Scene Challenges in Crowd-Sourced 3D Gaussian Splatting

ICASSP 2025accepted

We introduce Crowd-Sourced Splatting (CSS), a novel 3D Gaussian Splatting (3DGS) pipeline designed to overcome the challenges of pose-free scene reconstruction using crowd-sourced imagery. The dream of reconstructing historically significant but inaccessible scenes from collections of photographs ha…

Cited by 0SourceScholar
2025

Map-Free Visual Relocalization Enhanced by Instance Knowledge and Depth Knowledge

ICASSP 2025accepted

Map-free visual relocalization computes camera pose using only a query image and a reference image. Therefore, it is hindered by challenges in feature-point matching and the absence of scale information in monocular images. These issues may cause significant rotational and metric errors, leading to…

Cited by 0SourceScholar
2025

On Designing General and Expressive Quantum Graph Neural Networks with Applications to MILP Instance Representation

ICLR 2025poster

Graph-structured data is ubiquitous, and graph learning models have recently been extended to address complex problems like mixed-integer linear programming (MILP). However, studies have shown that the vanilla message-passing based graph neural networks (GNNs) suffer inherent limitations in learning…

Cited by 1SourcePDFScholar
2025

Tensor Network: from the Perspective of AI4Science and Science4AI

IJCAI 2025

Tensor network has been a promising numerical tool for computational problems across science and AI. For their emerging and fast development especially in the intersection between AI and science, this paper tries to present a compact review, regarding both their applications and its own recent techn

Cited by 0SourcePDFScholar
2025

Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated Content

CVPR 2025poster

Existing DeepFake detection techniques primarily focus on facial manipulations, such as face-swapping or lip-syncing. However, advancements in text-to-video (T2V) and image-to-video (I2V) generative models now allow fully AI-generated synthetic content and seamless background alterations, challengin…

Cited by 3SourcePDFScholar
2025

UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions

ACL 2025finding

Handling unanswerable questions (UAQ) is crucial for LLMs, as it helps prevent misleading responses in complex situations. While previous studies have built several datasets to assess LLMs’ performance on UAQ, these datasets lack factual knowledge support, which limits the evaluation of LLMs’ abilit…

2025

UniCO: On Unified Combinatorial Optimization via Problem Reduction to Matrix-Encoded General TSP

ICLR 2025poster

Various neural solvers have been devised for combinatorial optimization (CO), which are often tailored for specific problem types, e.g., TSP, CVRP and SAT, etc. Yet, it remains an open question how to achieve universality regarding problem representing and learning with a general framework. This pap…

Cited by 1SourcePDFScholar
2024

Circuit Design and Efficient Simulation of Quantum Inner Product and Empirical Studies of Its Effect on Near-Term Hybrid Quantum-Classic Machine Learning

CVPR 2024poster

For the essential operation namely inner product (IP) as widely adopted in classic computing e.g. matrix multiplication its quantum counterpart: quantum inner product (QIP) has also been recently theoretically explored with a verifiable lower complexity on quantum computers. However it remains uncle…

2024

Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation

ICASSP 2024accepted

Medical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation.…

Cited by 0SourceScholar
2024

Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use

CVPR 2024poster

From content moderation to wildlife conservation the number of applications that require models to recognize nuanced or subjective visual concepts is growing. Traditionally developing classifiers for such concepts requires substantial manual effort measured in hours days or even months to identify a…

Cited by 8SourcePDFScholar
2024

Node2ket: Efficient High-Dimensional Network Embedding in Quantum Hilbert Space

ICLR 2024poster

Network embedding (NE) is a prominent technique for network analysis where the nodes are represented as vectorized embeddings in a continuous space. Existing works tend to resort to the low-dimensional embedding space for efficiency and less risk of over-fitting. In this paper, we explore a new NE p…

Cited by 2SourcePDFScholar
2024

Towards LLM4QPE: Unsupervised Pretraining of Quantum Property Estimation and A Benchmark

ICLR 2024spotlight

Estimating the properties of quantum systems such as quantum phase has been critical in addressing the essential quantum many-body problems in physics and chemistry. Deep learning models have been recently introduced to property estimation, surpassing conventional statistical approaches. However, t…

Cited by 3SourcePDFScholar
2023

Dynamic Obstacle Avoidance for Cable-Driven Parallel Robots With Mobile Bases via Sim-to-Real Reinforcement Learning

RA-L 2023

A Cable-Driven Parallel Robot (CDPR) with Mobile Bases (MBs) can modify its geometric architecture and is suitable for manipulation tasks in constrained environments. In manipulation tasks, a CDPR with MBs inevitably encounters obstacles, including dynamic obstacles. However, the high dimensional st

Cited by 31SourceScholar
2023

Online Visual SLAM Adaptation against Catastrophic Forgetting with Cycle-Consistent Contrastive Learning

ICRA 2023poster

Visual SLAM (Simultaneous Localisation and Mapping) aims to simultaneously estimate camera poses and depth maps from navigation videos captured. While recent deep learning based methods have achieved great success on this task, they tend to work well on source domain data and suffer from performance…

Cited by 3SourceScholar
2022

Boosting the Performance of Generic Deep Neural Network Frameworks with Log-supermodular CRFs

NeurIPS 2022accept

Historically, conditional random fields (CRFs) were popular tools in a variety of application areas from computer vision to natural language processing, but due to their higher computational cost and weaker practical performance, they have, in many situations, fallen out of favor and been replaced b…

Cited by 0SourcePDFScholar
2022

Data-Driven Kinematic Control Scheme for Cable-Driven Parallel Robots Allowing Collisions

IROS 2022poster

Cable-Driven Parallel Robots (CDPRs) have been proposed for a variety of applications such as material handling, rehabilitation, and instrumentation. However, the collision-free constraint of CDPRs limits the workspace of CDPRs and the feasible position of anchor points. To address the collision-fre…

Cited by 10SourceScholar
2022

Enhancing Sequential Recommendation with Graph Contrastive Learning

IJCAI 2022poster

The sequential recommendation systems capture users' dynamic behavior patterns to predict their next interaction behaviors. Most existing sequential recommendation methods only exploit the local context information of an individual interaction sequence and learn model parameters solely based on the…

Cited by 74SourcePDFScholar