← Search

Jun Hu

26 accepted papers

2026

DGP: A Dual-Granularity Prompting Framework for Fraud Detection with Graph-Enhanced LLMs

AAAI 2026technical

Real-world fraud detection applications benefit from graph learning techniques that jointly exploit node features—often rich in textual data—and graph structural information. Recently, Graph-Enhanced LLMs have emerged as a promising graph learning approach that converts graph information into prompt

Cited by 0SourcePDFScholar
2026

Echoless Label-Based Pre-computation for Memory-Efficient Heterogeneous Graph Learning

AAAI 2026technical

Heterogeneous Graph Neural Networks (HGNNs) are widely used for deep learning on heterogeneous graphs. Typical end-to-end HGNNs require repetitive message passing during training, limiting efficiency for large-scale real-world graphs. Pre-computation-based HGNNs address this by performing message pa

Cited by 0SourcePDFScholar
2026

NTSFormer: A Self-Teaching Graph Transformer for Multimodal Isolated Cold-Start Node Classification

AAAI 2026technical

Isolated cold-start node classification on multimodal graphs is challenging because such nodes have no edges and often have missing modalities (e.g., absent text or image features). Existing methods address structural isolation by degrading graph learning models to multilayer perceptrons (MLPs) for

Cited by 0SourcePDFScholar
2025

Adapting Precomputed Features for Efficient Graph Condensation

ICML 2025poster

Graph Neural Networks (GNNs) face significant computational challenges when handling large-scale graphs. To address this, Graph Condensation (GC) methods aim to compress large graphs into smaller, synthetic ones that are more manageable for GNN training. Recently, trajectory matching methods have sh…

2025

CateEA: Enhancing Entity Alignment via Implicit Category Supervision

COLING 2025main

Entity Alignment (EA) is essential for integrating Knowledge Graphs (KGs) by matching equivalent entities across diverse KGs. With the rise of multi-modal KGs, which emerged to better depict real-world KGs by integrating visual, textual, and structured data, Multi-Modal Entity Alignment (MMEA) has b…

Cited by 0SourcePDFScholar
2025

Effective and Efficient Masked Image Generation Models

ICML 2025poster

Although masked image generation models and masked diffusion models are designed with different motivations and objectives, we observe that they can be unified within a single framework. Building upon this insight, we carefully explore the design space of training and sampling, identifying key facto…

2025

Let Modalities Teach Each Other: Modal-Collaborative Knowledge Extraction and Fusion for Multimodal Knowledge Graph Completion

NAACL 2025findings

Multimodal knowledge graph completion (MKGC) aims to predict missing triples in MKGs using multimodal information. Recent research typically either extracts information from each modality separately to predict, then ensembles the predictions at the decision stage, or projects multiple modalities int…

Cited by 0SourcePDFScholar
2025

MIXVPR++: Enhanced Visual Place Recognition With Hierarchical-Region Feature-Mixer and Adaptive Gabor Texture Fuser

RA-L 2025

Visual Place Recognition (VPR) is crucial for various computer vision and robotics applications. Traditional VPR techniques relying on handcrafted features, have been enhanced by using Convolutional Neural Networks (CNNs). Recently, MixVPR has set new benchmarks in VPR by using advanced feature aggr

Cited by 2SourceScholar
2025

Modality-Independent Graph Neural Networks with Global Transformers for Multimodal Recommendation

AAAI 2025technical

Multimodal recommendation systems can learn users' preferences from existing user-item interactions as well as the semantics of multimodal data associated with items. Many existing methods model this through a multimodal user-item graph, approaching multimodal recommendation as a graph learning task…

2025

Robust 4D Radar-Aided Inertial Navigation for Aerial Vehicles

ICRA 2025

While LiDAR and cameras are becoming ubiquitous for unmanned aerial vehicles (UAVs) but can be ineffective in challenging environments, 4D millimeter-wave (MMW) radars that can provide robust 3D ranging and Doppler velocity measurements are less exploited for aerial navigation. In this paper, we dev

Cited by 0SourceScholar
2025

Synergizing LLMs with Global Label Propagation for Multimodal Fake News Detection

ACL 2025long

Large Language Models (LLMs) can assist multimodal fake news detection by predicting pseudo labels. However, LLM-generated pseudo labels alone demonstrate poor performance compared to traditional detection methods, making their effective integration non-trivial. In this paper, we propose Global Labe…

2025

Towards Emotion Co-regulation with LLM-powered Socially Assistive Robots: Integrating LLM Prompts and Robotic Behaviors to Support Parent-Neurodivergent Child Dyads

IROS 2025

Socially Assistive Robotics (SAR) has shown promise in supporting emotion regulation for neurodivergent children. Recently, there has been increasing interest in leveraging advanced technologies to assist parents in co-regulating emotions with their children. However, limited research has explored t

Cited by 2SourceScholar
2025

Understanding Individual Agent Importance in Multi-Agent System via Counterfactual Reasoning

AAAI 2025technical

Explaining multi-agent systems (MAS) is urgent as these systems become increasingly prevalent in various applications. Previous work has provided explanations for the actions or states of agents, yet falls short in understanding the blackboxed agent’s importance within a MAS and the overall team str…

Cited by 0SourcePDFScholar
2024

Class-Specific Semantic Generation and Reconstruction Learning for Open Set Recognition

IJCAI 2024poster

Open set recognition is a crucial research theme for open-environment machine learning. For this problem, a common solution is to learn compact representations of known classes and identify unknown samples by measuring deviations from these known classes. However, the aforementioned methods (1) lack…

2024

Consistency Training with Learnable Data Augmentation for Graph Anomaly Detection with Limited Supervision

ICLR 2024spotlight

Graph Anomaly Detection (GAD) has surfaced as a significant field of research, predominantly due to its substantial influence in production environments. Although existing approaches for node anomaly detection have shown effectiveness, they have yet to fully address two major challenges: operating i…

2024

Dual Rank-1 Tensor Attention Module for Convolutional Neural Networks

ICASSP 2024accepted

Channel-spatial attention mechanisms have been extensively investigated in computer vision. However, it is still a difficult problem that how to efficiently utilize global and local contextual information laid in a feature tensor to generate an accurate 3D attention map. This paper proposes a novel…

Cited by 1SourceScholar
2024

Efficient Saliency Encoding for Visual Place Recognition: Introducing the Lightweight Pooling-Centric Saliency-Aware VPR Method

RA-L 2024

The paper introduces a novel Visual Place Recognition (VPR) method called Lightweight Pooling-centric Saliency-aware VPR (LPS-VPR), a high-performance VPR method capable of exploiting saliency information without computational burden. The key contribution of the method is a pooling-based saliency en

Cited by 6SourceScholar
2024

Square-Root Inverse Filter-based GNSS-Visual-Inertial Navigation

ICRA 2024poster

While Global Navigation Satellite System (GNSS) is often used to provide global positioning if available, its intermittency and/or inaccuracy calls for fusion with other sensors. In this paper, we develop a novel GNSS-Visual-Inertial Navigation System (GVINS) that fuses visual, inertial, and raw GNS…

Cited by 0SourceScholar
2024

Tools Identification By On-Board Adaptation of Vision-and-Language Models

AAAI 2024technical

A robotic workshop assistant has been a long-standing grand challenge for robotics, speech, computer vision, and artificial intelligence (AI) research. We revisit the goal of visual identification of tools from human queries in the current era of Large Vision-and-Language models (like GPT-4). We fin…

Cited by 1SourcePDFScholar
2023

Tracking Multiple Deformable Objects in Egocentric Videos

CVPR 2023poster

Most existing multiple object tracking (MOT) methods that solely rely on appearance features struggle in tracking highly deformable objects. Other MOT methods that use motion clues to associate identities across frames have difficulty handling egocentric videos effectively or efficiently. In this wo…

Cited by 14SourcePDFScholar
2022

1D-LRF Aided Visual-Inertial Odometry for High-Altitude MAV Flight

ICRA 2022poster

This paper addresses the problem of visual-inertial odometry (VIO) with a downward facing monocular camera when a micro aerial vehicle (MAV) flying at high altitude (over 100 meters). It is important to note that large scene depth causes visual motion constraints significantly less informative than…

Cited by 10SourceScholar
2021

Dialogue Disentanglement in Software Engineering: How Far are We?

IJCAI 2021poster

Despite the valuable information contained in software chat messages, disentangling them into distinct conversations is an essential prerequisite for any in-depth analyses that utilize this information. To provide a better understanding of the current state-of-the-art, we evaluate five popular dialo…

2019

Sell-corpus: an Open Source Multiple Accented Chinese-english Speech Corpus for L2 English Learning Assessment

ICASSP 2019accepted

We present SELL-CORPUS, a multiple accented speech corpus for L2 English learning in China, aiming at the potential research of multiple accented acoustic model, mispronunciation detection and pronunciation assessment for future nationwide oral English tests. Our corpus contains 31.6 hour speech rec…

Cited by 0SourceScholar