← Search

Xin Su

21 accepted papers

2026

Distill-SynthKG: Distilling Knowledge Graph Synthesis Workflow for Improved Coverage and Efficiency

ICLR 2026poster

Document-level knowledge graph (KG) construction faces a fundamental scaling challenge: existing methods either rely on expensive large language models (LLMs), making them economically unviable for large-scale corpora, or employ smaller models that produce incomplete and inconsistent graphs. We iden…

Cited by 0SourceScholar
2026

HPGS-SLAM: Hybrid Point-Guided Dense Visual SLAM With Online Mapping via Gaussian Splatting

RA-L 2026

In this letter, we introduce HPGS-SLAM, a real-time RGB-D SLAM system guided by hybrid point features (combining traditional and learned point features), enabling high-precision tracking and online dense mapping with photorealistic reconstruction. HPGS-SLAM consists of two main components: (1) a lig

Cited by 1SourceScholar
2025

CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction

ACL 2025long

The paper focuses on the interpretability of Grammatical Error Correction (GEC) evaluation metrics, which received little attention in previous studies. To bridge the gap, we introduce **CLEME2.0**, a reference-based metric describing four fundamental aspects of GEC systems: hit-correction, wrong-co…

2025

CMIF-VIO: A Novel Cross Modal Interaction Framework for Visual Inertial Odometry

RA-L 2025

Visual Inertial Odometry (VIO) estimates predicted trajectories through self motion. With the popularization of artificial intelligence, deep learning-based VIO methods have shown better performance than traditional geometry-based VIO methods. However, in deep learning methods, how to better achieve

Cited by 5SourceScholar
2025

DAST: Context-Aware Compression in LLMs via Dynamic Allocation of Soft Tokens

ACL 2025finding

Large Language Models (LLMs) face computational inefficiencies and redundant processing when handling long context inputs, prompting a focus on compression techniques. While existing semantic vector-based compression methods achieve promising performance, these methods fail to account for the intrin…

Cited by 0SourcePDFScholar
2025

Dike: Enhancing Fairness and Efficiency in GPU Clusters for Deep Learning

ICASSP 2025accepted

The advent of deep learning (DL) has transformed signal interpretation, enabling more efficient solutions to complex signal processing problems. DL workloads in signal processing typically share the computational resources of GPU clusters. However, the unpredictable nature of the duration of the DL…

Cited by 0SourceScholar
2025

HypCAD: Geometry-Enhanced Hyperbolic Contrastive Learning for CAD Model Retrieval

ICASSP 2025accepted

Retrieving CAD models for real-world object scans enhances object-level mapping, providing a nuanced spatial understanding crucial for precise interactions in robotics or mixed reality. Commonly, CAD model retrieval is performed by matching features learned in Euclidean space. However, learning disc…

Cited by 0SourceScholar
2025

Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction

ICASSP 2025accepted

Chinese grammatical error correction (CGEC) aims to detect and correct errors in the input Chinese sentences. Recently, Pre-trained Language Models (PLMS) have been employed to improve the performance. However, current approaches ignore that correction difficulty varies across different instances an…

Cited by 0SourceScholar
2025

SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs

ICML 2025oral

Multimodal retrieval-augmented generation (RAG) plays a crucial role in domains such as knowledge-based visual question answering (KB-VQA), where models should effectively integrate additional knowledge to generate a response. However, existing vision and language models (VLMs) are not inherently de…

Cited by 5SourcePDFScholar
2024

HEGN: Hierarchical Equivariant Graph Neural Network for 9DoF Point Cloud Registration

ICRA 2024poster

Given its wide application in robotics, point cloud registration is a widely researched topic. Conventional methods aim to find a rotation and translation that align two point clouds in 6 degrees of freedom (DoF). However, certain tasks in robotics, such as category-level pose estimation, involve no…

Cited by 1SourceScholar
2024

HPF-SLAM: An Efficient Visual SLAM System Leveraging Hybrid Point Features

ICRA 2024poster

Visual SLAM is an essential tool in diverse applications such as robot perception and extended reality, where feature-based methods are prevalent due to their accuracy and robustness. However, existing methods employ either hand-crafted or solely learnable point features and are thus limited by the…

Cited by 1SourceScholar
2024

Semi-Structured Chain-of-Thought: Integrating Multiple Sources of Knowledge for Improved Language Model Reasoning

NAACL 2024long

An important open question in the use of large language models for knowledge-intensive tasks is how to effectively integrate knowledge from three sources: the model’s parametric memory, external structured knowledge, and external unstructured knowledge. Most existing prompting methods either rely on…

2023

Fusing Temporal Graphs into Transformers for Time-Sensitive Question Answering

EMNLP 2023long findings

Answering time-sensitive questions from long documents requires temporal reasoning over the times in questions and documents. An important open question is whether large language models can perform such reasoning solely using a provided text document, or whether they can benefit from additional temp…

Cited by 0SourceScholar
2023

Modeling Action Spatiotemporal Relationships Using Graph-Based Class-Level Attention Network for Long-Term Action Detection

IROS 2023poster

In recent years, Action Detection has become an active research topic in various fields such as human-robot interaction and assistive robots. Most of the previous methods in this field focus on temporally processing the action representation, without considering the dependencies among the action cla…

Cited by 6SourceScholar
2021

DymSLAM: 4D Dynamic Scene Reconstruction Based on Geometrical Motion Segmentation

RA-L 2021

Most SLAM (Simultaneous Localization and Mapping) algorithms are based on the assumption that the scene is static. However, in practice, most real scenes usually contain moving objects. In this letter, we introduce DymSLAM, a dynamic stereo visual SLAM system being capable of reconstructing a 4D (3D

Cited by 50SourceScholar
2020

Adaptability Preserving Domain Decomposition for Stabilizing Sim2Real Reinforcement Learning

IROS 2020poster

In sim-to-real transfer of Reinforcement Learning (RL) policies for robot tasks, Domain Randomization (DR) is a widely used technique for improving adaptability. However, in DR there is a conflict between adaptability and training stability, and heavy DR tends to result in instability or even failur…

Cited by 6SourceScholar
2018

Graph-based Transforms for Predictive Light Field Compression based on Super-Pixels

ICASSP 2018accepted

In this paper, we explore the use of graph-based transforms to capture correlation in light fields. We consider a scheme in which view synthesis is used as a first step to exploit inter-view correlation. Local graph-based transforms (GT) are then considered for energy compaction of the residue signa…

Cited by 0SourceScholar