← Search

Lei Shi

56 accepted papers

2026

Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal Retrieval

AAAI 2026technical

Cross-modal hashing has emerged as a pivotal solution for efficient retrieval across diverse modalities, such as images and texts, by mapping them into compact binary hash spaces. However, in real-world scenarios, the modalities data is often missing or misaligned. Existing methods are most rely on

Cited by 0SourcePDFScholar
2026

Constraint-Augmented Mongolian-Chinese Neural Machine Translation Based on Dynamic Feedback Alignment (Student Abstract)

AAAI 2026technical

The scarcity of parallel corpora for Mongolian and Chinese constrains the performance of Mongolian-Chinese neural machine translation (NMT), particularly manifesting in inadequate accuracy in translating specialized terminology. To address this limitation, this study adopts a lexically constrained a

Cited by 0SourcePDFScholar
2026

MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative Recommendation

AAAI 2026technical

Generative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language M

Cited by 0SourcePDFScholar
2026

Whole-Body Impedance Coordinative Control for a Wheel-Legged Robot on Uncertain Terrain

RA-L 2026

This article proposes a whole-body impedance coordinative control framework for a wheel-legged humanoid robot to achieve adaptability on complex terrains while maintaining the robot's upper body stability. The framework contains a bi-level control strategy. The outer level is a variable-damping impe

Cited by 0SourceScholar
2025

A Fairness-Oriented Control Framework for Safety-Critical Multi-Robot Systems: Alternative Authority Control

ICRA 2025

This paper proposes a fair control framework for multi-robot systems, which integrates the newly introduced Alternative Authority Control (AAC) and Flexible Control Barrier Function (F-CBF). Control authority refers to a single robot which can plan its trajectory while considering others as moving o

Cited by 1SourceScholar
2025

ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrieval

EMNLP 2025

Many contemporary data-driven research efforts in the natural sciences, such as chemistry and materials science, require large-scale, high-performance entity recognition from scientific datasets. Large language models (LLMs) have increasingly been adopted to solve the entity recognition task, with t

Cited by 0SourcePDFScholar
2025

Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models

EMNLP 2025

Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Understanding these internal mechanisms is critical - not only to move beyond surface-level performance, but also for mode

2025

CFPT: Empowering Time Series Forecasting through Cross-Frequency Interaction and Periodic-Aware Timestamp Modeling

ICML 2025poster

Long-term time series forecasting has been widely studied, yet two aspects remain insufficiently explored: the interaction learning between different frequency components and the exploitation of periodic characteristics inherent in timestamps. To address the above issues, we propose **CFPT**, a nov…

2025

Can Classic GNNs Be Strong Baselines for Graph-level Tasks? Simple Architectures Meet Excellence

ICML 2025poster

Message-passing Graph Neural Networks (GNNs) are often criticized for their limited expressiveness, issues like over-smoothing and over-squashing, and challenges in capturing long-range dependencies. Conversely, Graph Transformers (GTs) are regarded as superior due to their employment of global atte…

2025

Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval

NeurIPS 2025poster

The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existi…

Cited by 0SourceScholar
2025

EVICheck: Evidence-Driven Independent Reasoning and Combined Verification Method for Fact-Checking

IJCAI 2025

Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have demonstrated significant potential in automated fact-checking. However, existing methods face limitations in insufficient evidence utilization and lack of explicit verification criteria. Specifically, these approaches aggrega

2025

FasterGold-DETR: An Efficient End-to-End Fire Detection Model via Gather-and-Distribute Mechanism

ICASSP 2025accepted

Fire detection technology based on deep learning methods has become a prevalent practice. However, the performance of current YOLO-based detection models is limited by NMS, and DETR-based detection models struggle with real-time performance. To address these challenges, a new fire detection model, F…

Cited by 0SourceScholar
2025

IWRN:A Robust Blind Watermarking Method for Artwork Image Copyright Protection Against Noise Attack

AAAI 2025technical

Adding imperceptible watermarks to artwork images, such as paintings and photographs, can effectively safeguard the copyright of these images without compromising their usability. However, existing blind watermarking techniques encounter two major challenges in addressing this task: imperceptibility…

2025

Learning to Rank for In-Context Example Retrieval

NeurIPS 2025poster

Recent advances in retrieval-based in-context learning (ICL) train the retriever using a classification objective, which categorizes in-context examples (ICEs) into the most useful and the rest based on absolute scores. However, during inference, ICEs are retrieved by score ranking rather than class…

Cited by 0SourcecodeScholar
2025

Leveraging semantic similarity for experimentation with AI-generated treatments

NeurIPS 2025poster

Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main methodological challenge in this setting is representing these high-dimensional treatments without losing their semantic m…

Cited by 0SourceScholar
2025

Leveraging the Dual Capabilities of LLM: LLM-Enhanced Text Mapping Model for Personality Detection

AAAI 2025technical

Personality detection aims to deduce a user’s personality from their published posts. The goal of this task is to map posts to specific personality types. Existing methods encode post information to obtain user vectors, which are then mapped to personality labels. However, existing methods face two…

2025

Node Identifiers: Compact, Discrete Representations for Efficient Graph Learning

ICLR 2025poster

We present a novel end-to-end framework that generates highly compact (typically 6-15 dimensions), discrete (int4 type), and interpretable node representations—termed node identifiers (node IDs)—to tackle inference challenges on large-scale graphs. By employing vector quantization, we compress conti…

2025

OSTAR: Optimized Statistical Text-classifier with Adversarial Resistance

NeurIPS 2025poster

The advancements in generative models and the real-world attack of machine-generated text(MGT) create a demand for more robust detection methods. The existing MGT detection methods for adversarial environments primarily consist of manually designed statistical-based methods and fine-tuned classifi…

Cited by 0SourcecodeScholar
2025

Online Iterative Self-Alignment for Radiology Report Generation

ACL 2025long

Radiology Report Generation (RRG) is an important research topic for relieving radiologists’ heavy workload. Existing RRG models mainly rely on supervised fine-tuning (SFT) based on different model architectures using data pairs of radiological images and corresponding radiologist-annotated reports.…

Cited by 0SourcePDFScholar
2025

Radiology Report Generation via Multi-objective Preference Optimization

AAAI 2025technical

Automatic Radiology Report Generation (RRG) is an important topic for alleviating the substantial workload of radiologists. Existing RRG approaches rely on supervised regression based on different architectures or additional knowledge injection, while the generated report may not align optimally wit…

Cited by 2SourcePDFScholar
2025

SmartEraser: Remove Anything from Images using Masked-Region Guidance

CVPR 2025poster

Object removal has so far been dominated by the mask-and-inpaint paradigm, where the masked region is excluded from the input, leaving models relying on unmasked areas to inpaint the missing region. However, this approach lacks contextual information for the masked area, often resulting in unstable…

Cited by 2SourcePDFScholar
2025

StrucFormer: Structural Prior Guided Transformer for Mobile Crowdsensing Data Inference

ICASSP 2025accepted

The inherent constraint of the "human-in-the-loop" sensing mechanism, imposes mobile crowdsensing with high dynamics and uncertainty, ultimately leading to the issue of incomplete data collection. Current data inference solutions in mobile crowdsensing can be broadly categorized as low-rank models a…

Cited by 0SourceScholar
2025

THGNets: Constrained Temporal Hypergraphs and Graph Neural Networks in Hyperbolic Space for Information Diffusion Prediction

AAAI 2025technical

Information diffusion prediction aims to predict the next infected user in the information diffusion, which is a critical task to understand how information spreads on social platforms. Existing methods mainly focus on the sequences or topology structure in euclidean space. However, they fail to suf…

Cited by 0SourcePDFScholar
2025

Unaligned Message-Passing and Contextualized-Pretraining for Robust Geo-Entity Resolution

AAAI 2025technical

Geo-entity resolution involves linking records that refer to the same entities across different spatial datasets, which underpins location-based services. Given the varying quality of geo-data, this task is known to be challenging, as directly comparing the semantic-centric representations of two en…

2024

An End-To-End Graph Attention Network Hashing for Cross-Modal Retrieval

NeurIPS 2024poster

Due to its low storage cost and fast search speed, cross-modal retrieval based on hashing has attracted widespread attention and is widely used in real-world applications of social media search. However, most existing hashing methods are often limited by uncomprehensive feature representations and s…

Cited by 1SourcePDFScholar
2024

Classic GNNs are Strong Baselines: Reassessing GNNs for Node Classification

NeurIPS 2024poster

Graph Transformers (GTs) have recently emerged as popular alternatives to traditional message-passing Graph Neural Networks (GNNs), due to their theoretically superior expressiveness and impressive performance reported on standard node classification benchmarks, often significantly outperforming GNN…

2024

Enhancing Graph Transformers with Hierarchical Distance Structural Encoding

NeurIPS 2024poster

Graph transformers need strong inductive biases to derive meaningful attention scores. Yet, current methods often fall short in capturing longer ranges, hierarchical structures, or community structures, which are common in various graphs such as molecules, social networks, and citation networks. Thi…

2024

LEGENT: Open Platform for Embodied Agents

ACL 2024system demonstrations

Despite advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), their integration into language-grounded, human-like embodied agents remains incomplete, hindering complex real-life task performance in 3D environments. Existing integrations often feature limited open-sourcing…

Cited by 9SourcePDFScholar
2024

Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition

ACL 2024long

Recent work on dialogue-based collaborative plan acquisition (CPA) has suggested that Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge. Although ToM was claimed to be important for effective collaboration, its real impact on…

Cited by 6SourcePDFScholar
2024

Novel Lightweight Lower Limb Exoskeleton Design for Single-Motor Sequential Assistance of Knee & Ankle Joints in Real World

RA-L 2024

In this article, we introduce a lightweight lower limb exoskeleton that provides auxiliary torque to both the ankle and knee joints during the stance phase of gait in a real-world environment, using one quasi-direct-drive (QDD) motor .This lightweight exoskeleton incorporates a novel driving mechani

Cited by 17SourceScholar
2024

Self-Supervised Representation Learning with Meta Comprehensive Regularization

AAAI 2024technical

Self-Supervised Learning (SSL) methods harness the concept of semantic invariance by utilizing data augmentation strategies to produce similar representations for different deformations of the same input. Essentially, the model captures the shared information among multiple augmented views of sample…

Cited by 6SourcePDFScholar
2024

Structural Information Enhanced Graph Representation for Link Prediction

AAAI 2024technical

Link prediction is a fundamental task of graph machine learning, and Graph Neural Network (GNN) based methods have become the mainstream approach due to their good performance. However, the typical practice learns node representations through neighborhood aggregation, lacking awareness of the struct…

Cited by 5SourcePDFScholar
2024

Using Surrogates in Covariate-adjusted Response-adaptive Randomization Experiments with Delayed Outcomes

NeurIPS 2024poster

Covariate-adjusted response-adaptive randomization (CARA) designs are gaining increasing attention. These designs combine the advantages of randomized experiments with the ability to adaptively revise treatment allocations based on data collected across multiple stages, enhancing estimation efficien…

Cited by 1SourcePDFScholar
2023

Differential Dynamic Programming based Hybrid Manipulation Strategy for Dynamic Grasping

ICRA 2023poster

To fully explore the potential of robots for dexterous manipulation, this paper presents a whole dynamic grasping process to achieve fluent grasping of a target object by the robot end-effector. The process starts from the phase of approaching the object over the phases of colliding with the object…

Cited by 7SourceScholar
2023

Improving Self-supervised Molecular Representation Learning using Persistent Homology

NeurIPS 2023poster

Self-supervised learning (SSL) has great potential for molecular representation learning given the complexity of molecular graphs, the large amounts of unlabelled data available, the considerable cost of obtaining labels experimentally, and the hence often only small training datasets. The importanc…

2022

Communication-Efficient Topologies for Decentralized Learning with $O(1)$ Consensus Rate

NeurIPS 2022accept

Decentralized optimization is an emerging paradigm in distributed learning in which agents achieve network-wide solutions by peer-to-peer communication without the central server. Since communication tends to be slower than computation, when each agent communicates with only a few neighboring agent…

2022

Exact-likelihood User Intention Estimation for Scene-compliant Shared-control Navigation

ICRA 2022poster

A predictive model for mobility systems capable of understanding the trajectory a user intends to follow in the environment is proposed. Understanding user intention is paramount for any shared-control navigation strategy between a user and an active robotic agent. Equally important however is being…

Cited by 3SourceScholar
2021

AdaSGN: Adapting Joint Number and Model Size for Efficient Skeleton-Based Action Recognition

ICCV 2021poster

Existing methods for skeleton-based action recognition mainly focus on improving the recognition accuracy, whereas the efficiency of the model is rarely considered. Recently, there are some works trying to speed up the skeleton modeling by designing light-weight modules. However, in addition to the…

Cited by 67PDFcodeScholar
2021

Multi-modal Scene-compliant User Intention Estimation in Navigation

IROS 2021poster

A multi-modal framework to generate user intention distributions when operating a mobile vehicle is proposed in this work. The model learns from past observed trajectories and leverages traversability information derived from the visual surroundings to produce a set of future trajectories, suitable…

Cited by 8SourceScholar
2020

Decoupling GCN with DropGraph Module for Skeleton-Based Action Recognition

ECCV 2020poster

In skeleton-based action recognition, graph convolutional networks (GCNs) have achieved remarkable success. Nevertheless, how to efficiently model the spatial-temporal skeleton graph without introducing extra computation burden is a challenging problem for industrial deployment. In this paper, we re…

2020

Multi-Layer Content Interaction Through Quaternion Product for Visual Question Answering

ICASSP 2020accepted

Multi-modality fusion technologies have greatly improved the performance of neural network-based Video Description/Caption, Visual Question Answering (VQA) and Audio Visual Scene-aware Dialog (AVSD) over the recent years. Most previous approaches only explore the last layers of multiple layer featur…

Cited by 0SourceScholar
2019

Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition

CVPR 2019poster

In skeleton-based action recognition, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have achieved remarkable performance. However, in existing GCN-based methods, the topology of the graph is set manually, and it is fixed over all layers and input…

Cited by 2123PDFcodeScholar
2017

Real-time 3D human tracking for mobile robots with multisensors

ICRA 2017poster

Acquiring the accurate 3-D position of a target person around a robot provides fundamental and valuable information that is applicable to a wide range of robotic tasks, including home service, navigation and entertainment. This paper presents a real-time robotic 3-D human tracking system which combi…

Cited by 34SourceScholar
2016

Constrained sampling of 2.5D probabilistic maps for augmented inference

IROS 2016poster

This work exploits modeling spatial correlation in 2.5D data using Gaussian Processes (GPs), and produces constrained sampling realizations on these models to improve certainty in the predictions by means of integrating additional sparse information. Data organized in 2.5D such as elevation and thic…

Cited by 3SourceScholar
2016

Understand scene categories by objects: A semantic regularized scene classifier using Convolutional Neural Networks

ICRA 2016

Scene classification is a fundamental perception task for environmental understanding in today's robotics. In this paper, we have attempted to exploit the use of popular machine learning technique of deep learning to enhance scene understanding, particularly in robotics applications. As scene images

Cited by 109SourceScholar