← Search

Tong Li

29 accepted papers

2026

Learning Attribute–Affordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding

ICML 2026poster

This paper pays attention to open-vocabulary 3D object affordance grounding (OVAG), which aims to localize affordance regions on 3D objects by leveraging interaction images or textual instructions. Most existing methods treat interaction images as sources of external affordance knowledge and align t…

Cited by 0SourceScholar
2026

PANKRAG: ENHANCING GRAPH RETRIEVAL VIA GLOBALLY AWARE QUERY RESOLUTION AND DEPENDENCY-AWARE RERANKING MECHANISM

ICASSP 2026poster

Recent graph-based RAG approaches leverage knowledge graphs by extracting entities from a query to fetch their associated relationships and metadata. However, relying solely on entity extraction often results in the misinterpretation or omission of latent critical information and relationships. This…

Cited by 0SourcePDFScholar
2026

Raise One and Infer Three: Toward Reasoning- and Memory-Augmented Diffusion Policy Generalization

IJCAI 2026

Diffusion policy has shown impressive performance in robotic manipulation tasks while struggling with out-of-distribution shifts and limited demonstrations. Recent advances primarily focus on improving geometric or perceptual representations for diffusion policy. However, these approaches rely heavi

Cited by 0Scholar
2025

Adaptive Dual Guidance Knowledge Distillation

AAAI 2025technical

Knowledge distillation (KD) aims to improve the performance of lightweight student networks under the guidance of pre-trained teachers. However, the large capacity gap between teachers and students limits the distillation gains. Previous methods addressing this problem have two weaknesses. First, mo…

Cited by 0SourcePDFScholar
2025

Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation

EMNLP 2025

Suicide remains a major global mental health challenge, and early intervention hinges on recognizing signs of suicidal ideation. In private conversations, such ideation is often expressed in subtle or conflicted ways, making detection especially difficult. Existing data sets are mainly based on publ

Cited by 0SourcePDFScholar
2025

FuzzyMIL: Decoupling Pathological Phenotypes through Deep Fuzzy Clustering for Efficient Whole Slide Image Analysis

ICASSP 2025accepted

In Multiple Instance Learning (MIL) for Whole Slide Image (WSI) analysis, attention mechanisms are often employed to weigh the importance of different instances. However, global attention may lead to feature homogenization and overlook tissue differences. Adding local attention can capture these var…

Cited by 0SourceScholar
2025

Multimodal Integrated Prediction and Decision-making with Adaptive Interaction Modality Explorations

IROS 2025

Navigating dense and dynamic environments poses a significant challenge for autonomous driving systems, owing to the intricate nature of multimodal interaction, wherein the actions of various traffic participants and the autonomous vehicle are complex and implicitly coupled. In this paper, we propos

Cited by 7SourcecodeScholar
2025

PaSTS: Parameter-affined Seasonal-Trend Synthesis for Multi-dimensional Long-Term Time Series Forecasting within LLM

ICASSP 2025accepted

Large Language Models (LLMs) have demonstrated remarkable performance across various domains, showcasing significant potential for long-term time series forecasting (LTSF), and consequently attracting substantial research interest. In LTSF, temporal decomposition has been widely adopted in existing…

Cited by 0SourceScholar
2025

Positive2Negative: Breaking the Information-Lossy Barrier in Self-Supervised Single Image Denoising

CVPR 2025poster

Image denoising enhances image quality, serving as a foundational technique across various computational photography applications. The obstacle to clean image acquisition in real scenarios necessitates the development of self-supervised image denoising methods only depending on noisy images, especia…

2025

Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models

ICLR 2025poster

Backdoor unalignment attacks against Large Language Models (LLMs) enable the stealthy compromise of safety alignment using a hidden trigger while evading normal safety auditing. These attacks pose significant threats to the applications of LLMs in the real-world Large Language Model as a Service (LL…

2025

Prompt-Guided Internal States for Hallucination Detection of Large Language Models

ACL 2025long

Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of tasks in different domains. However, they sometimes generate responses that are logically coherent but factually incorrect or misleading, which is known as LLM hallucinations. Data-driven supervised methods tr…

2024

BadActs: A Universal Backdoor Defense in the Activation Space

ACL 2024findings

Backdoor attacks pose an increasingly severe security threat to Deep Neural Networks (DNNs) during their development stage. In response, backdoor sample purification has emerged as a promising defense mechanism, aiming to eliminate backdoor triggers while preserving the integrity of the clean conten…

2024

Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation

CVPR 2024highlight

Recently an audio-visual segmentation (AVS) task has been introduced aiming to group pixels with sounding objects within a given video. This task necessitates a first-ever audio-driven pixel-level understanding of the scene posing significant challenges. In this paper we propose an innovative audio-…

2024

CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation

CVPR 2024poster

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion. The applications of listener agent generation in virtual interaction have promoted many works achieving diverse and fine-grained…

Cited by 4SourcePDFScholar
2024

MASTER: Market-Guided Stock Transformer for Stock Price Forecasting

AAAI 2024technical

Stock price forecasting has remained an extremely challenging problem for many decades due to the high volatility of the stock market. Recent efforts have been devoted to modeling complex stock correlations toward joint stock price forecasting. Existing works share a common neural architecture that…

2024

Theoretical Modeling and Bio-inspired Trajectory Optimization of A Multiple-locomotion Origami Robot

IROS 2024poster

Recent research on mobile robots has focused on increasing their adaptability to unpredictable and unstructured environments using soft materials and structures. However, the determination of key design parameters and control over these compliant robots are predominantly iterated through experiments…

Cited by 1SourceScholar
2024

Using Adaptive Bandit Experiments to Increase and Investigate Engagement in Mental Health

AAAI 2024technical

Digital mental health (DMH) interventions, such as text-message-based lessons and activities, offer immense potential for accessible mental health support. While these interventions can be effective, real-world experimental testing can further enhance their design and impact. Adaptive experimentatio…

2024

You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception

ECCV 2024poster

"Human-centric perception (detection, segmentation, pose estimation, and attribute analysis) is a long-standing problem for computer vision. This paper introduces a unified and versatile framework (HQNet) for single-stage multi-person multi-task human-centric perception (HCP). Our approach centers o…

2023

A Sequence-to-Sequence&Set Model for Text-to-Table Generation

ACL 2023findings

Recently, the text-to-table generation task has attracted increasing attention due to its wide applications. In this aspect, the dominant model formalizes this task as a sequence-to-sequence generation task and serializes each table into a token sequence during training by concatenating all rows in…

2023

Contrastive Learning at the Relation and Event Level for Rumor Detection

ICASSP 2023accepted

Existing studies for rumor detection rely heavily on a large number of labeled data to operate in a fully-supervised manner. However, manual data annotation in realistic cases is very expensive and time-consuming. In this paper, we propose a novel self-supervised Relation-Event based Contrastive Lea…

Cited by 0SourceScholar
2023

Single Depth-image 3D Reflection Symmetry and Shape Prediction

ICCV 2023poster

In this paper, we present Iterative Symmetry Completion Network (ISCNet), a single depth-image shape completion method that exploits reflective symmetry cues to obtain more detailed shapes. The efficacy of single depth-image shape completion methods is often sensitive to the accuracy of the symmetry…

Cited by 7PDFScholar
2023

Towards Fine-Grained Explainability for Heterogeneous Graph Neural Network

AAAI 2023technical

Heterogeneous graph neural networks (HGNs) are prominent approaches to node classification tasks on heterogeneous graphs. Despite the superior performance, insights about the predictions made from HGNs are obscure to humans. Existing explainability techniques are mainly proposed for GNNs on homogene…

2022

Estimation of Upper Limb Kinematics with a Magnetometer-Free Egocentric Visual-Inertial System

ICRA 2022poster

Most human activities in daily living or professional work rely on upper body motion. Measuring upper body motion is essential for many applications such as health evaluation, rehabilitation, human power augmentation, skill transferring, etc. Computer vision-based systems have been widely used to di…

Cited by 5SourceScholar
2022

UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

EMNLP 2022main

Geometry problem solving is a well-recognized testbed for evaluating the high-level multi-modal reasoning capability of deep models. In most existing works, two main geometry problems: calculation and proving, are usually treated as two specific tasks, hindering a deep model to unify its reasoning c…

2021

C-Learning: Horizon-Aware Cumulative Accessibility Estimation

ICLR 2021poster

Multi-goal reaching is an important problem in reinforcement learning needed to achieve algorithmic generalization. Despite recent advances in this field, current algorithms suffer from three major challenges: high sample complexity, learning only a single way of reaching the goals, and difficultie…

2021

Self-Supervised Detection of Contextual Synonyms in a Multi-Class Setting: Phenotype Annotation Use Case

EMNLP 2021main

Contextualised word embeddings is a powerful tool to detect contextual synonyms. However, most of the current state-of-the-art (SOTA) deep learning concept extraction methods remain supervised and underexploit the potential of the context. In this paper, we propose a self-supervised pre-training app…

Cited by 17SourcePDFScholar
2020

Multi-View Joint Graph Representation Learning for Urban Region Embedding

IJCAI 2020poster

The increasing amount of urban data enable us to investigate urban dynamics, assist urban planning, and eventually, make our cities more livable and sustainable. In this paper, we focus on learning an embedding space from urban data for urban regions. For the first time, we propose a multi-view join…