← Search

Xingyu Li

37 accepted papers

2026

AI in the Wild: A Meta-Analytic Evaluation of Depression Detection from Social Media Data

AAAI 2026technical

As AI moves into high-stakes, human-centered settings, we still lack clear evidence on when and why these systems succeed or fail. This meta-analysis synthesizes all empirical studies published between 2022 and 2025 that use social-media data to predict depression, quantifying pooled accuracy and te

Cited by 0SourcePDFScholar
2026

Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Reasoning Large Language Models

ICLR 2026poster

Reasoning Large Language Models (RLLMs) excelling in complex tasks present unique challenges for digital watermarking, as existing methods often disrupt logical coherence or incur high computational costs. Token-based watermarking techniques can corrupt the reasoning flow by applying pseudo-random…

Cited by 0SourceScholar
2026

FELP: Fast and Effective Autonomous Flight on Large-Scale and Cluttered Environments Based on Unified Linear Parametric Map

ICRA 2026poster

Current AAV autonomous flights exhibit efficient performance in both indoor and field environments. However, they often face significant challenges in large-scale and cluttered environments, where the vast amount of captured data can lead to computation and storage bottlenecks. Additionally, the existin…

Cited by 0SourceScholar
2026

From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning

AAAI 2026technical

LiDAR-based 3D object detection models often struggle to generalize to real-world environments due to limited object diversity in existing datasets. To tackle it, we introduce the first generalized cross-domain few-shot (GCFS) task in 3D object detection, aiming to adapt a source-pretrained model to

Cited by 0SourcePDFScholar
2026

LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation

AAAI 2026technical

As large language models (LLMs) scale, their inference incurs substantial computational resources, exposing them to energy-latency attacks, where crafted prompts induce high energy and latency cost. Existing attack methods aim to prolong output by delaying the generation of termination symbols. Howe

Cited by 0SourcePDFScholar
2026

ModalPatch: A Plug-And-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop

ICRA 2026poster

Multi-modal 3D object detection is pivotal for autonomous driving, integrating complementary sensors like LiDAR and cameras. However, its real-world reliability is challenged by transient data interruptions and missing, where modalities can momentarily drop due to hardware glitches, adverse weather,…

2026

SynerDetect: Hierarchical Synergistic Learning for Generalizable AI-Generated Image Detection

AAAI 2026technical

The rapid advancement of generative models, which produce increasingly realistic synthetic images, urgently demands robust and generalizable detection methods. Consequently, research has largely pivoted to leveraging large-scale Vision Foundation Models (VFMs) for enhanced generalization. However, e

Cited by 0SourcePDFScholar
2025

Accelerating Convergence in Bounding Box Regression with a Refined IoU Loss Function

ICASSP 2025accepted

Bounding box regression (BBR) is a critical component in object detection, significantly influencing the accuracy of object localization. However, existing Intersection over Union (IoU)-based loss functions encounter two primary challenges: (i) The penalty factor configuration often results in the e…

Cited by 0SourceScholar
2025

Cockroach's Turning Strategy Enhanced Hexapod Robot with Flexible Torso

IROS 2025

The design and control of hexapod robots have become an active research field due to the ability to achieve adaptive and stable multi-terrain locomotion. However, existing hexapod robots focus on the integration of flexible pitch joints to enhance their obstacle-crossing and slope-climbing abilities

Cited by 0SourceScholar
2025

FELP:Fast and Effective Autonomous Flight on Large-Scale and Cluttered Environments Based on Unified Linear Parametric Map

RA-L 2025

Current UAV autonomous flights exhibit efficient performance in both indoor and field environments. However, they often face significant challenges in large-scale and cluttered environments, where the vast amount of captured data can lead to computation and storage bottlenecks. Additionally, the exi

Cited by 1SourceScholar
2025

GeoSafe: A Unified Unconstrained Multi-DOF Optimization Framework for Multi-UAV Cooperative Hoisting and Obstacle Avoidance

IROS 2025

In warehouse logistics and post-disaster rescue, multi-UAV payload transport must navigate tight spaces, such as 1.2m × 0.8m aisles and collapsed pipelines as narrow as 0.6m. Traditional four-DOF (translation and scaling) trajectory planning struggles under such constraints. To overcome this, we pro

Cited by 1SourceScholar
2025

InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling

ICLR 2025poster

Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often produce results lacking realism and fidelity. In this work, we introduce *InterMask*, a novel framework for generating human interact…

Cited by 5SourcePDFScholar
2025

Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach

ACL 2025long

Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals, and is essential for understanding multimodal content. In the era of rapidly growing multimodal content and social media, MCR is particularly crucial…

2025

Real-Time Occupancy Grid Mapping Using RMM on Large-scale and Unstructured Environments

IROS 2025

Occupancy mapping is crucial for distinguishing between known and unknown regions, which plays a significant role in the autonomous exploration of unmanned aerial vehicles (UAVs). However, the construction of high-quality maps is still a challenge. The challenge comes from the following factors. The

Cited by 0SourceScholar
2024

Adaptive Confidence Multi-View Hashing for Multimedia Retrieval

ICASSP 2024accepted

The multi-view hash method converts heterogeneous data from multiple views into binary hash codes, which is one of the critical technologies in multimedia retrieval. However, the current methods mainly explore the complementarity among multiple views while lacking confidence in learning and fusion.…

Cited by 0SourceScholar
2024

Autonomous Blood Suction for Robot-Assisted Surgery: A Sim-to-Real Reinforcement Learning Approach

RA-L 2024

Recent applications of deep reinforcement learning (DRL) in surgical autonomy have shown promising results in automating various surgical sub-tasks. While most of these studies consider the rigid and soft body dynamics in the surgery such as tissue deformation, only a few have investigated the situa

Cited by 16SourceScholar
2024

EMG-Based Intention Detection Using Deep Learning for Shared Control in Upper-Limb Assistive Exoskeletons

RA-L 2024

In the field of human-robot interaction, surface electromyography (sEMG) provides a valuable tool for measuring active muscular effort. While numerous studies have investigated real-time control of upper extremity exoskeletons based on user intention and task-specific movements, the prediction of bo

Cited by 53SourceScholar
2024

Evaluating Gait Symmetry with a Smart Robotic Walker: A Novel Approach to Mobility Assessment

IROS 2024poster

Gait asymmetry, a consequence of various neurological or physical conditions such as aging and stroke, detrimentally impacts bipedal locomotion, causing biomechanical alterations, increasing the risk of falls and reducing quality of life. Addressing this critical issue, this paper introduces a novel…

Cited by 3SourcecodeScholar
2024

Federated Text-driven Prompt Generation for Vision-Language Models

ICLR 2024poster

Prompt learning for vision-language models, e.g., CoOp, has shown great success in adapting CLIP to different downstream tasks, making it a promising solution for federated learning due to computational reasons. Existing prompt learning techniques replace hand-crafted text prompts with learned vecto…

Cited by 11SourcePDFScholar
2024

MODDP: A Multi-modal Open-domain Chinese Dataset for Dialogue Discourse Parsing

ACL 2024findings

Dialogue discourse parsing (DDP) aims to capture the relations between utterances in the dialogue. In everyday real-world scenarios, dialogues are typically multi-modal and cover open-domain topics. However, most existing widely used benchmark datasets for DDP contain only textual modality and are d…

2024

Stability and Generalization for Stochastic Recursive Momentum-based Algorithms for (Strongly-)Convex One to $K$-Level Stochastic Optimizations

ICML 2024poster

STOchastic Recursive Momentum (STORM)-based algorithms have been widely developed to solve one to $K$-level ($K \geq 3$) stochastic optimization problems. Specifically, they use estimators to mitigate the biased gradient issue and achieve near-optimal convergence results. However, there is relativel…

Cited by 0SourcePDFScholar
2023

Boosting Multi-modal Model Performance with Adaptive Gradient Modulation

ICCV 2023poster

While the field of multi-modal learning keeps growing fast, the deficiency of the standard joint training paradigm has become clear through recent studies. They attribute the sub-optimal performance of the jointly trained model to the modality competition phenomenon. Existing works attempt to improv…

Cited by 29PDFcodeScholar
2023

Cross-View Geo-Localization via Learning Disentangled Geometric Layout Correspondence

AAAI 2023technical

Cross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and different capturing time between two views. Despite these difficu…

2023

How To Prevent the Poor Performance Clients for Personalized Federated Learning?

CVPR 2023poster

Personalized federated learning (pFL) collaboratively trains personalized models, which provides a customized model solution for individual clients in the presence of heterogeneous distributed local data. Although many recent studies have applied various algorithms to enhance personalization in pFL,…

Cited by 19SourcePDFScholar
2023

Improving Adversarial Robustness with Self-Paced Hard-Class Pair Reweighting

AAAI 2023technical

Deep Neural Networks are vulnerable to adversarial attacks. Among many defense strategies, adversarial training with untargeted attacks is one of the most effective methods. Theoretically, adversarial perturbation in untargeted attacks can be added along arbitrary directions and the predicted labels…

2022

A Domain-Adapted Machine Learning Approach for Visual Evaluation and Interpretation of Robot-Assisted Surgery Skills

RA-L 2022

In this study, we present an intuitive machine learning-based approach to evaluate and interpret surgical skills level of a participant working with robotic platforms. The proposed method is domain-adapted, i.e., jointly utilizes an end-to-end learning approach for smoothness detection and domain kn

Cited by 13SourceScholar
2022

DDDM: A Brain-Inspired Framework for Robust Classification

IJCAI 2022poster

Despite their outstanding performance in a broad spectrum of real-world tasks, deep artificial neural networks are sensitive to input noises, particularly adversarial perturbations. On the contrary, human and animal brains are much less vulnerable. In contrast to the one-shot inference performed by…

2022

Generalized Federated Learning via Sharpness Aware Minimization

ICML 2022spotlight

Federated Learning (FL) is a promising framework for performing privacy-preserving, distributed learning with a set of clients. However, the data distribution among clients often exhibits non-IID, i.e., distribution shift, which makes efficient optimization difficult. To tackle this problem, many FL…

Cited by 182SourcePDFScholar
2022

Generating Diverse and Natural 3D Human Motions From Text

CVPR 2022poster

Automated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem wi…

Cited by 615PDFcodeScholar
2022

Object Wake-Up: 3D Object Rigging from a Single Image

ECCV 2022poster

"Given a single chair image, could we wake it up by reconstructing its 3D shape and skeleton, as well as animating its plausible articulations and motions, similar to that of human modeling? It is a new problem that not only goes beyond image-based object reconstruction but also involves articulated…

Cited by 7SourcePDFScholar
2022

SHAPE: An Unified Approach to Evaluate the Contribution and Cooperation of Individual Modalities

IJCAI 2022poster

As deep learning advances, there is an ever-growing demand for models capable of synthesizing information from multi-modal resources to address the complex tasks raised from real-life applications. Recently, many large multi-modal datasets have been collected, on which researchers actively explore d…

2021

Deep Neural Skill Assessment and Transfer: Application to Robotic Surgery Training

IROS 2021poster

Due to the high sensitivity and complexity of robotic surgery tasks, acquiring appropriate skill levels by trainee surgeons through an effective training process is very important and affects the patient’s safety and the quality of surgical outcomes. With the advanced deep learning technology and th…

Cited by 16SourceScholar
2021

Learning with Instance-Dependent Label Noise: A Sample Sieve Approach

ICLR 2021poster

Human-annotated labels are often prone to noise, and the presence of such noise will degrade the performance of the resulting deep neural network (DNN) models. Much of the literature (with several recent exceptions) of learning with noisy labels focuses on the case when the label noise is independen…

2015

Blind stain decomposition for histo-pathology images using circular nature of chroma components

ICASSP 2015accepted

In this paper, we present a novel approach to achieve blind stain decomposition in histo-pathology images. The method is based on stain color estimation, followed by stain absorbing vector generation and matrix computation. Unlike conventional approaches adopting linear processing algorithms to anal…

Cited by 0SourceScholar