← Search

Xin Zhou

70 accepted papers

2026

Benchmarking Real-Time Question Answering via Executable Code Workflows

IJCAI 2026

Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and therefore fail to capture the temporal dynamics of information and the continuously evolving nature of real-world knowledge.

Cited by 0Scholar
2026

Flying in Clutter on Monocular RGB by Learning in 3D Radiance Fields With Domain Adaptation

RA-L 2026

Modern autonomous navigation systems predominantly rely on lidar and depth cameras. However, a fundamental question remains: Can flying robots navigate in clutter using solely monocular RGB images? Given the prohibitive costs of real-world data collection, learning policies in simulation offers a pr

Cited by 4SourceScholar
2026

PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding

CVPR 2026

Scene-level point cloud understanding remains challenging due to diverse geometries, imbalanced category distributions, and highly varied spatial layouts. Existing methods improve object-level performance but rely on static network parameters during inference, limiting their adaptability to dynamic

Cited by 0SourcecodeScholar
2026

TransFR: Transferable Federated Recommendation with Adapter Tuning on Pre-trained Language Models

AAAI 2026technical

Federated recommendations (FRs), facilitating multiple local clients to collectively learn a global model without disclosing user private data, have emerged as a prevalent on-device service. In conventional FRs, a dominant paradigm is to utilize discrete identities to represent clients and items, wh

Cited by 0SourcePDFScholar
2026

USE: A Unified Model for Universal Sound Separation and Extraction

AAAI 2026technical

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require precisely specified clues to achieve optimal performance. This

Cited by 0SourcePDFScholar
2026

UniFuture: A 4D Driving World Model for Future Generation and Perception

ICRA 2026poster

We present UniFuture, a unified 4D Driving World Model designed to simulate the dynamic evolution of the 3D physical world. Unlike existing driving world models that focus solely on 2D pixel-level video generation (lacking geometry) or static perception (lacking temporal dynamics), our approach brid…

2026

VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric Videos

CVPR 2026

With the great advancement of video generation models, a growing number of content creators and researchers are leveraging these technologies to produce large volumes of human-centric videos for content creation and customized data generation for specific tasks. Although existing video generation mo

Cited by 0SourcecodeScholar
2026

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

CVPR 2026

Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA, a training-free identify-then-guide framework for improved numerical alignment. NUMINA identifies prompt-layout inconsi

Cited by 0SourcecodeScholar
2025

AuscMLLM: Bridging Classification and Reasoning in Heart Sound Analysis with a Multimodal Large Language Model

ICASSP 2025accepted

This study introduces a multimodal large language model capable of not only accomplishing various heart sound tasks but also providing reasoning, marking an advancement in the field of medical diagnostics. The model’s innovation stems from a collaboration with experts to collect a novel dataset desi…

Cited by 0SourceScholar
2025

Both Supply and Precision: Sample Debias and Ranking Consistency Joint Learning for Large Scale Pre-Ranking System

AAAI 2025technical

Cascade ranking architecture, composed of matching, pre-ranking, ranking and re-ranking stages, is usually adopted to balance the efficiency and effectiveness in real-world recommendation system (RS). As the middle stage of RS, pre-ranking aims to quickly filter out the low-quality items selected at…

Cited by 0SourcePDFScholar
2025

ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge

EMNLP 2025

We introduce ESGenius , a comprehensive benchmark for evaluating and enhancing the proficiency of Large Language Models (LLMs) in Environmental, Social, and Governance (ESG) and sustainability-focused question answering. ESGenius comprises two key components: (i) ESGenius-QA , a collection of 1,136

2025

EffiQA: Efficient Question-Answering with Strategic Multi-Model Collaboration on Knowledge Graphs

COLING 2025main

While large language models (LLMs) have shown remarkable capabilities in natural language processing, they struggle with complex, multi-step reasoning tasks involving knowledge graphs (KGs). Existing approaches that integrate LLMs and KGs either underutilize the reasoning abilities of LLMs or suffer…

Cited by 4SourcePDFScholar
2025

EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates

ICASSP 2025accepted

Active Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Add…

Cited by 0SourceScholar
2025

FACT: Fast and Active Coordinate Initialization for Vision-Based Drone Swarms

RA-L 2025

Coordinate initialization is the first step in accomplishing collaborative tasks within robot swarms, determining the quality of tasks. However, fast and robust coordinate initialization in vision-based drone swarms remains elusive. To this end, our letter proposes a complete system for initial rela

Cited by 1SourcecodeScholar
2025

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

ICCV 2025poster

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which involves interpreting and reasoning about the driving environment. In this paper, we…

2025

KINND: A Keyframe Insertion Framework via Neural Network Decision-Making for VSLAM

RA-L 2025

Keyframe insertion is critical for the performance and robustness of SLAM systems. However, traditional heuristic-based methods often lead to suboptimal keyframe selection, compromising the accuracy of localization and mapping. To address this, we propose KINND, a lightweight neural network-based fr

Cited by 4SourceScholar
2025

More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models

NeurIPS 2025poster

Generative depth estimation methods leverage the rich visual priors stored in pretrained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, parameter updates during training lead to catastrophic degradation in the image generation capability of the pretrained mo…

Cited by 0SourceScholar
2025

S2R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

ACL 2025long

Recent studies have demonstrated the effectiveness of LLM test-time scaling. However, existing approaches to incentivize LLMs’ deep thinking abilities generally require large-scale data or significant training efforts. Meanwhile, it remains unclear how to improve the thinking abilities of less power…

2025

UniTac-NV: A Unified Tactile Representation For Non-Vision-Based Tactile Sensors *

IROS 2025

Generalizable algorithms for tactile sensing remain underexplored, primarily due to the diversity of sensor modalities. Recently, many methods for cross-sensor transfer between optical (vision-based) tactile sensors have been investigated, yet little work focus on non-optical tactile sensors. To add

Cited by 0SourceScholar
2025

Variable-Friction In-Hand Manipulation for Arbitrary Objects via Diffusion-Based Imitation Learning

ICRA 2025

Dexterous in-hand manipulation (IHM) for arbitrary objects is challenging due to the rich and subtle contact process. Variable-friction manipulation is an alternative approach to dexterity, previously demonstrating robust and versatile 2D IHM capabilities with only two single-joint fingers. However,

Cited by 2SourceScholar
2025

kNN-CL: Enhancing Continual Learning with Nearest Neighbor Retrieval

ICASSP 2025accepted

Continual learning aims to learn new tasks sequentially without forgetting previously acquired knowledge. However, catastrophic forgetting remains a significant challenge. In this paper, we introduce kNN-CL, a simple yet effective approach that harnesses k-nearest neighbors (kNN) to mitigate forgett…

Cited by 0SourceScholar
2025

sEMG-Based Joint Angle Estimation via Hierarchical Spiking Attentional Feature Decomposition Network

RA-L 2025

Surface electromyography (sEMG) has demonstrated significant potential in simultaneous and proportional control (SPC). However, existing algorithms for predicting joint angles based on sEMG often suffer from high inference costs or are limited to specific subjects rather than multi-subject scenarios

Cited by 2SourcecodeScholar
2024

A Unified Framework for 3D Scene Understanding

NeurIPS 2024poster

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are typically tailored to a specific task, limiting their underst…

2024

Better Zero-Shot Reasoning with Role-Play Prompting

NAACL 2024long

Modern large language models (LLMs) exhibit a remarkable capacity for role-playing, enabling them to embody not only human characters but also non-human entities. This versatility allows them to simulate complex human-like interactions and behaviors within various contexts, as well as to emulate spe…

2024

Dual-View Whitening on Pre-trained Text Embeddings for Sequential Recommendation

AAAI 2024technical

Recent advances in sequential recommendation models have demonstrated the efficacy of integrating pre-trained text embeddings with item ID embeddings to achieve superior performance. However, our study takes a unique perspective by exclusively focusing on the untapped potential of text embeddings, o…

Cited by 8SourcePDFScholar
2024

Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud Analysis

CVPR 2024poster

Point cloud analysis has achieved outstanding performance by transferring point cloud pre-trained models. However existing methods for model adaptation usually update all model parameters i.e. full fine-tuning paradigm which is inefficient as it relies on high computational costs (e.g. training GPU…

2024

HSDreport: Heart Sound Diagnosis with Echocardiography Reports

EMNLP 2024finding

Heart sound auscultation holds significant importance in the diagnosis of congenital heart disease. However, existing methods for Heart Sound Diagnosis (HSD) tasks are predominantly limited to a few fixed categories, framing the HSD task as a rigid classification problem that does not fully align wi…

Cited by 0SourcePDFScholar
2024

LongHeads: Multi-Head Attention is Secretly a Long Context Processor

EMNLP 2024finding

Large language models (LLMs) have achieved impressive performance in numerous domains but often struggle to process lengthy inputs effectively and efficiently due to limited length generalization and attention’s quadratic computational demands. Many sought to mitigate this by restricting the attenti…

2024

Making Harmful Behaviors Unlearnable for Large Language Models

ACL 2024findings

Large language models (LLMs) have shown great potential to empower various domains and are often customized by fine-tuning for the requirements of different applications. However, the powerful learning ability of LLMs not only enables them to learn new tasks but also makes them vulnerable to learnin…

2024

PointMamba: A Simple State Space Model for Point Cloud Analysis

NeurIPS 2024poster

Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity method with global modeling appealing. In this paper, we pr…

2024

Preserving Relative Localization of FoV-Limited Drone Swarm via Active Mutual Observation

IROS 2024poster

Relative state estimation is crucial for vision-based swarms to estimate and compensate for the unavoidable drift of visual odometry. For autonomous drones equipped with the most compact sensor setting — a stereo camera that provides a limited field of view (FoV), the demand for mutual observation f…

Cited by 1SourcecodeScholar
2024

Rewriting the Code: A Simple Method for Large Language Model Augmented Code Search

ACL 2024long

In code search, the Generation-Augmented Retrieval (GAR) framework, which generates exemplar code snippets to augment queries, has emerged as a promising strategy to address the principal challenge of modality misalignment between code snippets and natural language queries, particularly with the dem…

2024

Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs

EMNLP 2024main

Retrieval-Augmented Generation (RAG) significantly improved the ability of Large Language Models (LLMs) to solve knowledge-intensive tasks. While existing research seeks to enhance RAG performance by retrieving higher-quality documents or designing RAG-specific LLMs, the internal mechanisms within L…

2023

Coarse-to-fine Few-shot Learning for Named Entity Recognition

ACL 2023findings

Recently, Few-shot Named Entity Recognition has received wide attention with the growing need for NER models to learn new classes with minimized annotation costs. However, one common yet understudied situation is to transfer a model trained with coarse-grained classes to recognize fine-grained class…

2023

Continuous Estimation of Lower Limb Joint Angles From Multi-Stream Signals Based on Knowledge Tracing

RA-L 2023

Multi-stream signals are increasingly being used in robot-assisted rehabilitation training, where the timely and accurate prediction of a patient's motor intentions is frequently required to provide simultaneous and proportional control strategies. However, existing methods for motion intent predict

Cited by 21SourceScholar
2023

Diffusion Motion: Generate Text-Guided 3D Human Motion by Diffusion Model

ICASSP 2023accepted

We propose a simple and novel method for generating 3D human motion from complex natural language sentences, which describe different velocity, direction and composition of all kinds of actions. Different from existing methods that use classical generative architecture, we apply the Denoising Diffus…

Cited by 0SourceScholar
2023

Do We Really Need Complicated Model Architectures For Temporal Networks?

ICLR 2023top-5%

Recurrent neural network (RNN) and self-attention mechanism (SAM) are the de facto methods to extract spatial-temporal information for temporal graph learning. Interestingly, we found that although both RNN and SAM could lead to a good performance, in practice neither of them is always necessary. In…

Cited by 156SourcePDFScholar
2023

InstaGrasp: An Entirely 3D Printed Adaptive Gripper with TPU Soft Elements and Minimal Assembly Time

IROS 2023poster

Fabricating existing and popular open-source adaptive robotic grippers commonly involves using multiple professional machines, purchasing a wide range of parts, and tedious, time-consuming assembly processes. This poses a significant barrier to entry for some robotics researchers and drives others t…

Cited by 4SourceScholar
2023

Learning “O” Helps for Learning More: Handling the Unlabeled Entity Problem for Class-incremental NER

ACL 2023long

As the categories of named entities rapidly increase, the deployed NER models are required to keep updating toward recognizing more entity types, creating a demand for class-incremental learning for NER. Considering the privacy concerns and storage constraints, the standard paradigm for class-increm…

2023

PlanarTrack: A Large-scale Challenging Benchmark for Planar Object Tracking

ICCV 2023poster

Planar object tracking is a critical computer vision problem and has drawn increasing interest owing to its key roles in robotics, augmented reality, etc. Despite rapid progress, its further development, especially in the deep learning era, is largely hindered due to the lack of large-scale challeng…

Cited by 5PDFScholar
2023

Raising The Limit of Image Rescaling Using Auxiliary Encoding

ICASSP 2023accepted

Normalizing flow models using invertible neural networks (INN) have been widely investigated for successful generative image super-resolution (SR) by learning the transformation between the normal distribution of latent variable z and the conditional distribution of high-resolution (HR) images gave…

Cited by 0SourceScholar
2023

Tactile Identification of Object Shapes via In-Hand Manipulation with A Minimalistic Barometric Tactile Sensor Array

ICRA 2023poster

With the goal of providing an alternative to optical and other tactile sensors, we set out to stress test the object shape identification capabilities of barometric tactile arrays in robotic manipulation tasks. These sensors are superior to optical devices in terms of form factor, ease of fabricatio…

Cited by 3SourceScholar
2023

TextMixer: Mixing Multiple Inputs for Privacy-Preserving Inference

EMNLP 2023long findings

Pre-trained language models (PLMs) are often deployed as cloud services, enabling users to upload textual data and perform inference remotely. However, users' personal text often contains sensitive information, and sharing such data directly with the service providers can lead to serious privacy l…

Cited by 0SourceScholar
2023

TextObfuscator: Making Pre-trained Language Model a Privacy Protector via Obfuscating Word Representations

ACL 2023findings

In real-world applications, pre-trained language models are typically deployed on the cloud, allowing clients to upload data and perform compute-intensive inference remotely. To avoid sharing sensitive data directly with service providers, clients can upload numerical representations rather than pla…

2023

Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty View

EMNLP 2023long main

Recently, many studies have illustrated the robustness problem of Named Entity Recognition (NER) systems: the NER models often rely on superficial entity patterns for predictions, without considering evidence from the context. Consequently, even state-of-the-art NER models generalize poorly to out-o…

Cited by 0SourceScholar
2022

A Novel Method for Detecting Misclassifications of the Locomotion Mode in Lower-Limb Exoskeleton Robot Control

RA-L 2022

Lower-limb exoskeleton robots can support hemiplegic patients’ affected limbs and assist in their rehabilitation. In order to set effective control strategies, it is necessary to obtain the user’s motion intention accurately and timeously. These requirements pose many challenges. The surface electro

Cited by 23SourceScholar
2022

ASM-Loc: Action-Aware Segment Modeling for Weakly-Supervised Temporal Action Localization

CVPR 2022poster

Weakly-supervised temporal action localization aims to recognize and localize action segments in untrimmed videos given only video-level action labels for training. Without the boundary information of action segments, existing methods mostly rely on multiple instance learning (MIL), where the predic…

Cited by 120PDFcodeScholar
2022

E-TRoll: Tactile Sensing and Classification via A Simple Robotic Gripper for Extended Rolling Manipulations

IROS 2022poster

Robotic tactile sensing provides a method of recognizing objects and their properties where vision fails. Prior work on tactile perception in robotic manipulation has frequently focused on exploratory procedures (EPs). However, the also-human-inspired technique of in-hand-manipulation can glean rich…

Cited by 8SourceScholar
2022

LFKQG: A Controlled Generation Framework with Local Fine-tuning for Question Generation over Knowledge Bases

COLING 2022main

Question generation over knowledge bases (KBQG) aims at generating natural questions about a subgraph, which can be answered by a given answer entity. Existing KBQG models still face two main challenges: (1) Most models often focus on the most relevant part of the answer entity, while neglecting the…

Cited by 7SourcePDFScholar
2022

Making Parameter-efficient Tuning More Efficient: A Unified Framework for Classification Tasks

COLING 2022main

Large pre-trained language models (PLMs) have demonstrated superior performance in industrial applications. Recent studies have explored parameter-efficient PLM tuning, which only updates a small amount of task-specific parameters while achieving both high efficiency and comparable performance again…

2022

ProofInfer: Generating Proof via Iterative Hierarchical Inference

EMNLP 2022main

Proof generation focuses on deductive reasoning: given a hypothesis and a set of theories, including some supporting facts and logical rules expressed in natural language, the model generates a proof tree indicating how to deduce the hypothesis from given theories.Current models with state-of-the-ar…

2022

Searching for Optimal Subword Tokenization in Cross-domain NER

IJCAI 2022poster

Input distribution shift is one of the vital problems in unsupervised domain adaptation (UDA). The most popular UDA approaches focus on domain-invariant representation learning, trying to align the features from different domains into a similar feature distribution. However, these approaches ignore…

2022

Template-free Prompt Tuning for Few-shot NER

NAACL 2022long

Prompt-based methods have been successfully applied in sentence-level few-shot learning tasks, mostly owing to the sophisticated design of templates and label words. However, when applied to token-level labeling tasks such as NER, it would be time-consuming to enumerate the template queries over all…

2022

TextFusion: Privacy-Preserving Pre-trained Model Inference via Token Fusion

EMNLP 2022main

Recently, more and more pre-trained language models are released as a cloud service. It allows users who lack computing resources to perform inference with a powerful model by uploading data to the cloud. The plain text may contain private information, as the result, users prefer to do partial compu…

2021

Dual-Objective Collision-Free Path Optimization of Arc Welding Robot

RA-L 2021

In order to solve the dual-objective path optimization problem of arc welding robot, a complete and novel methodology is proposed, including modeling, collision-free path search and global path optimization. First of all, the grid method is utilized to accurately model the welding workpieces and sur

Cited by 19SourceScholar
2021

EGO-Planner: An ESDF-Free Gradient-Based Local Planner for Quadrotors

RA-L 2021

Gradient-based planners are widely used for quadrotor local planning, in which a Euclidean Signed Distance Field (ESDF) is crucial for evaluating gradient magnitude and direction. Nevertheless, computing such a field has much redundancy since the trajectory optimization procedure only covers a very

Cited by 455SourcecodeScholar
2021

EGO-Swarm: A Fully Autonomous and Decentralized Quadrotor Swarm System in Cluttered Environments

ICRA 2021poster

This paper presents a decentralized and asynchronous systematic solution for multi-robot autonomous navigation in unknown obstacle-rich scenes using merely onboard resources. The planning system is formulated under gradient-based local planning framework, where collision avoidance is achieved by for…

Cited by 197SourcecodeScholar
2021

No Need for Interactions: Robust Model-Based Imitation Learning using Neural ODE

ICRA 2021poster

Interactions with either environments or expert policies during training are needed for most of the current imitation learning (IL) algorithms. For IL problems with no interactions, a typical approach is Behavior Cloning (BC). However, BC-like methods tend to be affected by distribution shift. To mi…

Cited by 9SourcecodeScholar
2021

TGK-Planner: An Efficient Topology Guided Kinodynamic Planner for Autonomous Quadrotors

RA-L 2021

In this letter, we propose a lightweight yet effective Topology Guided Kinodynamic planner (TGK-Planner) for quadrotor aggressive flights with limited onboard computing resources. The proposed system follows the traditional hierarchical planning workflow, with novel designs to improve the robustness

Cited by 38SourcecodeScholar
2020

Alternating Minimization Based Trajectory Generation for Quadrotor Aggressive Flight

RA-L 2020

With much research has been conducted into trajectory planning for quadrotors, planning with spatial and temporal optimal trajectories in real-time is still challenging. In this letter, we propose a framework for large-scale waypoint-based polynomial trajectory generation, with highlights on its sup

Cited by 38SourcecodeScholar
2018

Prediction of Satisfied User Ratio for Compressed Video

ICASSP 2018accepted

A large-scale video quality dataset called the VideoSet has been constructed recently to measure human subjective experience of H.264 coded video in terms of the just-noticeable-difference (JND). It measures the first three JND points of 5-second video of resolution 1080p, 720p, 540p and 360p. Based…

Cited by 0SourceScholar