← Search

Xin Xu

80 accepted papers

2026

A Hierarchical Vision-Language and Reinforcement Learning Framework for Robotic Task and Motion Planning in Collaborative Manipulation

RA-L 2026

Vision-language-action models (VLAs) use an end-to-end learning architecture, which can realize the integration of visual perception, semantic understanding and motion control. However, when tackling with the dynamic or long-horizon tasks, VLAs have poor robustness and real-time adjustment ability a

Cited by 2SourceScholar
2026

BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses

ICLR 2026poster

Existing studies on bias mitigation methods for large language models (LLMs) use diverse baselines and metrics to evaluate debiasing performance, leading to inconsistent comparisons among them. Moreover, their evaluations are mostly based on the comparison between LLMs' probabilities of biased and u…

Cited by 0SourcecodeScholar
2026

Domain-Aware Suppression and Aggregation for Federated DG ReID

AAAI 2026technical

Federated domain generalization in person re-identification (FedDG-ReID) aims to learn a privacy-preserving server model from decentralized client source domains that generalizes to unseen domains. Existing approaches enhance the generalizability of the server model by increasing the diversity of c

Cited by 0SourcePDFScholar
2026

Efficient Hierarchical Reinforcement Learning with Dynamic Kolmogorov–Arnold Network for Long-Horizon Robotic Manipulation

ICRA 2026poster

Long-horizon robotic manipulation remains a critical challenge in robotics. Hierarchical reinforcement learning offers a promising solution, but often suffers from an imbalance dilemma: simplifying skill learning increases the complexity of planning, thereby expanding the solution space and computat…

Cited by 0Scholar
2026

FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts

CVPR 2026

Federated Domain Generalization for Person Re-Identification (FedDG-ReID) aims to learn domain-invariant representations from decentralized data. Although Vision Transformers (ViTs) are widely adopted, their global attention often fails to distinguish pedestrians from high similarity backgrounds or

Cited by 0SourcecodeScholar
2026

ICLR: Inter-Chrominance and Luminance Interaction for Natural Color Restoration in Low-Light Image Enhancement

AAAI 2026technical

Low-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling of chrominance and luminance. However, for the interaction of

Cited by 0SourcePDFScholar
2026

MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution

AAAI 2026technical

Chinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has adva

Cited by 0SourcePDFScholar
2026

MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task

ICLR 2026poster

Mathematical reasoning represents a critical frontier in advancing large language models (LLMs). While step-by-step approaches have emerged as the dominant paradigm for mathematical problem-solving in LLMs, the quality of reasoning steps in training data fundamentally constrains model performance. R…

Cited by 0SourceScholar
2026

Neuro-evolutionary Continual Reinforcement Learning

ICML 2026spotlight

Deploying robots in open‑ended real‑world environments demands continual learning capabilities to adapt to an ever-expanding range of tasks. This requires retaining previously acquired skills without forgetting while effectively leveraging prior knowledge to learn new ones. Inspired by neuroscience,…

Cited by 0SourceScholar
2026

On Predictability of Reinforcement Learning Dynamics for Large Language Models

ICLR 2026poster

Recent advances in reasoning capabilities of large language models (LLMs) are largely driven by reinforcement learning (RL), yet the underlying parameter dynamics during RL training remain poorly understood. This work identifies two fundamental properties of RL-induced parameter updates in LLMs: (1)…

Cited by 0SourcecodeScholar
2026

PRISM: Sequence Modeling as Parallel Residual Iteration

ICML 2026poster

Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models. Existing efficient architectures are theoretically bounded by shallow, single-step linear updates, while powerful iterative methods like Test-Time Training …

Cited by 0SourceScholar
2026

Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward

ICLR 2026poster

Enhancing the multimodal reasoning capabilities of Multimodal Large Language Models (MLLMs) is a challenging task that has attracted increasing attention in the community. Recently, several studies have applied Reinforcement Learning with Verifiable Rewards (RLVR) to the multimodal domain in order t…

Cited by 0SourcecodeScholar
2026

RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection

CVPR 2026

Underwater images commonly suffer from foreground-background ambiguity, loss of structural details, and severely reduced contrast, which collectively make underwater object detection (UOD) an inherently challenging task. To handle this issue, we present a residual-guided hierarchical calibration net

Cited by 0SourcecodeScholar
2026

Receding Horizon Reinforcement Learning with Autoregressive Model for Motion Control of Autonomous Vehicles

ICRA 2026poster

This paper presents a model-based reinforcement learning (MBRL) approach with a receding horizon mechanism to optimize the lateral trajectory-tracking performance of autonomous vehicles (AVs). Accurate modeling of complex vehicle dynamics and adaptation to dynamic environments with limited data pose…

Cited by 0Scholar
2026

S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary Segmentation

CVPR 2026

Open-vocabulary semantic segmentation extends pixel-level recognition to arbitrary text-described categories. Despite strong global semantic understanding, vision-language models such as CLIP exhibit limited spatial precision and semantic ambiguity across large vocabularies, constraining their effec

Cited by 0SourceScholar
2026

Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners

ICLR 2026poster

Reinforcement Learning with Verifiable Reward (RLVR) effectively solves complex tasks but demands extremely long context lengths during training, leading to substantial computational costs. While multi-stage training can partially mitigate this, starting with overly short contexts often causes irrev…

Cited by 0SourcecodeScholar
2026

USER: A Unified and Extensible System for Online Real-World Policy Learning in Embodied AI

RSS 2026poster

Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitrarily accelerated, cheaply reset, or massively replicated, which makes scalable data collection, heterogeneous deployment, a…

Cited by 0SourceScholar
2026

VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection

CVPR 2026

Closed-set object detection in remote sensing imagery has made significant progress, but achieving high detection accuracy remains challenging. Vision-Language Models (VLMs), which possess rich prior knowledge, offer a promising solution to this challenge. However, most existing VLMs are designed fo

Cited by 0SourcecodeScholar
2026

VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models

ICLR 2026poster

Large reasoning models such as OpenAI o1 and DeepSeek-R1 have demonstrated remarkable performance in complex reasoning tasks. A critical component of their training is the incorporation of reference-based reward systems within reinforcement learning (RL), where model outputs are evaluated against gr…

Cited by 0SourcecodeScholar
2026

WENETSPEECH-CHUAN: A LARGE-SCALE SICHUANESE CORPUS WITH RICH ANNOTATION FOR DIALECTAL SPEECH PROCESSING

ICASSP 2026poster

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constru…

Cited by 0SourcePDFScholar
2026

WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional Annotation

AAAI 2026technical

The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximate

Cited by 0SourcePDFScholar
2026

When Birds Meet Fish: Vision-Force Fusion for Autonomous Underwater Docking in Cross-Domain Avian-Aquatic Collaboration

ICRA 2026poster

Unmanned aerial–aquatic vehicles (UAAVs) provide cross-domain adaptability and broad visions, while autonomous underwater vehicles (AUVs) support long-duration operations. This work integrates the two by developing a rapid underwater docking and releasing system. An autonomous clamping mechanism is …

Cited by 0Scholar
2025

Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception

ICLR 2025poster

With success in image generation, generative diffusion models are increasingly adopted for discriminative scenarios because generating pixels is a unified and natural perception interface. Although directly re-purposing their generative denoising process has established promising progress in special…

2025

CKnowEdit: A New Chinese Knowledge Editing Dataset for Linguistics, Facts, and Logic Error Correction in LLMs

ACL 2025long

Chinese, as a linguistic system rich in depth and complexity, is characterized by distinctive elements such as ancient poetry, proverbs, idioms, and other cultural constructs. However, current Large Language Models (LLMs) face limitations in these specialized domains, highlighting the need for the d…

2025

Can LLMs Solve Longer Math Word Problems Better?

ICLR 2025poster

Math Word Problems (MWPs) play a vital role in assessing the capabilities of Large Language Models (LLMs), yet current research primarily focuses on questions with concise contexts. The impact of longer contexts on mathematical reasoning remains under-explored. This study pioneers the investigation…

2025

Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning

IROS 2025

In offline reinforcement learning, value overestimation caused by out-of-distribution (OOD) actions significantly limits policy performance. Recently, diffusion models have been leveraged for their strong distribution-matching capabilities, enforcing conservatism through behavior policy constraints.

Cited by 0SourceScholar
2025

Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration

ICASSP 2025accepted

Video-based person re-identification (ReID) has become increasingly important due to its applications in video surveillance applications. By employing events in video-based person ReID, more motion information can be provided between continuous frames to improve recognition accuracy. Previous approa…

Cited by 5SourceScholar
2025

Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection

ICML 2025oral

One-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent.…

2025

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

NeurIPS 2025poster

Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers,…

Cited by 0SourcecodeScholar
2025

Learning Predictive Control with Online Modeling for Agile Maneuvering of Autonomous Vehicles

IROS 2025

The agile maneuvering control of autonomous vehicles (AVs) requires the tracking of reference trajectories characterized by high acceleration, sharp curvature, considerable disturbances, and significant time-varying, all while ensuring stability and accuracy. The inherent uncertainty and time-varyin

Cited by 0SourceScholar
2025

MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation

ICASSP 2025accepted

The technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primarily due to the difficulty of balancing performance and cost with existing objective and subjective evaluation methods. In…

Cited by 0SourceScholar
2025

Offline Reinforcement Learning with Koopman Operators for Control of Soft Robots

IROS 2025

Soft robots are promising to offer flexibility in environmental interaction tasks through compliant deformations. However, the infinite degrees of freedom and high nonlinearity of dynamics pose significant challenges in dynamic modeling and control in soft robots. While online reinforcement learning

Cited by 0SourceScholar
2025

S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical Reasoners

AAAI 2025technical

Self-correction is a novel method that can stimulate the potential reasoning abilities of large language models (LLMs). It involves detecting and correcting errors during the inference process when LLMs solve reasoning problems. However, recent works do not regard self-correction as a spontaneous an…

Cited by 7SourcePDFScholar
2025

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification

ACL 2025long

Chain-of-Thought (CoT) prompting has become the de facto method to elicit reasoning capabilities from large language models (LLMs). However, to mitigate hallucinations in CoT that are notoriously difficult to detect, current methods such as process reward models (PRMs) or self-consistency operate as…

2025

Skill Expansion and Composition in Parameter Space

ICLR 2025poster

Humans excel at reusing prior knowledge to address new challenges and developing skills while solving problems. This paradigm becomes increasingly popular in the development of autonomous agents, as it develops systems that can self-evolve in response to new challenges like human beings. However, pr…

2025

The Parables of the Mustard Seed and the Yeast: Extremely Low-Budget, High-Performance Nighttime Semantic Segmentation

AAAI 2025technical

Nighttime Semantic Segmentation (NSS) is essential to many cutting-edge vision applications. However, existing technologies overly rely on massive labeled data, whose annotation is time-consuming and laborious. In this paper, we pioneer a new task focusing on exploring the potential of training stra…

Cited by 0SourcePDFScholar
2025

TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-Identification

AAAI 2025technical

Unsupervised visible-infrared person re-identification (US-VI-ReID) seeks to match infrared and visible images of the same individual without the use of annotations. Current methods typically derive cross-modal correspondences through a single global feature matching process for generating pseudo la…

2025

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

ICLR 2025poster

Large Language Models (LLMs) have made significant strides in mathematical reasoning, underscoring the need for a comprehensive and fair evaluation of their capabilities. However, existing benchmarks often fall short, either lacking extensive coverage of undergraduate-level mathematical problems or…

Cited by 1SourcePDFScholar
2025

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

ICML 2025poster

Large language models (LLMs) have demonstrated remarkable capabilities in solving complex reasoning tasks, particularly in mathematics. However, the domain of physics reasoning presents unique challenges that have received significantly less attention. Existing benchmarks often fall short in evaluat…

2025

Versatile Distributed Maneuvering With Generalized Formations Using Guiding Vector Fields

ICRA 2025

This paper presents a unified approach to realize versatile distributed maneuvering with generalized formations. Specifically, we decompose the robots' maneuvers into two independent components, i.e., interception and enclosing, which are parameterized by two independent virtual coordinates. Treatin

Cited by 0SourceScholar
2025

WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning

EMNLP 2025

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in the multimodal symbolic music domain remain largely unexplored.We introduce WildScore, the first in-the-wild multimodal sy

2024

Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning

ICML 2024poster

In offline reinforcement learning, the challenge of out-of-distribution (OOD) is pronounced. To address this, existing methods often constrain the learned policy through policy regularization. However, these methods often suffer from the issue of unnecessary conservativeness, hampering policy improv…

2024

Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View

ACL 2024long

As Natural Language Processing (NLP) systems are increasingly employed in intricate social environments, a pressing query emerges: *Can these NLP systems mirror human-esque collaborative intelligence, in a multi-agent society consisting of multiple large language models (LLMs)?* This paper probes th…

2024

FMRNet: Image Deraining via Frequency Mutual Revision

AAAI 2024technical

The wavelet transform has emerged as a powerful tool in deciphering structural information within images. And now, the latest research suggests that combining the prowess of wavelet transform with neural networks can lead to unparalleled image deraining results. By harnessing the strengths of both t…

2024

InfPose: Real-Time Infrared Multi-Human Pose Estimation for Edge Devices Based on Encoder-Decoder CNN Architecture

RA-L 2024

Despite its remarkable performance, RGB-based Multi-human Pose Estimation (MPE) technology has many practical limitations, such as nighttime and smoggy environments. Infrared imaging is a valid substitution in these scenarios but needs an efficient and fast method for MPE. This letter aims to design

Cited by 6SourceScholar
2024

Learning-Based Near-Optimal Motion Planning for Intelligent Vehicles With Uncertain Dynamics

RA-L 2024

Motion planning has been an important research topic in achieving safe and flexible maneuvers for intelligent vehicles. However, it remains challenging to realize efficient and optimal planning in the presence of uncertain model dynamics. In this paper, a sparse kernel-based reinforcement learning (

Cited by 6SourceScholar
2024

M3-GMN: A Multi-environment, Multi-LiDAR, Multi-task dataset for Grid Map based Navigation

IROS 2024

In this paper, we propose a multi-environment, multi-LiDAR, multi-task dataset to promote the grid map-based navigation capability for autonomous vehicles. The dataset comprises structured and unstructured environmental data captured by different types of LiDAR and contains various challenging scena

Cited by 2SourcecodeScholar
2024

Mutuality Attribute Makes Better Video Anomaly Detection

ICASSP 2024accepted

Video anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep featur…

Cited by 0SourceScholar
2024

RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization

NeurIPS 2024poster

The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded datasets. However, the acoustic mismatch between simulated and re…

2024

TaskLAMA: Probing the Complex Task Understanding of Language Models

AAAI 2024technical

Structured Complex Task Decomposition (SCTD) is the problem of breaking down a complex real-world task (such as planning a wedding) into a directed acyclic graph over individual steps that contribute to achieving the task, with edges specifying temporal dependencies between steps. SCTD is an i…

Cited by 8SourcePDFScholar
2023

BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information

NeurIPS 2023poster

Automated reasoning with unstructured natural text is a key requirement for many potential applications of NLP and for developing robust AI systems. Recently, Language Models (LMs) have demonstrated complex reasoning capacities even without any finetuning. However, existing evaluation for automated…

Cited by 39SourcePDFScholar
2023

CSGP: Closed-Loop Safe Grasp Planning via Attention-Based Deep Reinforcement Learning From Demonstrations

RA-L 2023

Grasping is at the core of many robotic manipulation tasks. Despite the recent progress, closed-loop grasp planning in stacked scenes is still unsatisfactory, in terms of efficiency, stability, and most importantly, safety. In this letter, we present CSGP, a closed-loop safe grasp planning approach

Cited by 8SourceScholar
2023

DDK: A Deep Koopman Approach for Longitudinal and Lateral Control of Autonomous Ground Vehicles

ICRA 2023poster

Autonomous driving has attracted lots of attention in recent years. For some tasks, e.g., trajectory prediction, motion planning, and trajectory tracking, an accurate vehicle model can reduce the difficulty of these tasks and improve task completion performance. Prior works focused on parameter esti…

Cited by 8SourceScholar
2023

From Generation to Suppression: Towards Effective Irregular Glow Removal for Nighttime Visibility Enhancement

IJCAI 2023poster

Most existing Low-Light Image Enhancement (LLIE) methods are primarily designed to improve brightness in dark regions, which suffer from severe degradation in nighttime images. However, these methods have limited exploration in another major visibility damage, the glow effects in real night scenes.…

Cited by 5SourcePDFScholar
2023

LAMBADA: Backward Chaining for Automated Reasoning in Natural Language

ACL 2023long

Remarkable progress has been made on automated reasoning with natural text, by using Large Language Models (LLMs) and methods such as Chain-of-Thought prompting and Selection-Inference. These techniques search for proofs in the forward direction from axioms to the conclusion, which suffers from a co…

Cited by 86SourcePDFScholar
2023

Neural Koopman Pooling: Control-Inspired Temporal Dynamics Encoding for Skeleton-Based Action Recognition

CVPR 2023poster

Skeleton-based human action recognition is becoming increasingly important in a variety of fields. Most existing works train a CNN or GCN based backbone to extract spatial-temporal features, and use temporal average/max pooling to aggregate the information. However, these pooling methods fail to cap…

2023

RSFNet: A White-Box Image Retouching Approach using Region-Specific Color Filters

ICCV 2023poster

Retouching images is an essential aspect of enhancing the visual appeal of photos. Although users often share common aesthetic preferences, their retouching methods may vary based on their individual preferences. Therefore, there is a need for white-box approaches that produce satisfying results and…

Cited by 14PDFcodeScholar
2023

Schema-adaptable Knowledge Graph Construction

EMNLP 2023long findings

Conventional Knowledge Graph Construction (KGC) approaches typically follow the static information extraction paradigm with a closed set of pre-defined schema. As a result, such approaches fall short when applied to dynamic scenarios or domains, whereas a new type of knowledge emerges. This necessit…

Cited by 0SourcecodeScholar
2023

StereoVAE: A lightweight stereo-matching system using embedded GPUs

ICRA 2023poster

We propose a lightweight system for stereo-matching using embedded graphic processing units (GPUs). The proposed system overcomes the trade-off between accuracy and processing speed in stereo matching, thus further improving the matching accuracy while ensuring real-time processing. The basic idea i…

Cited by 8SourceScholar
2022

Barrier Function-based Safe Reinforcement Learning for Formation Control of Mobile Robots

ICRA 2022poster

Distributed model predictive control (DMPC) concerns how to online control multiple robotic systems with constraints effectively. However, the nonlinearity, nonconvexity, and strong interconnections of dynamic system models and constraints can make the real-time and real-world DMPC implementations n…

Cited by 13SourceScholar
2022

HEA-D: A Hybrid Evolutionary Algorithm for Diversified Top-k Weight Clique Search Problem

IJCAI 2022poster

The diversified top-k weight clique (DTKWC) search problem is an important generalization of the diversified top-k clique (DTKC) search problem with extensive applications, which extends the DTKC search problem by taking into account the weight of vertices. In this paper, we formulate DTKWC search p…

2022

M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

Recent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologi…

Cited by 0SourceScholar
2022

Multiview Long-Short Spatial Contrastive Learning For 3D Medical Image Analysis

ICASSP 2022accepted

The success of supervised deep learning heavily depends on large labeled datasets whose construction is often challenging in medical image analysis. Contrastive learning, a variant of self-supervised learning, is a potential solution to alleviate the strong demand for data annotation. In this work,…

Cited by 0SourceScholar
2022

Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve Refinement

ICASSP 2022accepted

Deep learning networks with deeper layers become a trend for their good performance but lacks the potential for real-time mobile deployment. Another challenge for paired training networks is the limited generalization capacity caused by the sample bias. To overcome these two challenges, we propose a…

Cited by 0SourceScholar
2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2022

SymmetryGrasp: Symmetry-Aware Antipodal Grasp Detection From Single-View RGB-D Images

RA-L 2022

Symmetry is ubiquitous in everyday objects. Humans tend to grasp objects by recognizing the symmetric regions. In this letter, we investigate how symmetry could boost robotic grasp detection. To this end, we present a learning-based method for detecting grasp from single-view RGB-D images. The key i

Cited by 11SourceScholar
2022

Towards Realistic Low-resource Relation Extraction: A Benchmark with Empirical Baseline Study

EMNLP 2022finding

This paper presents an empirical study to build relation extraction systems in low-resource settings. Based upon recent pre-trained language models, we comprehensively investigate three schemes to evaluate the performance in low-resource settings: (i) different types of prompt-based methods with few…

2022

WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition

ICASSP 2022accepted

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of…

Cited by 0SourceScholar
2021

StablePose: Learning 6D Object Poses From Geometrically Stable Patches

CVPR 2021poster

We introduce the concept of geometric stability to the problem of 6D object pose estimation and propose to learn pose inference based on geometrically stable patches extracted from observed 3D point clouds. According to the theory of geometric stability analysis, a minimal set of three planar/cylind…

Cited by 46PDFScholar
2021

The Multi-Speaker Multi-Style Voice Cloning Challenge 2021

ICASSP 2021accepted

The Multi-speaker Multi-style Voice Cloning Challenge (M2VoC) aims to provide a common sizable dataset as well as a fair testbed for the benchmarking of the popular voice cloning task. Specifically, we formulate the challenge to adapt an average TTS model to the stylistic target voice with limited d…

Cited by 0SourceScholar
2021

Two-stream 2D/3D Residual Networks for Learning Robot Manipulations from Human Demonstration Videos

ICRA 2021poster

Learning manipulation skills from observing human demonstration videos is a promising aspect for intelligent robotic systems. Recent advances in video to command provide an end-to-end approach to translate a video into robot plans. However, the general video captioning methods focus more on the unde…

Cited by 9SourceScholar