← Search

Lei He

49 accepted papers

2026

Griffin: Aerial-Ground Cooperative Detection and Tracking Dataset and Benchmark

AAAI 2026technical

While cooperative perception can overcome the limitations of single-vehicle systems, the practical implementation of vehicle-to-vehicle and vehicle-to-infrastructure systems is often impeded by significant economic barriers. Aerial-ground cooperation (AGC), which pairs ground vehicles with drones, p

Cited by 0SourcePDFScholar
2026

Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference

AAAI 2026technical

We introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where exper

Cited by 0SourcePDFScholar
2026

SIGN: Safety-Aware Image-Goal Navigation for Autonomous Drones Via Reinforcement Learning

ICRA 2026poster

Image-goal navigation (ImageNav) tasks a robot with autonomously exploring an unknown environment and reaching a location that visually matches a given target image. While prior works primarily study ImageNav for ground robots, enabling this capability for autonomous drones is substantially more cha…

2026

SIGN: Safety-Aware Image-Goal Navigation for Autonomous Drones via Reinforcement Learning

RA-L 2026

Image-goal navigation (ImageNav) tasks a robot with autonomously exploring an unknown environment and reaching a location that visually matches a given target image. While prior works primarily study ImageNav for ground robots, enabling this capability for autonomous drones is substantially more cha

Cited by 1SourcecodeScholar
2026

Topological Active Inference for Task Disambiguation

ICML 2026poster

In open-ended domains, natural language instructions are often *underspecified*, mapping to multiple valid yet functionally distinct latent intents. While Large Language Models (LLMs) excel at generation, their ability to resolve such *task ambiguity* through interaction is currently hampered by *se…

Cited by 0SourceScholar
2026

TransforMARS: Fault-Tolerant Self-Reconfiguration for Arbitrary-Shaped Modular Aerial Robot Systems

ICRA 2026poster

Modular Aerial Robot Systems (MARS) consist of multiple drone modules that are physically bound together to form a single structure for flight. Exploiting structural redundancy, MARS can be reconfigured into different formations to mitigate unit or rotor failures and maintain stable flight. Prior wo…

Cited by 0codeScholar
2025

Controllable Traffic Simulation through LLM-Guided Hierarchical Reasoning and Refinement

IROS 2025

Evaluating autonomous driving systems in complex and diverse traffic scenarios through controllable simulation is essential to ensure their safety and reliability. However, existing traffic simulation methods face challenges in their controllability. To address this, we propose a novel diffusion-bas

Cited by 1SourceScholar
2025

Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation

AAAI 2025technical

Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core…

2025

Hierarchical End-to-End Autonomous Driving: Integrating BEV Perception with Deep Reinforcement Learning

ICRA 2025

End-to-end autonomous driving offers a stream-lined alternative to the traditional modular pipeline, integrating perception, prediction, and planning within a single framework. While Deep Reinforcement Learning (DRL) has recently gained traction in this domain, existing approaches often overlook the

Cited by 8SourceScholar
2025

PodAgent: A Comprehensive Framework for Podcast Generation

ACL 2025finding

Existing automatic audio generation methods struggle to generate podcast-like audio programs effectively. The key challenges lie in in-depth content generation, appropriate and expressive voice production. This paper proposed PodAgent, a comprehensive framework for creating audio programs. PodAgent…

2025

USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature Decorrelation

AAAI 2025technical

Contrastive learning has achieved great success in skeleton-based representation learning recently. However, the prevailing methods are predominantly negative-based, necessitating additional momentum encoder and memory bank to get negative samples, which increases the difficulty of model training. F…

2025

Unveiling the Black Box: Independent Functional Module Evaluation for Bird's-Eye-View Perception Model

ICRA 2025

End-to-end models are emerging as the mainstream in autonomous driving perception. However, the inability to meticulously deconstruct their internal mechanisms results in diminished development efficacy and impedes the establishment of trust. Pioneering in the issue, we present the Independent Funct

Cited by 1SourceScholar
2025

Vision-Driven 2D Supervised Fine-Tuning Framework for Bird's Eye View Perception

IROS 2025

Visual bird’s eye view (BEV) perception, dute to its excellent perceptual capabilities, is progressively replacing costly LiDAR-based perception systems, especially in the realm of urban intelligent driving. However, this type of perception still relies on LiDAR data to construct ground truth databa

Cited by 2SourceScholar
2025

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

ICASSP 2025accepted

Style voice conversion aims to transform the speaking style of source speech into a desired style while keeping the original speaker’s identity. However, previous style voice conversion approaches primarily focus on well-defined domains such as emotional aspects, limiting their practical application…

Cited by 0SourceScholar
2024

CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations

NeurIPS 2024poster

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conver…

2024

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

ICLR 2024spotlight

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large TTS systems usually quantize speech into discrete tokens and use language models…

2024

NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

ICML 2024oral

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall shorts in speech quality, similarity, and prosody. Considering that speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant…

Cited by 172SourcePDFScholar
2024

PromptTTS 2: Describing and Generating Voices with Text Prompt

ICLR 2024poster

Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods relying on speech prompts (reference speech) for voice variability, using text prompts (descriptions) is more user-friendly…

2024

Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis

ICASSP 2024accepted

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-supervised style enhancing method with VQ-VAE-based pre-training for expressive a…

Cited by 0SourceScholar
2023

AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models

NeurIPS 2023poster

Audio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio. Recently, some diffusion-based methods achieved zero-shot audio editing by using a diffusion and denoising process conditioned on the text descripti…

2023

Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation

ICASSP 2023accepted

Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from the speech of the source language to the speech of the target language are very rare. To add…

Cited by 0SourceScholar
2023

KEPL: Knowledge Enhanced Prompt Learning for Chinese Hypernym-Hyponym Extraction

EMNLP 2023long main

Modeling hypernym-hyponym ("is-a") relations is very important for many natural language processing (NLP) tasks, such as classification, natural language inference and relation extraction. Existing work on is-a relation extraction is mostly in the English language environment. Due to the flexibility…

Cited by 0SourceScholar
2023

LeanSpeech: The Microsoft Lightweight Speech Synthesis System for Limmits Challenge 2023

ICASSP 2023accepted

This paper describes the Microsoft Text-to-Speech (TTS) system: LeanSpeech for LIMMITS (Lightweight, Multi-speaker, Multi-lingual Indic TTS) Challenge 2023<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>, which is part of ICASSP2023 to encourage…

Cited by 0SourceScholar
2023

VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing

AAAI 2023technical

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech synthesis. To ensure the translated speech to be well aligned with t…

2022

BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio Synthesis

NeurIPS 2022accept

Binaural audio plays a significant role in constructing immersive augmented and virtual realities. As it is expensive to record binaural audio from the real world, synthesizing them from mono audio has attracted increasing attention. This synthesis process involves not only the basic physical warpin…

2022

ConCL: Concept Contrastive Learning for Dense Prediction Pre-training in Pathology Images

ECCV 2022poster

"Detecting and segmenting objects within whole slide images is essential in computational pathology workflow. Self-supervised learning (SSL) is appealing to such annotation-heavy tasks. Despite the extensive benchmarks in natural images for dense tasks, such studies are, unfortunately, absent in cur…

2022

Diff-Net: Image Feature Difference Based High-Definition Map Change Detection for Autonomous Driving

ICRA 2022poster

Up-to-date High-Definition (HD) maps are essential for self-driving cars. To achieve constantly updated HD maps, we present a deep neural network (DNN), Diff-Net, to detect changes in them. Compared to traditional methods based on object detectors, the essential design in our work is a parallel feat…

Cited by 8SourceScholar
2022

Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward Network

ICASSP 2022accepted

FastSpeech, as a feed-forward transformer based TTS, can avoid the slow serial, autoregressive inference to generate the target mel-spectrogram in a parallel way. As a non-autoregressive TTS, the latency and computation load in inference is shifted from vocoder to transformer where the efficiency is…

Cited by 0SourceScholar
2022

Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training

ICASSP 2022accepted

Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, which invariably results in slow inference speed. Previous approaches aim to opt…

Cited by 0SourceScholar
2022

Prosodyspeech: Towards Advanced Prosody Model for Neural Text-to-Speech

ICASSP 2022accepted

This paper proposes ProsodySpeech, a novel prosody model to enhance encoder-decoder neural Text-To-Speech (TTS), to generate high expressive and personalized speech even with very limited training data. First, we use a Prosody Extractor built from a large speech corpus with various speakers to gener…

Cited by 0SourceScholar
2022

TreeMoCo: Contrastive Neuron Morphology Representation Learning

NeurIPS 2022accept

Morphology of neuron trees is a key indicator to delineate neuronal cell-types, analyze brain development process, and evaluate pathological changes in neurological diseases. Traditional analysis mostly relies on heuristic features and visual inspections. A quantitative, informative, and comprehensi…

2021

CLMM-Net: Robust Cascaded LiDAR Map Matching based on Multi-Level Intensity Map

IROS 2021poster

LiDAR map matching(LMM) is a critical localization technique in autonomous driving while existing methods have problems in terms of both accuracy and robustness when driving in the scenes with poor structure information (e.g. highways). This paper put forward a multi-level intensity map based cascad…

Cited by 0SourceScholar
2021

Exploring Forensic Dental Identification with Deep Learning

NeurIPS 2021poster

Dental forensic identification targets to identify persons with dental traces. The task is vital for the investigation of criminal scenes and mass disasters because of the resistance of dental structures and the wide-existence of dental imaging. However, no widely accepted automated solution is ava…

2021

KLMo: Knowledge Graph Enhanced Pretrained Language Model with Fine-Grained Relationships

EMNLP 2021finding

Interactions between entities in knowledge graph (KG) provide rich knowledge for language representation learning. However, existing knowledge-enhanced pretrained language models (PLMs) only focus on entity information and ignore the fine-grained relationships between entities. In this work, we prop…

2021

Oral-3D: Reconstructing the 3D Structure of Oral Cavity from Panoramic X-ray

AAAI 2021technical

Panoramic X-ray (PX) provides a 2D picture of the patient's mouth in a panoramic view to help dentists observe the invisible disease inside the gum. However, it provides limited 2D information compared with cone-beam computed tomography (CBCT), another dental imaging method that generates a 3D pictu…

Cited by 41SourcePDFScholar
2020

Adaptation of RNN Transducer with Text-To-Speech Technology for Keyword Spotting

ICASSP 2020accepted

With the advent of recurrent neural network transducer (RNN-T) model, the performance of keyword spotting (KWS) systems has greatly improved. However, the KWS systems, employed for wake-word detection, still rely on the availability of keyword specific training data for achieving reasonable performa…

Cited by 0SourceScholar
2020

Improving Prosody with Linguistic and Bert Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTS

ICASSP 2020accepted

Recent advances of neural TTS have made "human parity" synthesized speech possible when a large amount of studio-quality training data from a voice talent is available. However, with only limited, casual recordings from an ordinary speaker, human-like TTS is still a big challenge, in addition to oth…

Cited by 0SourceScholar
2020

Integrated moment-based LGMD and deep reinforcement learning for UAV obstacle avoidance

ICRA 2020poster

In this paper, a bio-inspired monocular vision perception method combined with a learning-based reaction local planner for obstacle avoidance of micro UAVs is presented. The system is more computationally efficient than other vision-based perception and navigation methods such as SLAM and optical fl…

Cited by 43SourceScholar
2020

Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker Adaptation

ICASSP 2020accepted

We propose to use the personalized speech synthesis and the neural language generator to synthesize content relevant personalized speech for rapid speaker adaptation. It has two distinct aspects: First, it relieves the general data sparsity issue in rapid adaptation via making use of additional synt…

Cited by 32SourceScholar
2019

Learning Latent Representations for Style Control and Transfer in End-to-end Speech Synthesis

ICASSP 2019accepted

In this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good properties such as disentangling, scaling, and combination,…

Cited by 0SourceScholar
2015

Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesis

ICASSP 2015accepted

In DNN-based TTS synthesis, DNNs hidden layers can be viewed as deep transformation for linguistic features and the output layers as representation of acoustic space to regress the transformed linguistic features to acoustic parameters. The deep-layered architectures of DNN can not only represent hi…

Cited by 0SourceScholar
2015

Word embedding for recurrent neural network based TTS synthesis

ICASSP 2015accepted

The current state of the art TTS synthesis can produce synthesized speech with highly decent quality if rich segmental and suprasegmental information are given. However, some suprasegmental features, e.g., Tone and Break (TOBI), are time consuming due to being manually labeled with a high inconsiste…

Cited by 57SourceScholar