← Search

Xun Gong

17 accepted papers

2026

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

CVPR 2026

Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals--without sacrificing interpretability and traceability--remai

Cited by 0SourceScholar
2026

Zero-Shot Image Denoising via Hybrid Prior-Guided Pseudo Sample Generation

CVPR 2026

Zero-shot image denoising has gained prominence in recent years, as it inherently relies on the intrinsic priors of images rather than learning from external data. Nevertheless, most existing methods either fail to fully exploit global priors, or do not properly preserve the fine-grained details gov

Cited by 0SourceScholar
2025

Gate to the Vessel: Residual Experts Restore What SAM Overlooks

NeurIPS 2025poster

Foundation segmentation models like Segment Anything (SAM) exhibit strong generalization on natural images but struggle with localized failures in medical imaging, especially on fine-grained structures such as vessels with complex morphology and indistinct boundaries. To address this, we propose Fin…

Cited by 0SourceScholar
2025

Harnessing and Evaluating the Intrinsic Extrapolation Ability of Large Language Models for Vehicle Trajectory Prediction

NAACL 2025long

Emergent abilities of large language models (LLMs) have significantly advanced their application in autonomous vehicle (AV) research. Safe integration of LLMs into vehicles, however, necessitates their thorough understanding of dynamic traffic environments. Towards this end, this study introduces a…

Cited by 0SourcePDFScholar
2025

Learning Normals of Noisy Points by Local Gradient-Aware Surface Filtering

ICCV 2025poster

Estimating normals for noisy point clouds is a persistent challenge in 3D geometry processing, particularly for end-to-end oriented normal estimation. Existing methods generally address relatively clean data and rely on supervised priors to fit local surfaces within specific neighborhoods. In this p…

2025

RepFace: Refining Closed-Set Noise with Progressive Label Correction for Face Recognition

AAAI 2025technical

Face recognition has made remarkable strides, driven by the expanding scale of datasets, advancements in various backbone and discriminative losses. However, face recognition performance is heavily affected by the label noise, especially closed-set noise. While numerous studies have focused on handl…

Cited by 0SourcePDFScholar
2025

VA-GS: Enhancing the Geometric Representation of Gaussian Splatting via View Alignment

NeurIPS 2025poster

3D Gaussian Splatting has recently emerged as an efficient solution for high-quality and real-time novel view synthesis. However, its capability for accurate surface reconstruction remains underexplored. Due to the discrete and unstructured nature of Gaussians, supervision based solely on image rend…

Cited by 0SourcecodeScholar
2023

Adversarial Driving Behavior Generation Incorporating Human Risk Cognition for Autonomous Vehicle Evaluation

IROS 2023poster

Autonomous vehicle (AV) evaluation has been the subject of increased interest in recent years both in industry and in academia. This paper focuses on the development of a novel framework for generating adversarial driving behavior of background vehicle interfering against the AV to expose effective…

Cited by 1SourceScholar
2023

Factorized AED: Factorized Attention-Based Encoder-Decoder for Text-Only Domain Adaptive ASR

ICASSP 2023accepted

End-to-end automatic speech recognition (ASR) systems have gained popularity given their simplified architecture and promising results. However, text-only domain adaptation remains a big challenge for E2E systems. Text-to-speech (TTS) based approaches fine-tune ASR models by synthesized speech with…

Cited by 0SourceScholar
2023

Joint Discriminator and Transfer Based Fast Domain Adaptation For End-To-End Speech Recognition

ICASSP 2023accepted

Adapting End-to-End (E2E) models to unseen domains is still a big challenge since training E2E models requires lots of paired audio and text training data. We propose a novel domain adaptation framework for the E2E model, which only uses the text of the target domain. Moreover, the proposed methods…

Cited by 0SourceScholar
2023

LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer

ICASSP 2023accepted

Traditional automatic speech recognition (ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is more practical in real scenarios. Simply attending longer transcription history for a vanilla neural transducer model shows…

Cited by 0SourceScholar
2022

BoundaryFace: A Mining Framework with Noise Label Self-Correction for Face Recognition

ECCV 2022poster

"Face recognition has made tremendous progress in recent years due to the advances in loss functions and the explosive growth in training sets size. A properly designed loss is seen as key to extract discriminative features for classification. Several margin-based losses have been proposed as altern…

2022

The Sjtu System For Multimodal Information Based Speech Processing Challenge 2021

ICASSP 2022accepted

This paper describes the SJTU system for ICASSP Multi-modal Information based Speech Processing Challenge (MISP) 2021. To solve the speech recognition problem in real complex environments where time-synchronized near- and far-field signals are available for training an enhancement frontend. We build…

Cited by 0SourceScholar
2020

Text Adaptation for Speaker Verification with Speaker-Text Factorized Embeddings

ICASSP 2020accepted

Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully collecting data with the target speech content, such data c…

Cited by 0SourceScholar