← Search

Shengwu Xiong

27 accepted papers

2026

Driving with Advice: Large Model as Motion Advisor for Joint Planning

AAAI 2026technical

We address the challenge of integrating high-level semantic reasoning with low-level trajectory planning in end-to-end autonomous driving, where most existing frameworks decouple perception, decision-making, and control, leading to limited interpretability and poor instruction compliance. To bridge

Cited by 0SourcePDFScholar
2026

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. Existing benchmarks are usually constructed in a task-oriented manner, without a guarantee…

Cited by 0SourceScholar
2026

ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-Labeling

AAAI 2026technical

Existing approaches for the problem of ultrasound image segmentation, whether supervised or semi-supervised, are typically specialized for specific anatomical structures or tasks, limiting their practical utility in clinical settings. In this paper, we pioneer the task of universal semi-supervised u

Cited by 0SourcePDFScholar
2025

Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot Segmentation

ICCV 2025poster

This paper studies the few-shot segmentation (FSS) task, which aims to segment objects belonging to unseen categories in a query image by learning a model on a small number of well-annotated support samples. Our analysis of two mainstream FSS paradigms reveals that the predictions made by prototype…

2025

Enhancing Zero-Shot Relation Extraction through Staged Interaction with Large Language Models

ICASSP 2025accepted

Zero-shot Relation Triplet Extraction (ZeroRTE) is a challenging yet valuable task that extracts relation triplets from unstructured texts for new relation types, significantly reducing the time and effort needed for data labeling. With the advancement of the zero-shot capabilities of large language…

Cited by 0SourceScholar
2025

GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free Approach

AAAI 2025technical

A good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-deco…

Cited by 0SourcePDFScholar
2025

Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label Generation

AAAI 2025technical

Efficiently applying fully supervised learning to virtual try-on tasks is challenging due to the lack of paired ground truth in available training samples. Recent works have achieved virtual try-ons by employing self-supervised learning-based inpainting paradigms. However, this approach is heavily d…

Cited by 0SourcePDFScholar
2025

Mask Does Not Matter: A Unified Latent Diffusion-Enhanced Framework for Mask-Free Virtual Try-On

IJCAI 2025

A good virtual try-on model should introduce minimal redundant conditional information to avoid instability and increase inference efficiency. Existing methods rely on inpainting masks to guide the generation of the object, but the masks, generated by unstable human parsers, often produce unreliable

Cited by 0SourcePDFScholar
2025

Mitigating Occlusions in Virtual Try-On via A Simple-Yet-Effective Mask-Free Framework

NeurIPS 2025poster

This paper investigates the occlusion problems in virtual try-on (VTON) tasks. According to how they affect the try-on results, the occlusion issues of existing VTON methods can be grouped into two categories: (1) Inherent Occlusions, which are the ghosts of the clothing from reference input images…

Cited by 0SourceScholar
2025

ROME: Radar Sparsity Improvement and Omnimodal Enhancement for 3D Object Detection in Bird's Eye Views

ICASSP 2025accepted

Combining omnimodal feature interaction using LiDAR, surround-view camera, and Radar to form a network has a great guarantee for the safety of autonomous driving, but most of the current omnimodal fusion methods focus on the interaction enhancement of LiDAR and surround-view camera, ignoring the foc…

Cited by 0SourceScholar
2025

RealisID: Scale-Robust and Fine-Controllable Identity Customization via Local and Global Complementation

AAAI 2025technical

Recently, the success of text-to-image synthesis has greatly advanced the development of identity customization techniques, whose main goal is to produce realistic identity-specific photographs based on text prompts and reference face images. However, it is difficult for existing identity customizat…

2024

Content-Style Decoupling for Unsupervised Makeup Transfer without Generating Pseudo Ground Truth

CVPR 2024poster

The absence of real targets to guide the model training is one of the main problems with the makeup transfer task. Most existing methods tackle this problem by synthesizing pseudo ground truths (PGTs). However the generated PGTs are often sub-optimal and their imprecision will eventually lead to per…

2024

CycleVTON: A Cycle Mapping Framework for Parser-Free Virtual Try-On

AAAI 2024technical

Image-based virtual try-on aims to transfer a target clothing onto a specific person. A significant challenge is arbitrarily matched clothing and person lack corresponding ground truth to supervised learning. A recent pioneering work leveraged an improved cycleGAN to enable one network to generate t…

Cited by 2SourcePDFScholar
2024

IFNET: Integrating Data Augmentation and Decoupled Attention Fusion for 3D Object Detection

ICASSP 2024accepted

LiDAR is a key sensor for accurately sensing of the environment in autonomous driving. While existing 3D object detection methods generally rely on data augmentation and feature fusion to improve performance, the challenge of dealing with sample imbalance is often overlooked. We design a novel 3D de…

Cited by 0SourceScholar
2024

SHMT: Self-supervised Hierarchical Makeup Transfer via Latent Diffusion Models

NeurIPS 2024poster

This paper studies the challenging task of makeup transfer, which aims to apply diverse makeup styles precisely and naturally to a given facial image. Due to the absence of paired data, current methods typically synthesize sub-optimal pseudo ground truths to guide the model training, resulting in l…

2023

ESPT: A Self-Supervised Episodic Spatial Pretext Task for Improving Few-Shot Learning

AAAI 2023technical

Self-supervised learning (SSL) techniques have recently been integrated into the few-shot learning (FSL) framework and have shown promising results in improving the few-shot image classification performance. However, existing SSL approaches used in FSL typically seek the supervision signals from the…

2023

Greatness in Simplicity: Unified Self-Cycle Consistency for Parser-Free Virtual Try-On

NeurIPS 2023poster

Image-based virtual try-on tasks remain challenging, primarily due to inherent complexities associated with non-rigid garment deformation modeling and strong feature entanglement of clothing within human body. Recent groundbreaking formulations, such as in-painting, cycle consistency, and knowledge…

Cited by 8SourcePDFScholar
2022

A Frame Loss of Multiple Instance Learning for Weakly Supervised Sound Event Detection

ICASSP 2022accepted

Sound event detection(SED) consists of two subtasks: predicting the classes of sound events within an audio clip (audio tagging) and indicating the onset and offset times for each event (localization). One of the common approaches for SED with weak label is multiple instance learning (MIL) method. H…

Cited by 0SourceScholar
2022

SSAT: A Symmetric Semantic-Aware Transformer Network for Makeup Transfer and Removal

AAAI 2022technical

Makeup transfer is not only to extract the makeup style of the reference image, but also to render the makeup style to the semantic corresponding position of the target image. However, most existing methods focus on the former and ignore the latter, resulting in a failure to achieve desired results.…

Cited by 44SourcePDFScholar
2021

Accurate and Robust Stereo Direct Visual Odometry for Agricultural Environment

ICRA 2021poster

Vision-based localization and mapping in the agricultural environment is challenging due to the unstructured scene with unstable features, illumination variations, bumpy roads, and dynamic environmental objects. To address these challenges, we propose an accurate and robust stereo direct visual odom…

Cited by 10SourceScholar
2021

Benchmark Platform for Ultra-Fine-Grained Visual Categorization Beyond Human Performance

ICCV 2021poster

Deep learning methods have achieved remarkable success in fine-grained visual categorization. Such successful categorization at sub-ordinate level, e.g., different animal or plant species, however relies heavily on the visual differences that human can observe and the ground-truths are labelled on t…

Cited by 37PDFcodeScholar
2021

Neural Noise Embedding for End-To-End Speech Enhancement with Conditional Layer Normalization

ICASSP 2021accepted

Most of the deep learning based speech enhancement methods focus on the modeling of complicated relationship between the noisy speech and the clean speech without the consideration of noise information. In order to cope with various complex noise scenes, we introduce a novel enhancement architecture…

Cited by 0SourceScholar
2020

A Time-Frequency Network with Channel Attention and Non-Local Modules for Artificial Bandwidth Extension

ICASSP 2020accepted

Convolution neural networks (CNNs) have been achieving increasing attention for the artificial bandwidth extension (ABE) task recently. However, these methods use the flipped low-frequency phase to reconstruct speech signals, which may lead to the well-known invalid short-time Fourier Transform (STF…

Cited by 0SourceScholar
2019

Densely Connected Network with Time-frequency Dilated Convolution for Speech Enhancement

ICASSP 2019accepted

The data driven speech enhancement approaches using regression-based deep neural network usually result in enormous number of model parameters, which increase the computational load and the difficulty of model training. In order to improve the model efficiency, we propose a densely connected network…

Cited by 0SourceScholar
2018

A Second-Order Variational Framework for Joint Depth Map Estimation and Image Dehazing

ICASSP 2018accepted

Outdoor images captured in poor weather conditions (e.g., fog or haze) commonly suffer from reduced contrast and visibility. Increasing attention has recently been paid to single image dehazing, i.e., improving image contrast and visibility. It is generally thought that the dehazing performance high…

Cited by 0SourceScholar