← Search

Kun Song

13 accepted papers

2025

A Planning Framework for Complex Flipping Manipulation of Multiple Mobile Manipulators

RA-L 2025

During complex object manipulation, manipulator systems often face the configuration disconnectivity problem due to closed-chain constraints. Although regrasping can be adopted to guarantee connectivity, it introduces additional issues such as impact and efficiency. Therefore, regrasping numbers sho

Cited by 5SourceScholar
2025

Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update

AAAI 2025technical

Warning: This paper contains offensive content that may disturb some readers. Vision-language models (VLMs) demonstrate strong multimodal capabilities but have been found to be more susceptible to generating harmful content compared to their backbone large language models (LLMs). Our investigation r…

2025

P2 Explore: Efficient Exploration in Unknown Cluttered Environment with Floor Plan Prediction

IROS 2025

Robot exploration aims at the reconstruction of unknown environments, and it is important to achieve it with shorter paths. Traditional methods focus on optimizing the visiting order of frontiers based on current observations, which may lead to local-minimal results. Recently, by predicting the stru

Cited by 4SourcecodeScholar
2025

ProAPO: Progressively Automatic Prompt Optimization for Visual Classification

CVPR 2025poster

Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the prompt quality. While recent methods show that visual descriptions generated by large language models (LLMs) enhance the…

2024

FHT-Map: Feature-Based Hybrid Topological Map for Relocalization and Path Planning

RA-L 2024

Topological maps are favorable for their small storage compared to geometric maps. However, they are limited in relocalization and path planning capabilities. To solve the problem, a feature-based hybrid topological map (FHT-Map) is proposed along with a real-time map construction algorithm based on

Cited by 8SourcecodeScholar
2024

Multi-Robot Rendezvous in Unknown Environment With Limited Communication

RA-L 2024

Rendezvous aims at gathering all robots at a specific location, which is an important collaborative behavior for multi-robot systems. However, in an unknown environment, it is challenging to achieve rendezvous. Previous researches mainly focus on special scenarios where communication is not allowed

Cited by 4SourcecodeScholar
2024

RHAML: Rendezvous-Based Hierarchical Architecture for Mutual Localization

RA-L 2024

Mutual localization serves as the foundation for collaborative perception in multi-robot systems. Effectively utilizing limited onboard sensors for mutual localization between marker-less robots is worthwhile. However, due to inadequate consideration of large scale variations of the robot and locali

Cited by 2SourceScholar
2024

Robustly Train Normalizing Flows via KL Divergence Regularization

AAAI 2024technical

In this paper, we find that the training of Normalizing Flows (NFs) are easily affected by the outliers and a small number (or high dimensionality) of training samples. To solve this problem, we propose a Kullback–Leibler (KL) divergence regularization on the Jacobian matrix of NFs. We prove that su…

Cited by 4SourcePDFScholar
2024

Unveiling the Dynamics of Information Interplay in Supervised Learning

ICML 2024poster

In this paper, we use matrix information theory as an analytical tool to analyze the dynamics of the information interplay between data representations and classification head vectors in the supervised learning process. Specifically, inspired by the theory of Neural Collapse, we introduce matrix mut…

Cited by 3SourcePDFScholar
2023

DSPGAN: A Gan-Based Universal Vocoder for High-Fidelity TTS by Time-Frequency Domain Supervision from DSP

ICASSP 2023accepted

Recent development of neural vocoders based on the generative adversarial neural network (GAN) has shown obvious advantages of generating raw waveform conditioned on mel-spectrogram with fast inference speed and lightweight networks. Whereas, it is still challenging to train a universal neural vocod…

Cited by 0SourceScholar
2023

FD-Align: Feature Discrimination Alignment for Fine-tuning Pre-Trained Models in Few-Shot Learning

NeurIPS 2023poster

Due to the limited availability of data, existing few-shot learning methods trained from scratch fail to achieve satisfactory performance. In contrast, large-scale pre-trained models such as CLIP demonstrate remarkable few-shot and zero-shot capabilities. To enhance the performance of pre-trained mo…

2023

Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling

ICASSP 2023accepted

This paper aims to synthesize the target speaker’s speech with desired speaking style and emotion by transferring the style and emotion from reference speech recorded by other speakers. We address this challenging problem with a two-stage framework composed of a text-to-style-and-emotion (Text2SE) m…

Cited by 0SourceScholar
2021

The Multi-Speaker Multi-Style Voice Cloning Challenge 2021

ICASSP 2021accepted

The Multi-speaker Multi-style Voice Cloning Challenge (M2VoC) aims to provide a common sizable dataset as well as a fair testbed for the benchmarking of the popular voice cloning task. Specifically, we formulate the challenge to adapt an average TTS model to the stylistic target voice with limited d…

Cited by 0SourceScholar