← Search

Song Wu

20 accepted papers

2026

A Training-Free Framework for High-Fidelity Appearance Transfer via Diffusion Transformers

ICASSP 2026poster

Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic scene structure. We address this by proposing the first tra…

Cited by 0SourcePDFScholar
2026

APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt Tuning

ICLR 2026poster

Scene Graph Generation (SGG) is pivotal for structured visual understanding, yet it remains hindered by a fundamental limitation: the reliance on fixed, frozen semantic representations from pre-trained language models. These semantic priors, while beneficial in other domains, are inherently misalign…

Cited by 0SourcecodeScholar
2026

CPBA-LIWO: Continuous-Time LiDAR-Inertial-Wheel Odometry Based on Probabilistic Bundle Adjustment

ICRA 2026poster

LiDAR-based odometry is widely used in ground robot localization. However, current methods encounter challenges in accuracy and robustness due to structural degradation, system observational error, and accumulated error. To address the above issues, we propose CPBA-LIWO, a continuous-time LiDAR-Iner…

Cited by 0Scholar
2026

From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation

CVPR 2026

Estimating continuous optical flow is a fundamental yet challenging problem in dynamic visual perception. Event-based cameras, with microsecond latency and high dynamic range, capture brightness changes asynchronously, offering a unique opportunity to model motion with fine temporal precision. Howev

Cited by 0SourceScholar
2026

RLAP-CLIP: Continual Multimodal Learning with Prototype Adaptation and Difficulty-Aware Routing

ICLR 2026poster

Vision-language models, such as CLIP, achieve strong zero-shot performance through contrastive pre-training but face significant challenges in class-incremental image classification scenarios. When learning new tasks sequentially, current methods suffer from degradation in prototype quality due to p…

Cited by 0SourceScholar
2026

Views Attention Fusion of Granular-ball Fuzzy Representations Split for Improved Multi-view Clustering

AAAI 2026technical

Multi-View Clustering (MVC) is a pivotal multi-view learning paradigm widely adopted across various fields. Despite recent advances, existing methods primarily focus on enhancing the performance of fused multi-view representation, often neglecting the issue of Representation Degradation (RD) arising

Cited by 0SourcePDFScholar
2025

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

NeurIPS 2025poster

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference…

Cited by 0SourceScholar
2025

MSPA-LIO: LiDAR-Inertial Odometry with Multi-Scale Plane Adjustment

IROS 2025

Most current LiDAR-based odometry methods use point-to-local plane registration to constrain poses, ignoring the explicit plane structure in the environment. Due to noise interference and uneven distribution of point cloud, local planes are prone to tilt, resulting in registration errors. Therefore,

Cited by 0SourceScholar
2025

Sim-LLM: Optimizing LLM Inference at the Edge through Inter-Task KV Reuse

NeurIPS 2025poster

KV cache technology, by storing key-value pairs, helps reduce the computational overhead incurred by *large language models* (LLMs). It facilitates their deployment on resource-constrained edge computing nodes like edge servers. However, as the complexity and size of tasks increase, KV cache usage l…

Cited by 0SourcecodeScholar
2025

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

ICCV 2025poster

Recent advances in computational pathology have introduced whole slide image (WSI)-level multimodal large language models (MLLMs) for automated pathological analysis. However, current WSI-level MLLMs face two critical challenges: limited explainability in their decision-making process and insufficie…

Cited by 0SourcePDFScholar
2024

Arbitrary Style Transfer Based on Content Integrity and Style Consistency Enhancement

ICASSP 2024accepted

The existing arbitrary style transfer methods mainly suffer two challenges. One is content integrity, as most methods focus too much on style, resulting in incomplete content information and missing details. The other is style consistency, which requires more exploration of style information to alle…

Cited by 0SourceScholar
2024

Cross-View Contrastive Fusion for Enhanced Molecular Property Prediction

IJCAI 2024poster

Machine learning based molecular property prediction has been a hot topic in the field of computer aided drug discovery (CADD). However, current MPP methods face two prominent challenges: 1) single-view MPP methods do not sufficiently exploit the complementary information of molecular data across mu…

Cited by 1SourcePDFScholar
2024

E-Motion: Future Motion Simulation via Event Sequence Diffusion

NeurIPS 2024poster

Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict fu…

2024

LA-LIO: Robust Localizability-Aware LiDAR-Inertial Odometry for Challenging Scenes

IROS 2024poster

Modern robotic systems are increasingly deployed in complex and diverse environments, and reliable localization under challenging conditions becomes crucial for the safe and efficient operation of these systems. The odometry based on LiDAR is prone to system collapse caused by computational divergen…

Cited by 1SourceScholar
2024

Multi-Scale Fusion of Gated Neighborhood Attention Transformers for Single Image Deraining

ICASSP 2024accepted

Since the diverse geometric appearances and densities of rain streaks, local-global information is equally essential for single image deraining. Balancing local-global information becomes a challenge. Thus, a Multi-Scale Fusion of Gated Neighborhood Attention Transformers (MSF-GNAT) for single image…

Cited by 0SourceScholar
2022

Bi-Directional Normalization and Color Attention-Guided Generative Adversarial Network for Image Enhancement

ICASSP 2022accepted

Most existing image enhancement methods require paired images, and rarely consider the aesthetic quality. This paper proposes a bi-directional normalization and color attention-guided generative adversarial network (BNCAGAN) for unsupervised image enhancement. An auxiliary attention classifier (AAC)…

Cited by 0SourceScholar
2022

MS-ROCANet: Multi-Scale Residual Orthogonal-Channel Attention Network for Scene Text Detection

ICASSP 2022accepted

Deep neural networks-based scene text detection has obtained increasing attention in recent years. However, the existing scene text detection methods cannot effectively solve the problem of unclear text features. In this paper, a Multi-scale Residual Orthogonal-Channel Attention Network (MS-ROCANet)…

Cited by 0SourceScholar
2022

Video Interpolation by Event-Driven Anisotropic Adjustment of Optical Flow

ECCV 2022poster

"Video frame interpolation is a challenging task due to the ever-changing real-world scene. Previous methods often calculate the bi-directional optical flows and then predict the intermediate optical flows under the linear motion assumptions, leading to isotropic intermediate flow generation. Follow…

Cited by 15SourcePDFScholar
2021

PointLIE: Locally Invertible Embedding for Point Cloud Sampling and Recovery

IJCAI 2021poster

Point Cloud Sampling and Recovery (PCSR) is critical for massive real-time point cloud collection and processing since raw data usually requires large storage and computation. This paper addresses a fundamental problem in PCSR: How to downsample the dense point cloud with arbitrary scales while pres…

2018

Cell Subclass Identification in Single-Cell RNA-Sequencing Data Using Orthogonal Nonnegative Matrix Factorization

ICASSP 2018accepted

Identification of cell subclasses using single-cell RNA-Sequencing (scRNA-Seq) data is of paramount importance since it uncovers the hidden biological processes within the cell population. While the nonnegative matrix factorization (NMF) model has been reported to be effective in various unsupervise…

Cited by 0SourceScholar