← Search

Kai Wu

40 accepted papers

2026

$\texttt{MetaDistill}$: Unlocking the Performance Ceiling for Pretrained Optimizers

ICML 2026poster

Meta Black-Box Optimization (MetaBBO) has emerged as a promising paradigm by employing meta learning to automatically optimize the configurations of low-level black-box optimizers. Despite its potential, the generalization of MetaBBO remains significantly constrained when facing unseen, complex obje…

Cited by 0SourceScholar
2026

A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space

CVPR 2026

Innovative visual stylization is a cornerstone of artistic creation, yet generating novel and consistent visual styles remains a significant challenge. Existing generative approaches typically rely on lengthy textual prompts, reference images, or parameter-efficient fine-tuning to guide style-aware

Cited by 0SourcecodeScholar
2026

BOLT: Decision‑Aligned Distillation and Budget-Aware Routing for Constrained Multimodal QA on Robots

ICLR 2026poster

Robotic systems can require multimodal reasoning under stringent constraints of latency, memory, and energy. Standard instruction tuning and token-level distillation fail to deliver decision quality, reliability, and interpretability under these constraints. We introduce BOLT, a decision-aligned dis…

Cited by 0SourceScholar
2026

BaseReward: A Strong Baseline for Multimodal Reward Model

ICLR 2026poster

The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for achieving this goal, but a systematic guide for building state-of-the-art Multimodal Reward Models (MRMs) is currently l…

Cited by 0SourceScholar
2026

DiffVecMap: A Robust Online Vectorized HD Map Construction Method With a Diffusion Model

RA-L 2026

High-definition (HD) map construction plays a critical role in providing precise and comprehensive static environmental information for autonomous driving systems. However, the performance of existing methods degrades significantly under adverse conditions such as nighttime or rainy weather, where t

Cited by 0SourceScholar
2026

Group-wise Data Ordering: Enhancing Instruction Tuning of Large Language Models via Embedding Proximity

ICML 2026poster

Instruction tuning (IT) is a central mechanism for aligning large language models (LLMs) with user intent. In practice, randomly shuffling the training set is a simple yet surprisingly strong baseline. However, it overlooks latent structure, such as domain and reasoning depth, and thus interleaves h…

Cited by 0SourceScholar
2026

Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control

CVPR 2026

Physics-based humanoid control relies on training with motion datasets that have diverse data distributions. However, the fixed difficulty distribution of datasets limits the performance ceiling of the trained control policies. Additionally, the method of acquiring high-quality data through professi

Cited by 0SourceScholar
2026

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

CVPR 2026

MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from existing counterparts by three core features: 1) Systematic capab

Cited by 0SourcecodeScholar
2026

MedLesionVQA: A Multimodal Benchmark Emulating Clinical Visual Diagnosis for Body Surface Health

ICLR 2026poster

Body-surface health conditions, spanning diverse clinical departments, represent some of the most frequent diagnostic scenarios and a primary target for medical multimodal large language models (MLLMs). Yet existing medical benchmarks are either built from publicly available sources with limited ex…

Cited by 0SourceScholar
2026

Principled RL for Flow Matching Emerges From the Chunk-level Policy Optimization

ICML 2026poster

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive …

Cited by 0SourceScholar
2026

Stochastic Universal Adversarial Perturbations with Fixed Optimization Constraint and Ensured High-probability Transferability

AAAI 2026technical

Adversarial perturbations (APs) have become a great concern in image classification tasks. The most challenging branch, universal adversarial perturbations (UAPs), are exploited to fool most of the unseen samples. Such one-to-all perturbations have the merit of transferability, which has strong prac

Cited by 0SourcePDFScholar
2026

Textual Self-Attention Network: Test-Time Preference Optimization Through Textual Gradient-Based Attention

AAAI 2026technical

Large Language Models (LLMs) have demonstrated remarkable generalization capabilities, but aligning their outputs with human preferences typically requires expensive supervised fine-tuning. Recent test-time methods leverage textual feedback to overcome this, but they often critique and revise a sing

Cited by 0SourcePDFScholar
2026

VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction

CVPR 2026

Unifying multimodal understanding, generation and reconstruction representation in a single tokenizer remains a key challenge in building unified models. Previous research predominantly attempts to address this in a dual encoder paradigm, e.g., utilizing the separate encoders for understanding and g

Cited by 0SourcecodeScholar
2025

Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing

CVPR 2025poster

Due to control limitations in the denoising process and the lack of training, zero-shot video editing methods often struggle to meet user instructions, resulting in generated videos that are visually unappealing and fail to fully satisfy expectations. To address this problem, we propose Align-A-Vide…

Cited by 0SourcePDFScholar
2025

AutoSGNN: Automatic Propagation Mechanism Discovery for Spectral Graph Neural Networks

AAAI 2025technical

In real-world applications, spectral Graph Neural Networks (GNNs) are powerful tools for processing diverse types of graphs. However, a single GNN often struggles to handle different graph types—such as homogeneous and heterogeneous graphs—simultaneously. This challenge has led to the manual design…

2025

B2Opt: Learning to Optimize Black-box Optimization with Little Budget

AAAI 2025technical

The core challenge of high-dimensional and expensive black-box optimization (BBO) is how to obtain better performance faster with little function evaluation cost. The essence of the problem is how to design an efficient optimization strategy tailored to the target task. This paper designs a powerful…

Cited by 10SourcePDFScholar
2025

CustAny: Customizing Anything from A Single Example

CVPR 2025poster

Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, is still challenging.Object customization, using reference images and textual descriptions, is key to addressing this iss…

2025

Enhancing Zero-Shot Black-Box Optimization via Pretrained Models with Efficient Population Modeling, Interaction, and Stable Gradient Approximation

NeurIPS 2025poster

Zero-shot optimization aims to achieve both generalization and performance gains on solving previously unseen black-box optimization problems over SOTA methods without task-specific tuning. Pre-trained optimization models (POMs) address this challenge by learning a general mapping from task features…

Cited by 0SourceScholar
2025

PoisonedEye: Knowledge Poisoning Attack on Retrieval-Augmented Generation based Large Vision-Language Models

ICML 2025poster

Vision-Language Retrieval-Augmented Generation (VLRAG) systems have been widely applied to Large Vision-Language Models (LVLMs) to enhance their generation ability. However, the reliance on external multimodal knowledge databases renders VLRAG systems vulnerable to malicious poisoning attacks. In th…

Cited by 0SourcePDFScholar
2025

Synthetic Series-Symbol Data Generation for Time Series Foundation Models

NeurIPS 2025poster

Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as training data scarcity and imbalance continue to hinder their development. Inspired by complex dynamic system theories, we design a series-symbol data generation mechanism, enabling the…

Cited by 0SourcecodeScholar
2025

V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception

NeurIPS 2025spotlight

Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In…

Cited by 0SourceScholar
2025

VTON-HandFit: Virtual Try-on for Arbitrary Hand Pose Guided by Hand Priors Embedding

CVPR 2025poster

Although diffusion-based image virtual try-on has made considerable progress, emerging approaches still struggle to effectively address the issue of hand occlusion (i.e., clothing regions occluded by the hand part), leading to a notable degradation of the try-on performance. To tackle this issue wid…

2025

pFedRAG: A Personalized Federated Retrieval-Augmented Generation System with Depth-Adaptive Tiered Embedding Tuning

EMNLP 2025

Large Language Models (LLMs) can undergo hallucinations in specialized domains, and standard Retrieval-Augmented Generation (RAG) often falters due to general-purpose embeddings ill-suited for domain-specific terminology. Though domain-specific fine-tuning enhances retrieval, centralizing data intro

Cited by 0SourcePDFScholar
2024

Automated Loss function Search for Class-imbalanced Node Classification

ICML 2024poster

Class-imbalanced node classification tasks are prevalent in real-world scenarios. Due to the uneven distribution of nodes across different classes, learning high-quality node representations remains a challenging endeavor. The engineering of loss functions has shown promising potential in addressing…

Cited by 1SourcePDFScholar
2024

Biomimetic Crawling Robot Based on Dielectric Elastomer: Design, Modeling and Experiment

RA-L 2024

In this study, a bio-inspired crawling robot with multi-surface locomotion capability is developed. The robot is driven by Dielectric Elastomer Minimum Energy Structures (DEMES) and utilizes a three-dimensional scissor mechanism and electrostatic adhesion technology to achieve multi-surface crawling

Cited by 3SourceScholar
2024

Pretrained Optimization Model for Zero-Shot Black Box Optimization

NeurIPS 2024poster

Zero-shot optimization involves optimizing a target task that was not seen during training, aiming to provide the optimal solution without or with minimal adjustments to the optimizer. It is crucial to ensure reliable and robust performance in various applications. Current optimizers often struggle…

2024

Signed Graph Neural Ordinary Differential Equation for Modeling Continuous-Time Dynamics

AAAI 2024technical

Modeling continuous-time dynamics constitutes a foundational challenge, and uncovering inter-component correlations within complex systems holds promise for enhancing the efficacy of dynamic modeling. The prevailing approach of integrating graph neural networks with ordinary differential equations h…

2024

Tuning-Free Image Customization with Image and Text Guidance

ECCV 2024poster

"Despite significant advancements in image customization with diffusion models, current methods still have several limitations: 1) unintended changes in non-target areas when regenerating the entire image; 2) guidance solely by a reference image or text descriptions; and 3) time-consuming fine-tunin…

2024

UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation

IJCAI 2024poster

3D open-vocabulary scene understanding aims to recognize arbitrary novel categories beyond the base label space. However, existing works not only fail to fully utilize all the available modal information in the 3D domain but also lack sufficient granularity in representing the features of each modal…

2024

Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt

AAAI 2024technical

Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due…

2022

Adaptable Action-Aware Vital Models for Personalized Intelligent Patient Monitoring

ICRA 2022poster

Vital signs such as heart rate, oxygen saturation, and blood pressure are crucial information for healthcare workers to identify clinical deterioration of ward patients. Currently, medical devices monitor these vital signs and trigger alarms when the vital signs are not in the normal ranges based on…

Cited by 5SourceScholar
2022

Class-Aware Contrastive Semi-Supervised Learning

CVPR 2022poster

Pseudo-label-based semi-supervised learning (SSL) has achieved great success on raw data utilization. However, its training procedure suffers from confirmation bias due to the noise contained in self-generated artificial labels. Moreover, the model's judgment becomes noisier in real-world applicatio…

Cited by 139PDFcodeScholar
2022

SoftPatch: Unsupervised Anomaly Detection with Noisy Data

NeurIPS 2022accept

Although mainstream unsupervised anomaly detection (AD) algorithms perform well in academic datasets, their performance is limited in practical application due to the ideal experimental setting of clean training data. Training with noisy data is an inevitable problem in real-world anomaly detection…

2019

EV-Gait: Event-Based Robust Gait Recognition Using Dynamic Vision Sensors

CVPR 2019poster

In this paper, we introduce a new type of sensing modality, the Dynamic Vision Sensors (Event Cameras), for the task of gait recognition. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and muc…

Cited by 186PDFScholar
2015

Multi-source direction-of-arrival estimation in a reverberant environment using single acoustic vector sensor

ICASSP 2015accepted

We address the problem of estimating direction-of-arrivals (DOAs) for multiple sound sources using a single acoustic vector sensor (AVS) in an enclosed room environment. It is well-known that multi-source DOA estimation in an enclosed environment is challenging due to room reverberation, environment…

Cited by 0SourceScholar
2015

Single-channel speech enhancement in a transient noise environment by exploiting speech harmonicity

ICASSP 2015accepted

This paper focuses on the problem of single-channel noise reduction in a transient noise environment for speech enhancement application. A typical speech enhancement algorithm requires an estimate of the noise statistics. However, the problem of noise estimation is challenging when the statistics of…

Cited by 0SourceScholar