← Search

Xiaoyu Zhang

45 accepted papers

2026

$\texttt{MetaDistill}$: Unlocking the Performance Ceiling for Pretrained Optimizers

ICML 2026poster

Meta Black-Box Optimization (MetaBBO) has emerged as a promising paradigm by employing meta learning to automatically optimize the configurations of low-level black-box optimizers. Despite its potential, the generalization of MetaBBO remains significantly constrained when facing unseen, complex obje…

Cited by 0SourceScholar
2026

A General Highly Accurate Online Planning Method Integrating Large Language Models into Nested Rollout Policy Adaptation for Dialogue Tasks

AAAI 2026technical

In goal-oriented dialogue tasks, the main challenge is to steer the interaction towards a given goal within a limited number of turns. Existing approaches either rely on elaborate prompt engineering, whose effectiveness is heavily dependent on human experience, or integrate policy networks and pre-t

Cited by 0SourcePDFScholar
2026

Adaptive Legged Locomotion Via Online Learning for Model Predictive Control

ICRA 2026poster

We provide an algorithm for adaptive legged locomotion via online learning and model predictive control. The algorithm is composed of two interacting modules: model predictive control (MPC) and online learning of residual dynamics. The residual dynamics can represent modeling errors and external dis…

2026

AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors

AAAI 2026technical

As Convolutional Neural Networks (CNNs) continue to gain traction in deep learning, Winograd convolution has emerged as a key algorithm to enhance computational efficiency. Although ARM-based CPUs are increasingly prevalent in mobile devices, embedded systems and HPC servers, existing 2D Winograd co

Cited by 0SourcePDFScholar
2026

Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?

ICLR 2026poster

Stage lighting is a vital component in live music performances, shaping an engaging experience for both musicians and audiences. In recent years, Automatic Stage Lighting Control (ASLC) has attracted growing interest due to the high costs of hiring or training professional lighting engineers. Howeve…

Cited by 0SourcecodeScholar
2026

City-Scale Lane-Level Mapping From Crowdsourced Trajectories and Satellite Imagery

RA-L 2026

Lane-level maps are increasingly preferred over Standard-Definition (SD) and High-Definition (HD) maps, offering a better trade-off among detail richness, coverage breadth, and data freshness. However, constructing city-scale lane-level maps remains time-consuming and labor-intensive. To address the

Cited by 0SourceScholar
2026

CodeChemist: Test-Time Scaling for Low-Resource Code Generation via Functional Knowledge Transfer

ICML 2026poster

Code Large Language Models (CodeLLMs) have been widely adopted for Natural Language to Programming Language code generation, powering applications with large user bases. Their performance, however, varies sharply across programming languages (PLs) and is particularly suboptimal for low-resource PLs …

Cited by 0SourceScholar
2026

SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Cameras

CVPR 2026

High-quality 4D reconstruction enables photorealistic and immersive rendering of the dynamic real world. However, unlike static scenes that can be fully captured with a single camera, high-quality dynamic scenes typically require dense arrays of tens or even hundreds of synchronized cameras. Depende

Cited by 0SourcecodeScholar
2026

Stochastic Universal Adversarial Perturbations with Fixed Optimization Constraint and Ensured High-probability Transferability

AAAI 2026technical

Adversarial perturbations (APs) have become a great concern in image classification tasks. The most challenging branch, universal adversarial perturbations (UAPs), are exploited to fool most of the unseen samples. Such one-to-all perturbations have the merit of transferability, which has strong prac

Cited by 0SourcePDFScholar
2025

B2Opt: Learning to Optimize Black-box Optimization with Little Budget

AAAI 2025technical

The core challenge of high-dimensional and expensive black-box optimization (BBO) is how to obtain better performance faster with little function evaluation cost. The essence of the problem is how to design an efficient optimization strategy tailored to the target task. This paper designs a powerful…

Cited by 10SourcePDFScholar
2025

EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models

AAAI 2025technical

The rapid development of diffusion models has significantly advanced AI-generated content (AIGC), particularly in Text-to-Image (T2I) and Text-to-Video (T2V) generation. Text-based video editing, leveraging these generative capabilities, has emerged as a promising field, enabling precise modificatio…

2025

Enhancing Zero-Shot Black-Box Optimization via Pretrained Models with Efficient Population Modeling, Interaction, and Stable Gradient Approximation

NeurIPS 2025poster

Zero-shot optimization aims to achieve both generalization and performance gains on solving previously unseen black-box optimization problems over SOTA methods without task-specific tuning. Pre-trained optimization models (POMs) address this challenge by learning a general mapping from task features…

Cited by 0SourceScholar
2025

EvDetMAV: Generalized MAV Detection From Moving Event Cameras

RA-L 2025

Existing micro aerial vehicle (MAV) detection methods mainly rely on the target's appearance features in RGB images, whose diversity makes it difficult to achieve generalized MAV detection. We notice that different types of MAVs share the same distinctive features in event streams due to their high-

Cited by 4SourcecodeScholar
2025

Liberated-GS: 3D Gaussian Splatting Independent from SfM Point Clouds

ICCV 2025poster

3D Gaussian Splatting (3DGS) has demonstrated impressive performance in novel view synthesis and real-time rendering. However, it heavily relies on high-quality initial sparse points from Structure-from-Motion (SfM) which often struggles in textureless regions, degrading the geometry and visual qual…

Cited by 0SourcePDFScholar
2025

LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene

CVPR 2025poster

Humans perceive and comprehend their surroundings through information spanning multiple frequencies. In immersive scenes, people naturally scan their environment to grasp its overall structure while examining fine details of objects that capture their attention. However, current NeRF frameworks prim…

Cited by 1SourcePDFScholar
2025

PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs

ICCV 2025poster

Panoramic Image Generation (PIG) aims to create coherent images of arbitrary lengths. Most existing methods fall in the joint diffusion paradigm, but their complex and heuristic crop connection designs often limit their ability to achieve multilevel coherence. By deconstructing this challenge into i…

2025

PoisonedEye: Knowledge Poisoning Attack on Retrieval-Augmented Generation based Large Vision-Language Models

ICML 2025poster

Vision-Language Retrieval-Augmented Generation (VLRAG) systems have been widely applied to Large Vision-Language Models (LVLMs) to enhance their generation ability. However, the reliance on external multimodal knowledge databases renders VLRAG systems vulnerable to malicious poisoning attacks. In th…

Cited by 0SourcePDFScholar
2025

STAFF: Speculative Coreset Selection for Task-Specific Fine-tuning

ICLR 2025poster

Task-specific fine-tuning is essential for the deployment of large language models (LLMs), but it requires significant computational resources and time. Existing solutions have proposed coreset selection methods to improve data efficiency and reduce model training overhead, but they still have limit…

Cited by 2SourcePDFScholar
2025

Synthetic Series-Symbol Data Generation for Time Series Foundation Models

NeurIPS 2025poster

Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as training data scarcity and imbalance continue to hinder their development. Inspired by complex dynamic system theories, we design a series-symbol data generation mechanism, enabling the…

Cited by 0SourcecodeScholar
2025

The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation

ACL 2025long

Large Language Models (LLMs) have emerged as the new recommendation engines, surpassing traditional methods in both capability and scope, particularly in code generation. In this paper, we reveal a novel **provider bias** in LLMs: without explicit directives, these models show systematic preferences…

2025

Tile-wise vs. Image-wise: Random-Tile Loss and Training Paradigm for Gaussian Splatting

ICCV 2025poster

3D Gaussian Splatting (3DGS) has drawn significant attention for its advantages in rendering speed and quality. Most existing methods still rely on the image-wise loss and training paradigm because of its intuitive nature in the Splatting algorithm. However, image-wise loss lacks multi-view constrai…

Cited by 0SourcePDFScholar
2025

Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial Robustness

AAAI 2025technical

For an extensive period, Vision Transformers (ViTs) have been deemed unsuitable for attaining robust performance on small-scale datasets, with WideResNet models maintaining dominance in this domain. While WideResNet models have persistently set the state-of-the-art (SOTA) benchmarks for robust accur…

Cited by 0SourcePDFScholar
2024

Automated Loss function Search for Class-imbalanced Node Classification

ICML 2024poster

Class-imbalanced node classification tasks are prevalent in real-world scenarios. Due to the uneven distribution of nodes across different classes, learning high-quality node representations remains a challenging endeavor. The engineering of loss functions has shown promising potential in addressing…

Cited by 1SourcePDFScholar
2024

Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning

RSS 2024poster

A common failure mode for policies trained with imitation is compounding execution errors at test time. When the learned policy encounters states that are not present in the expert demonstrations, the policy fails, leading to degenerate behavior. The Dataset Aggregation, or DAgger approach to this p…

Cited by 15SourcePDFScholar
2024

Enhancing Vectorized Map Perception with Historical Rasterized Maps

ECCV 2024poster

"In autonomous driving, there is growing interest in end-to-end online vectorized map perception in bird’s-eye-view (BEV) space, with an expectation that it could replace traditional high-cost offline high-definition (HD) maps. However, the accuracy and robustness of these methods can be easily comp…

2024

How to Engage your Readers? Generating Guiding Questions to Promote Active Reading

ACL 2024long

Using questions in written text is an effective strategy to enhance readability. However, what makes an active reading question good, what the linguistic role of these questions is, and what is their impact on human reading remains understudied. We introduce GuidingQ, a dataset of 10K in-text questi…

2024

Leveraging Enhanced Queries of Point Sets for Vectorized Map Construction

ECCV 2024poster

"In autonomous driving, the high-definition (HD) map plays a crucial role in localization and planning. Recently, several methods have facilitated end-to-end online map construction in DETR-like frameworks. However, little attention has been paid to the potential capabilities of exploring the query…

2024

Pretrained Optimization Model for Zero-Shot Black Box Optimization

NeurIPS 2024poster

Zero-shot optimization involves optimizing a target task that was not seen during training, aiming to provide the optimal solution without or with minimal adjustments to the optimizer. It is crucial to ensure reliable and robust performance in various applications. Current optimizers often struggle…

2023

ERM-KTP: Knowledge-Level Machine Unlearning via Knowledge Transfer

CVPR 2023poster

Machine unlearning can fortify the privacy and security of machine learning applications. Unfortunately, the exact unlearning approaches are inefficient, and the approximate unlearning approaches are unsuitable for complicated CNNs. Moreover, the approximate approaches have serious security flaws be…

2023

Explaining Adversarial Robustness of Neural Networks from Clustering Effect Perspective

ICCV 2023poster

Adversarial training (AT) is the most commonly used mechanism to improve the robustness of deep neural networks. Recently, a novel adversarial attack against intermediate layers exploits the extra fragility of adversarially trained networks to output incorrect predictions. The result implies the ins…

Cited by 1PDFcodeScholar
2023

MUter: Machine Unlearning on Adversarially Trained Models

ICCV 2023poster

Machine unlearning is an emerging task of removing the influence of selected training datapoints from a trained model upon data deletion requests, which echoes the widely enforced data regulations mandating the Right to be Forgotten. Many unlearning methods have been proposed recently, achieving sig…

Cited by 27PDFScholar
2023

rPPG-Toolbox: Deep Remote PPG Toolbox

NeurIPS 2023poster

Camera-based physiological measurement is a fast growing field of computer vision. Remote photoplethysmography (rPPG) utilizes imaging devices (e.g., cameras) to measure the peripheral blood volume pulse (BVP) via photoplethysmography, and enables cardiac measurement via webcams and smartphones. How…

2022

HIRL: Hybrid Image Restoration Based on Hierarchical Deep Reinforcement Learning via Two-Step Analysis

ICASSP 2022accepted

The restoration of hybrid distorted images in real-world scenarios is still a difficult problem due to the fact that the degrading types and degrees are always unknown. Previous studies typically utilize multiple recovery tools to restore images. However, each tool adopted inevitably introduces addi…

Cited by 0SourceScholar
2022

JE2NET: Joint Exploitation and Exploration in Reinforcement Learning Based Image Restoration

ICASSP 2022accepted

Previous reinforcement learning (RL) based image restoration studies typically train RL agents to search for recovery tools from a constructed toolset and iteratively recover images. However, we argue that these agents rely on pre-trained RL models with fixed-length paths for restoration, which perf…

Cited by 0SourceScholar
2022

Robust Localization of Occluded Targets in Aerial Manipulation Via Range-Only Mapping

RA-L 2022

This letter studies the problem of target localization in aerial manipulation tasks. When an aerial robot flies close to a target to manipulate, the target would be occluded by the onboard robotic manipulator occasionally or for a long period of time. It is, however, necessary to continuously locali

Cited by 6SourceScholar
2022

SO-SLAM: Semantic Object SLAM With Scale Proportional and Symmetrical Texture Constraints

RA-L 2022

Object SLAM introduces the concept of objects into Simultaneous Localization and Mapping (SLAM) and helps understand indoor scenes for mobile robots and object-level interactive applications. The state-of-art object SLAM systems face challenges such as partial observations, occlusions, unobservable

Cited by 81SourcecodeScholar
2021

Inferring Camouflaged Objects by Texture-Aware Interactive Guidance Network

AAAI 2021technical

Camouflaged objects, similar to the background, show indefinable boundaries and deceptive textures, which increases the difficulty of detection task and makes the model rely on features with more information. Herein, we design a texture label to facilitate our network for accurate camouflaged object…

Cited by 126SourcePDFScholar
2021

Task-Space Decomposed Motion Planning Framework for Multi-Robot Loco-Manipulation

ICRA 2021poster

This paper introduces a novel task-space decomposed motion planning framework for multi-robot simultaneous locomotion and manipulation. When several manipulators hold an object, closed-chain kinematic constraints are formed, and it will make the motion planning problems challenging by inducing lower…

Cited by 12SourceScholar