← Search

Ziyu Wang

60 accepted papers

2026

Deep Incomplete Multi-View Clustering via Hierarchical Imputation and Alignment

AAAI 2026technical

Incomplete multi-view clustering (IMVC) aims to discover shared cluster structures from multi-view data with partial observations. The core challenges lie in accurately imputing missing views without introducing bias, while maintaining semantic consistency across views and compactness within cluster

Cited by 0SourcePDFScholar
2026

EVALUATING HIGH-RESOLUTION PIANO SUSTAIN PEDAL DEPTH ESTIMATION WITH MUSICALLY INFORMED METRICS

ICASSP 2026poster

Evaluation for continuous piano pedal depth estimation tasks remains incomplete when relying only on conventional frame-level metrics, which overlook musically important features such as direction-change boundaries and pedal curve contours. To provide more interpretable and musically meaningful insi…

Cited by 0SourcePDFScholar
2026

Fast Exploration Planning with Learning-Based Motion Time Prediction for Aerial Robots

ICRA 2026poster

Unmanned aerial vehicles (UAVs) have been widely employed to achieve autonomous exploration of 3D unknown environments. However, most existing algorithms suffer from low exploration efficiency caused by inaccurate motion time cost evaluation, which typically leads to the motion inconsistency during …

Cited by 0Scholar
2026

RIPNEON: Memory-Lite and Computation-Efficient Occupancy Mapping Via Block Read-Write and Key Grids Expansion

ICRA 2026poster

Mobile robot motion planning heavily relies on grid-based occupancy maps, while existing works require high memory usage and expensive updating overhead. In this work, we propose a memory-lite grid-block data structure and an efficient map updating algorithm for LiDAR-based online exploration-orient…

Cited by 0codeScholar
2026

WHEN NOISE LOWERS THE LOSS: RETHINKING LIKELIHOOD-BASED EVALUATION IN MUSIC LARGE LANGUAGE MODELS

ICASSP 2026poster

The rise of music large language models (LLMs) demands robust methods of evaluating output quality, especially in distinguishing high-quality compositions from "garbage music". Curiously, we observe that the standard cross-entropy loss -- a core training metric -- often decrease when models encounte…

Cited by 0SourcePDFScholar
2025

ArticuBot: Learning Universal Articulated Object Manipulation Policy via Large Scale Simulation

RSS 2025poster

This paper presents ArticuBot, in which a single learned policy enables a robotics system to open diverse categories of unseen articulated objects in the real world. This task has long been challenging for robotics due to the large variations in the geometry, size, and articulation types of such ob…

Cited by 0PDFScholar
2025

Efficient Fine-Grained Guidance for Diffusion Model Based Symbolic Music Generation

ICML 2025poster

Developing generative models to create or conditionally create symbolic music presents unique challenges due to the combination of limited data availability and the need for high precision in note pitch. To address these challenges, we introduce an efficient Fine-Grained Guidance (FGG) approach with…

Cited by 0SourcePDFScholar
2025

Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts

ICML 2025poster

Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transfor…

Cited by 0SourcePDFScholar
2025

On Probabilistic Truncation in Privacy-preserving Machine Learning

AAAI 2025technical

Probabilistic truncation has been widely used in a broad range of privacy-preserving machine learning (PPML) platforms, such as EdaBits (Crypto 20), ABY 2.0 (Usenix 21), Crypten (NIPS 21), Piranha-Falcon (Usenix 22), and Bicoptor (S&P 23), etc. In this work, we examine the problems of common proba…

2025

On Subjective Uncertainty Quantification and Calibration in Natural Language Generation

AISTATS 2025poster

Applications of large language models often involve the generation of free-form responses, in which case uncertainty quantification becomes challenging. This is due to the need to identify task-specific uncertainties (e.g., about the semantics) which appears difficult to define in general cases. Thi…

Cited by 0SourcecodeScholar
2025

Real-World Offline Reinforcement Learning from Vision Language Model Feedback

IROS 2025

Offline reinforcement learning can enable policy learning from pre-collected, sub-optimal datasets without online interactions. This makes it ideal for real-world robots and safety-critical scenarios, where collecting online data or expert demonstrations is slow, costly, and risky. However, most exi

Cited by 15SourceScholar
2025

Synthetic Video Enhances Physical Fidelity in Video Synthesis

ICCV 2025poster

We investigate how to enhance the physical fidelity of video generation models by leveraging synthetic videos generated via standard computer graphics techniques. These rendered videos respect real-world physics -- such as maintaining 3D consistency -- thereby serving as a valuable resource that can…

2025

TROI: Cross-Subject Pretraining with Sparse Voxel Selection for Enhanced fMRI Visual Decoding

ICASSP 2025accepted

fMRI (functional Magnetic Resonance Imaging) visual decoding involves decoding the original image from brain signals elicited by visual stimuli. This often relies on manually labeled ROIs (Regions of Interest) to select brain voxels. However, these ROIs can contain redundant information and noise, r…

Cited by 0SourceScholar
2025

Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization

NeurIPS 2025poster

We present a unified framework for automatic multitrack music arrangement that enables a single pre-trained symbolic music model to handle diverse arrangement scenarios, including reinterpretation, simplification, and additive generation. At its core is a segment-level reconstruction objective opera…

Cited by 0SourcecodeScholar
2025

Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints

ICLR 2025poster

We contribute an unsupervised method that effectively learns disentangled content and style representations from sequences of observations. Unlike most disentanglement algorithms that rely on domain-specific labels or knowledge, our method is based on the insight of domain-general statistical differ…

Cited by 0SourcePDFScholar
2024

Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement

CVPR 2024poster

Recently dataset distillation has paved the way towards efficient machine learning especially for image datasets. However the distillation for videos characterized by an exclusive temporal dimension remains an underexplored domain. In this work we provide the first systematic study of video distilla…

2024

DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation Learning

NeurIPS 2024spotlight

This paper introduces DiffTORI, which utilizes $\textbf{Diff}$erentiable $\textbf{T}$rajectory $\textbf{O}$ptimization as the policy representation to generate actions for deep $\textbf{R}$einforcement and $\textbf{I}$mitation learning. Trajectory optimization is a powerful and widely used algorithm…

2024

Distill Gold from Massive Ores: Bi-level Data Pruning towards Efficient Dataset Distillation

ECCV 2024poster

"Data-efficient learning has garnered significant attention, especially given the current trend of large multi-modal models. Recently, dataset distillation has become an effective approach by synthesizing data samples that are essential for network training. However, it remains to be explored which…

2024

Fast and Communication-Efficient Multi-UAV Exploration Via Voronoi Partition on Dynamic Topological Graph

IROS 2024

Efficient data transmission and reasonable task allocation are important to improve multi-robot exploration efficiency. However, most communication data types typically contain redundant information and thus require massive communication volume. Moreover, exploration-oriented task allocation is far

Cited by 21SourcecodeScholar
2024

Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception

NeurIPS 2024poster

Biological motion perception (BMP) refers to humans' ability to perceive and recognize the actions of living beings solely from their motion patterns, sometimes as minimal as those depicted on point-light displays. While humans excel at these tasks \textit{without any prior training}, current AI mod…

2024

Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective

ICML 2024poster

In-context learning (ICL) has emerged as a particularly remarkable characteristic of Large Language Models (LLM): given a pretrained LLM and an observed dataset, LLMs can make predictions for new data points from the same distribution without fine-tuning. Numerous works have postulated ICL as approx…

2024

Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling

NeurIPS 2024poster

In the realm of music AI, arranging rich and structured multi-track accompaniments from a simple lead sheet presents significant challenges. Such challenges include maintaining track cohesion, ensuring long-term coherence, and optimizing computational efficiency. In this paper, we introduce a novel…

2024

Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models

ICLR 2024spotlight

Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured **whole-song** generation. In this paper, we make the first attempt to model a full music piece under the realization of *compositional hierar…

Cited by 11SourcePDFScholar
2023

A constrained Bayesian approach to out-of-distribution prediction

UAI 2023poster

Consider the problem of out-of-distribution prediction given data from multiple environments. While a sufficiently diverse collection of training environments will facilitate the identification of an invariant predictor, with an optimal generalization performance, many applications only provide us w…

Cited by 0SourcePDFScholar
2023

Controllable Music Inpainting with Mixed-Level and Disentangled Representation

ICASSP 2023accepted

Music inpainting, which is to complete the missing part of a piece given some context, is an important task of automated music generation. In this study, we contribute a controllable inpainting model by combining the high expressivity of mixed-level, disentangled music representations and the strong…

Cited by 0SourceScholar
2023

HumanGen: Generating Human Radiance Fields With Explicit Priors

CVPR 2023poster

Recent years have witnessed the tremendous progress of 3D GANs for generating view-consistent radiance fields with photo-realism. Yet, high-quality generation of human radiance fields remains challenging, partially due to the limited human-related priors adopted in existing methods. We present Human…

Cited by 38SourcePDFScholar
2023

Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos

CVPR 2023poster

The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has inspired numerous attempts on dynamic scenes. Current techniques that utilize neural rendering for facilitating free-view videos (FVVs) are restricted to either offline rendering or are capable…

Cited by 66SourcePDFScholar
2023

Object-centric Learning with Cyclic Walks between Parts and Whole

NeurIPS 2023poster

Learning object-centric representations from complex natural environments enables both humans and machines with reasoning abilities from low-level perceptual features. To capture compositional entities of the scene, we proposed cyclic walks between perceptual features extracted from vision transform…

2022

Audio-To-Symbolic Arrangement Via Cross-Modal Music Representation Learning

ICASSP 2022accepted

Could we automatically derive the score of a piano accompaniment based on the audio of a pop song? This is the audio-to-symbolic arrangement problem we tackle in this paper. A good arrangement model should not only consider the audio content but also have prior knowledge of piano composition (so tha…

Cited by 0SourceScholar
2022

DynaMixer: A Vision MLP Architecture with Dynamic Mixing

ICML 2022spotlight

Recently, MLP-like vision models have achieved promising performances on mainstream visual recognition tasks. In contrast with vision transformers and CNNs, the success of MLP-like models shows that simple information fusion operations among tokens and channels can yield a good representation power…

2022

Kubric: A Scalable Dataset Generator

CVPR 2022poster

Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises a…

Cited by 249PDFcodeScholar
2022

Liftoff of A Motor-Driven Flapping Wing Rotorcraft with Mechanically Decoupled Wings

ICRA 2022poster

Flapping Wing Rotorcraft (FWR) combines flapping and rotating wing motion in one element. Such a hybrid design integrates the high-efficiency characteristics of the rotating wing and the high-lift feature of the flapping wing under low Reynolds number, providing a broader range of simultaneous lift…

Cited by 7SourceScholar
2022

SWEM: Towards Real-Time Video Object Segmentation With Sequential Weighted Expectation-Maximization

CVPR 2022poster

Matching-based methods, especially those based on space-time memory, are significantly ahead of other solutions in semi-supervised video object segmentation (VOS). However, continuously growing and redundant template features lead to an inefficient inference. To alleviate this, we propose a novel Se…

Cited by 58PDFcodeScholar
2021

Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization

ICLR 2021poster

Standard dynamics models for continuous control make use of feedforward computation to predict the conditional distribution of next state and reward given current state and action using a multivariate Gaussian with a diagonal covariance structure. This modeling choice assumes that different dimensio…

Cited by 53SourcePDFScholar
2021

Benchmarks for Deep Off-Policy Evaluation

ICLR 2021poster

Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability to learn offline is particularly important in many real-world domains, such as in healthcare, recommender systems, or ro…

2021

Fork or Fail: Cycle-Consistent Training with Many-to-One Mappings

AISTATS 2021poster

Cycle-consistent training is widely used for jointly learning a forward and inverse mapping between two domains of interest without the cumbersome requirement of collecting matched pairs within each domain. In this regard, the implicit assumption is that there exists (at least approximately) a groun…

2021

Scalable Quasi-Bayesian Inference for Instrumental Variable Regression

NeurIPS 2021poster

Recent years have witnessed an upsurge of interest in employing flexible machine learning models for instrumental variable (IV) regression, but the development of uncertainty quantification methodology is still lacking. In this work we present a scalable quasi-Bayesian procedure for IV regression,…

Cited by 11SourcePDFScholar
2020

A Wasserstein Minimum Velocity Approach to Learning Unnormalized Models

AISTATS 2020poster

Score matching provides an effective approach to learning flexible unnormalized models, but its scalability is limited by the need to evaluate a second-order derivative. In this paper, we present a scalable approximation to a general family of learning objectives including score matching, by observi…

2020

Critic Regularized Regression

NeurIPS 2020poster

Offline reinforcement learning (RL), also known as batch RL, offers the prospect of policy optimization from large pre-recorded datasets without online environment interaction. It addresses challenges with regard to the cost of data collection and safety, both of which are particularly pertinent to…

Cited by 378SourcePDFScholar
2020

Further Analysis of Outlier Detection with Deep Generative Models

NeurIPS 2020poster

The recent, counter-intuitive discovery that deep generative models (DGMs) can frequently assign a higher likelihood to outliers has implications for both outlier detection applications as well as our overall understanding of generative modeling. In this work, we present a possible explanation for t…

2020

Making Efficient Use of Demonstrations to Solve Hard Exploration Problems

ICLR 2020poster

This paper introduces R2D3, an agent that makes efficient use of demonstrations to solve hard exploration problems in partially observable environments with highly variable initial conditions. We also introduce a suite of eight tasks that combine these three properties, and show that R2D3 can solve…

Cited by 107SourceScholar
2020

RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning

NeurIPS 2020poster

Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to learn policies from offline datasets, thus overcoming concerns associated with online data collection in the real-world, in…

2020

Scaling data-driven robotics with reward sketching and batch reinforcement learning

RSS 2020poster

By harnessing a growing dataset of robot experience, we learn control policies for a diverse and increasing set of related manipulation tasks. To make this possible, we introduce reward sketching: an effective way of eliciting human preferences to learn the reward function for a new task. This rewar…

2020

Task-Relevant Adversarial Imitation Learning

CoRL 2020

We show that a critical vulnerability in adversarial imitation is the tendency of discriminator networks to learn spurious associations between visual features and expert labels. When the discriminator focuses on task-irrelevant features, it does not provide an informative reward signal, leading to

Cited by 0SourcePDFScholar
2019

Function Space Particle Optimization for Bayesian Neural Networks

ICLR 2019poster

While Bayesian neural networks (BNNs) have drawn increasing attention, their posterior inference remains challenging, due to the high-dimensional and over-parameterized nature. To address this issue, several highly flexible and scalable variational inference procedures based on the idea of particle…

2018

Learning an Embedding Space for Transferable Robot Skills

ICLR 2018poster

We present a method for reinforcement learning of closely related skills that are parameterized via a skill embedding space. We learn such skills by taking advantage of latent variables and exploiting a connection between reinforcement learning and variational inference. The main contribution of our…

Cited by 365SourcePDFScholar
2018

Playing hard exploration games by watching YouTube

NeurIPS 2018spotlight

Deep reinforcement learning methods traditionally struggle with tasks where environment rewards are particularly sparse. One successful method of guiding exploration in these domains is to imitate trajectories provided by a human demonstrator. However, these demonstrations are typically collected un…

Cited by 329SourcePDFScholar
2018

Reinforcement and Imitation Learning for Diverse Visuomotor Skills

RSS 2018poster

We propose a general model-free deep reinforcement learning method and apply it to robotic manipulation tasks. Our approach leverages a small amount of demonstration data to assist a reinforcement learning agent. We train end-to-end visuomotor policies to learn a direct mapping from RGB camera input…

Cited by 398SourcePDFScholar
2017

Parallel Multiscale Autoregressive Density Estimation

ICML 2017poster

PixelCNN achieves state-of-the-art results in density estimation for natural images. Although training is fast, inference is costly, requiring one network evaluation per pixel; O(N) for N pixels. This can be sped up by caching activations, but still involves generating each pixel sequentially. In th…

Cited by 261SourcePDFScholar
2017

Robust Imitation of Diverse Behaviors

NeurIPS 2017poster

Deep generative models have recently shown great promise in imitation learning for motor control. Given enough data, even supervised approaches can do one-shot imitation learning; however, they are vulnerable to cascading failures when the agent trajectory diverges from the demonstrations. Compared…

Cited by 266SourcePDFScholar
2017

Sample Efficient Actor-Critic with Experience Replay

ICLR 2017poster

This paper presents an actor-critic deep reinforcement learning agent with experience replay that is stable, sample efficient, and performs remarkably well on challenging environments, including the discrete 57-game Atari domain and several continuous control problems. To achieve this, the paper int…

Cited by 1079SourceScholar
2017

The Intentional Unintentional Agent: Learning to Solve Many Continuous Control Tasks Simultaneously

CoRL 2017

This paper introduces the Intentional Unintentional (IU) agent. This agent endows the deep deterministic policy gradients (DDPG) agent for continuous control with the ability to solve several tasks simultaneously. Learning to solve many tasks simultaneously has been a long-standing, core goal of art

Cited by 0SourcePDFScholar
2016

Dueling Network Architectures for Deep Reinforcement Learning

ICML 2016poster

In recent years there have been many successes of using deep representations in reinforcement learning. Still, many of these applications use conventional architectures, such as convolutional networks, LSTMs, or auto-encoders. In this paper, we present a new neural network architecture for model-fre…

Cited by 5833SourcePDFScholar