← Search

Yixiao Zhang

11 accepted papers

2026

From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition

CVPR 2026

Images can be viewed as layered compositions, foreground objects over background, with potential occlusions. This layered representation enables independent editing of elements, offering greater flexibility for content creation. Despite the progress in large generative models, decomposing a single i

Cited by 0SourceScholar
2024

Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls

IJCAI 2024poster

Controllable music generation plays a vital role in human-AI music co-creation. While Large Language Models (LLMs) have shown promise in generating high-quality music, their focus on autoregressive generation limits their utility in music editing tasks. To bridge this gap, To address this gap, we pr…

2024

Multi-robot Human-in-the-loop Control under Spatiotemporal Specifications

ICRA 2024poster

In this work, we present a coordination strategy tailored for scenarios involving multiple agents and tasks. We devise a range of tasks using signal temporal logic (STL), each earmarked for specific agents. These tasks are then imposed through control barrier function (CBF) constraints to ensure com…

Cited by 1SourceScholar
2024

MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models

IJCAI 2024poster

Recent advances in text-to-music generation models have opened new avenues in musical creativity. However, the task of editing these generated music remains a significant challenge. This paper introduces a novel approach to edit music generated by such models, enabling the modification of specific a…

2023

CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection

ICCV 2023poster

An increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited investigation of diverse types of tumors, the resulting models are often limited…

Cited by 230PDFcodeScholar
2023

SQUID: Deep Feature In-Painting for Unsupervised Anomaly Detection

CVPR 2023poster

Radiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. To exploit this structured information, we propose the use of Space-aware Memory Queues for In-painting and Detecting anomalies…

2022

Music Phrase Inpainting Using Long-Term Representation and Contrastive Loss

ICASSP 2022accepted

Deep generative modeling has already become the leading technique for music automation. However, long-term generation remains a challenging task as most methods fall short in preserving a natural structure and the overall musicality when the generation scope exceeds several beats. In this study, we…

Cited by 0SourceScholar
2021

Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real Images

ICCV 2021poster

While neural symbolic methods demonstrate impressive performance in visual question answering on synthetic images, their performance suffers on real images. We identify that the long-tail distribution of visual concepts and unequal importance of reasoning steps in real data are the two key obstacles…

Cited by 18PDFcodeScholar
2020

C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentation

CVPR 2020poster

3D convolution neural networks (CNN) have been proved very successful in parsing organs or tumours in 3D medical images, but it remains sophisticated and time-consuming to choose or design proper 3D networks given different task contexts. Recently, Neural Architecture Search (NAS) is proposed to sol…

Cited by 182PDFScholar
2017

Model-free control for soft manipulators based on reinforcement learning

IROS 2017poster

Most control methods of soft manipulators are developed based on physical models derived from mathematical analysis or learning methods. However, due to internal nonlinearity and external uncertain disturbances, it is difficult to build an accurate model, further, these methods lack robustness and p…

Cited by 79SourceScholar