← Search

Zhiyuan Zhao

27 accepted papers

2026

A Visual Reinforcement Learning-Based Separate Primitive Policy for Peg-in-Hole Tasks

RA-L 2026

For peg-in-hole tasks, humans rely on binocular visual perception to locate the peg above the hole surface and then proceed with insertion. This paper draws insights from this behavior to enable agents to learn efficient assembly strategies through visual reinforcement learning. Hence, we propose a

Cited by 0SourceScholar
2026

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

AAAI 2026technical

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these

Cited by 0SourcePDFScholar
2026

Exploring the Underwater World Segmentation without Extra Training

CVPR 2026

Accurate segmentation of marine organisms is vital for biodiversity monitoring and ecological assessment, yet existing datasets and models remain largely limited to terrestrial scenes. To bridge this gap, we introduce **AquaOV255**, the first large-scale and fine-grained underwater segmentation data

Cited by 0SourcecodeScholar
2026

IntroSVG: Learning from Rendering Feedback for Text-to-SVG Generation via an Introspective Generator-Critic Framework

CVPR 2026

Scalable Vector Graphics (SVG) are central to digital design due to their inherent scalability and editability. Despite significant advancements in content generation enabled by Visual Language Models (VLMs), existing text-to-SVG generation methods are limited by a core challenge: the autoregressive

Cited by 0SourceScholar
2026

MedMamba: Multi-View State Space Models with Adaptive Graph Learning for Medical Time Series Classification

ICML 2026poster

Medical time series are central to healthcare, enabling continuous monitoring and supporting timely clinical decisions. Despite recent progress, existing methods struggle to jointly model local-global dynamics and handle nonstationarities like baseline drift, while often failing to capture latent ch…

Cited by 0SourceScholar
2026

OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning

CVPR 2026

Document AI has advanced rapidly and is attracting increasing attention. Yet, while most efforts have focused on document layout analysis (DLA), its generative counterpart, layout generation, remains underexplored. Distinct from traditional graphic layout design and room layout planning, document la

Cited by 0SourcecodeScholar
2026

PromptEnhancer: Taming Your Rewriter for Text-to-Image Generation via Fine-Grained Reward

CVPR 2026

Recent text-to-image (T2I) diffusion models have achieved impressive progress in generating high-fidelity images, yet they often fail to faithfully follow complex user prompts, especially in attribute binding, negation, and compositional reasoning. To address this limitation, we propose PromptEnhanc

Cited by 0SourcecodeScholar
2026

Tackling Time-Series Forecasting Generalization via Mitigating Concept Drift

ICLR 2026poster

Time-series forecasting finds broad applications in real-world scenarios. Due to the dynamic nature of time series data, it is important for time-series forecasting models to handle potential distribution shifts over time. In this paper, we initially identify two types of distribution shifts in time…

Cited by 0SourcecodeScholar
2026

TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level Effectiveness

ICLR 2026poster

Time-series forecasting is an essential task with wide real-world applications across domains. While recent advances in deep learning have enabled time-series forecasting models with accurate predictions, there remains considerable debate over which architectures and design components, such as serie…

Cited by 0SourcecodeScholar
2026

UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition

CVPR 2026

This paper introduces UniMERNet, a high-accuracy, computation-efficient algorithm for Mathematical Expression Recognition (MER) across diverse real-world scenarios. To facilitate UniMERNet's training, we constructed UniMER-1M, a million-scale dataset whose unprecedented diversity endows the model wi

Cited by 0SourcecodeScholar
2025

FKAN-GMFNet: Fourier Kolmogorov-Arnold-based Group Multi-scale Fusion Network for Aneurysm Image Segmentation

ICASSP 2025accepted

KAN-based networks, while offering improved interpretability compared to traditional models used in medical image segmentation, often struggle with limited adaptability to diverse imaging environments, making them less ideal for such tasks. To address this issue, we propose a Fourier Kolmogorov–Arno…

Cited by 1SourceScholar
2025

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges

ACL 2025long

The widespread adoption of Large Language Models (LLMs) has heightened concerns about their security, particularly their vulnerability to jailbreak attacks that leverage crafted prompts to generate malicious outputs. While prior research has been conducted on general security capabilities of LLMs, t…

2025

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

CVPR 2025poster

Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have not been fairly and comprehensively evaluated due to the nar…

2025

PUO-Bench: A Panel Understanding and Operation Benchmark with A Privacy-Preserving Framework

NeurIPS 2025poster

Recent advancements in Vision-Language Models (VLMs) have enabled GUI agents to leverage visual features for interface understanding and operation in the digital world. However, limited research has addressed the interpretation and interaction with control panels in real-world settings. To bridge th…

Cited by 0SourceScholar
2025

Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling

ICLR 2025poster

Controllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. T…

Cited by 0SourcePDFScholar
2025

WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code

ACL 2025finding

With the rapid advancement of Generative AI technology, Multimodal Large Language Models(MLLMs) have the potential to act as AI software engineers capable of executing complex web application development. Considering that the model requires a confluence of multidimensional sub-capabilities to addres…

2024

LSTPrompt: Large Language Models as Zero-Shot Time Series Forecasters by Long-Short-Term Prompting

ACL 2024findings

Time-series forecasting (TSF) finds broad applications in real-world scenarios. Prompting off-the-shelf Large Language Models (LLMs) demonstrates strong zero-shot TSF capabilities while preserving computational efficiency. However, existing prompting methods oversimplify TSF as language next-token p…

2024

MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation

CVPR 2024highlight

We present MicroCinema a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage proces…

Cited by 15SourcePDFScholar
2024

PINNsFormer: A Transformer-Based Framework For Physics-Informed Neural Networks

ICLR 2024poster

Physics-Informed Neural Networks (PINNs) have emerged as a promising deep learning framework for approximating numerical solutions to partial differential equations (PDEs). However, conventional PINNs, relying on multilayer perceptrons (MLP), neglect the crucial temporal dependencies inherent in pra…

2024

Time-MMD: Multi-Domain Multimodal Dataset for Time Series Analysis

NeurIPS 2024poster

Time series data are ubiquitous across a wide range of real-world domains. While real-world time series analysis (TSA) requires human experts to integrate numerical series data with multimodal domain-specific knowledge, most existing TSA models rely solely on numerical data, overlooking the signific…

2024

Time-Series Forecasting for Out-of-Distribution Generalization Using Invariant Learning

ICML 2024poster

Time-series forecasting (TSF) finds broad applications in real-world scenarios. Due to the dynamic nature of time-series data, it is crucial for TSF models to preserve out-of-distribution (OOD) generalization abilities, as training and test sets represent historical and future data respectively. In…

2023

Filler Word Detection with Hard Category Mining and Inter-Category Focal Loss

ICASSP 2023accepted

Filler words like "um" or "uh" are common in spontaneous speech. It is desirable to automatically detect and remove them in recordings, as they affect the fluency, confidence, and professionalism of speech. Previous studies and our preliminary experiments reveal that the biggest challenge in filler…

Cited by 0SourceScholar
2020

Joint Time-Frequency and Time Domain Learning for Speech Enhancement

IJCAI 2020poster

For single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framewo…

Cited by 0SourcePDFScholar