← Search

Yuyao Zhang

26 accepted papers

2026

A Survey of Joint Online-Offline Fine-tuning for Large Language Models

IJCAI 2026

Post-training for Large Language Models (LLMs) can be mainly categorized into offline Supervised Fine-Tuning (SFT) for knowledge acquisition and online Reinforcement Fine-Tuning (RFT) for adaptive refinement. Current state-of-the-art approaches typically employ a sequential cold-start pipeline (SFT-

Cited by 0Scholar
2026

HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing

CVPR 2026

High-resolution image editing is essential for professional and creative applications, yet existing multimodal diffusion-based editors remain computationally inefficient and constrained to relatively low resolutions. Current approaches redundantly process the entire image canvas or rely on large-sca

Cited by 0SourceScholar
2026

Improving 2D Diffusion Models for 3D Medical Imaging with Inter‑Slice Consistent Stochasticity

ICLR 2026poster

3D medical imaging is in high demand and essential for clinical diagnosis and scientific research. Currently, diffusion models have become an effective tool for medical imaging reconstruction thanks to their ability to learn rich, high‑quality data priors. However, learning the 3D data distribution…

Cited by 0SourcecodeScholar
2026

NICE: Neural Implicit Craniofacial Model for Orthognathic Surgery Prediction

AAAI 2026technical

Orthognathic surgery is a crucial intervention for correcting dentofacial skeletal deformities to enhance occlusal functionality and facial aesthetics. Accurate postoperative facial appearance prediction remains challenging due to the complex nonlinear interactions between skeletal movements and fac

Cited by 0SourcePDFScholar
2026

Plug-and-Play Diffusion Meets ADMM: Dual-Variable Coupling for Robust Medical Image Reconstruction

ICML 2026poster

Plug-and-Play diffusion prior (PnPDP) frameworks have emerged as a powerful paradigm for solving imaging inverse problems by treating pretrained generative models as modular priors. However, we identify a critical flaw in prevailing PnP solvers (e.g., based on HQS or Proximal Gradient): they functio…

Cited by 0SourceScholar
2026

Reference-Free Meta-Learning for Generalized Implicit Neural Representation in Efficient MRI Reconstruction

ICML 2026poster

Implicit Neural Representation (INR) has emerged as a powerful paradigm for continuous MRI reconstruction. However, standard unsupervised INR requires time-consuming optimization from scratch for each scan, hindering clinical deployment. This work presents IPOD, a Reference-Free Meta-Learning framew…

Cited by 0SourceScholar
2026

Resolving Blind Inverse Problems under Dynamic Range Compression via Structured Forward Operator Modeling

ICML 2026poster

Recovering radiometric fidelity from unknown dynamic range compression (UDRC), such as low-light enhancement and HDR reconstruction, is a challenging blind inverse problem, due to the unknown forward model and irreversible information loss introduced by compression. To address this challenge, we fir…

Cited by 0SourceScholar
2026

Unsupervised Motion-Compensated Decomposition for Cardiac MRI Reconstruction via Neural Representation

AAAI 2026technical

Cardiac magnetic resonance (CMR) imaging is widely used to characterize cardiac morphology and function. To accelerate CMR imaging, various methods have been proposed to recover high-quality spatiotemporal CMR images from highly undersampled k-t space data. However, current CMR reconstruction techni

Cited by 0SourcePDFScholar
2026

Unsupervised Multi-Parameter Inverse Solving for Reducing Ring Artifacts in 3D X-Ray CBCT

AAAI 2026technical

Ring artifacts are prevalent in 3D cone-beam computed tomography (CBCT) due to non-ideal responses of X-ray detectors, substantially affecting image quality and diagnostic reliability. Existing state-of-the-art (SOTA) ring artifact reduction (RAR) methods rely on supervised learning with large-scale

Cited by 0SourcePDFScholar
2026

Zero-shot Implicit Neural Manifold Representation (INMR) for Ultra-high Temporal Resolution Dynamic MRI

AAAI 2026technical

Capturing accurate dynamic information of moving organs is essential for functional assessment using non-invasive imaging modalities. Achieving high temporal resolution visualization of physiological processes remains a critical challenge in dynamic magnetic resonance imaging (MRI) when reconstructi

Cited by 0SourcePDFScholar
2025

Hierarchical Document Refinement for Long-context Retrieval-augmented Generation

ACL 2025long

Real-world RAG applications often encounter long-context input scenarios, where redundant information and noise results in higher inference costs and reduced performance. To address these challenges, we propose LongRefiner, an efficient plug-and-play refiner that leverages the inherent structural ch…

2025

LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration

NeurIPS 2025poster

Text-to-image (T2I) generation has made remarkable progress, yet existing systems still lack intuitive control over spatial composition, object consistency, and multi-step editing. We present **LayerCraft**, a modular framework that uses large language models (LLMs) as autonomous agents to orchestra…

Cited by 0SourceScholar
2025

Moner: Motion Correction in Undersampled Radial MRI with Unsupervised Neural Representation

ICLR 2025spotlight

Motion correction (MoCo) in radial MRI is a particularly challenging problem due to the unpredictability of subject movement. Current state-of-the-art (SOTA) MoCo algorithms often rely on extensive high-quality MR images to pre-train neural networks, which constrains the solution space and leads to…

2025

Neuro-Symbolic Query Compiler

ACL 2025finding

Precise recognition of search intent in Retrieval-Augmented Generation (RAG) systems remains a challenging goal, especially under resource constraints and for complex queries with nested structures and dependencies. This paper presents **QCompiler**, a neuro-symbolic framework inspired by linguistic…

2025

PD3F: A Pluggable and Dynamic DoS-Defense Framework against resource consumption attacks targeting Large Language Models

EMNLP 2025

Large Language Models (LLMs), due to substantial computational requirements, are vulnerable to resource consumption attacks, which can severely degrade server performance or even cause crashes, as demonstrated by denial-of-service (DoS) attacks designed for LLMs. However, existing works lack mitigat

2025

SVRMamba: Slice-to-Volume Reconstruction from Multiple MRI Stacks with Slice Sequence Guided Mamba

AAAI 2025technical

In fetal magnetic resonance imaging (MRI), slice-to-volume reconstruction (SVR) involves the computational creation of a 3D volume from multiple stacks of 2D slices. This process is challenging due to slice misalignment and image noise. Current state-of-the-art (SOTA) SVR methods typically employ co…

Cited by 0SourcePDFScholar
2025

Search-o1: Agentic Search-Enhanced Large Reasoning Models

EMNLP 2025

Large reasoning models (LRMs) like OpenAI-o1 have demonstrated impressive long stepwise reasoning capabilities through large-scale reinforcement learning. However, their extended reasoning processes often suffer from knowledge insufficiency, leading to frequent uncertainties and potential errors. To

2025

Unsupervised Self-Prior Embedding Neural Representation for Iterative Sparse-View CT Reconstruction

AAAI 2025technical

Emerging unsupervised implicit neural representation (INR) methods, such as NeRP, NeAT, and SCOPE, have shown great potential in addressing sparse-view computed tomography (SVCT) inverse problems. While these INR-based methods perform well on relatively dense SVCT reconstructions, they struggle to a…

2024

Causally Aware Generative Adversarial Networks for Light Pollution Control

AAAI 2024technical

Artificial light plays an integral role in modern cities, significantly enhancing human productivity and the efficiency of civilization. However, excessive illumination can lead to light pollution, posing non-negligible threats to economic burdens, ecosystems, and human health. Despite its critical…

2024

Enabling Discriminative Reasoning in LLMs for Legal Judgment Prediction

EMNLP 2024finding

Legal judgment prediction is essential for enhancing judicial efficiency. In this work, we identify that existing large language models (LLMs) underperform in this domain due to challenges in understanding case complexities and distinguishing between similar charges. To adapt LLMs for effective lega…

2024

Synergistic Patch Pruning for Vision Transformer: Unifying Intra- & Inter-Layer Patch Importance

ICLR 2024poster

The Vision Transformer (ViT) has emerged as a powerful architecture for various computer vision tasks. Nonetheless, this comes with substantially heavier computational costs than Convolutional Neural Networks (CNNs). The attention mechanism in ViTs, which integrates information from different image…

Cited by 5SourcePDFScholar
2023

Unsupervised Polychromatic Neural Representation for CT Metal Artifact Reduction

NeurIPS 2023poster

Emerging neural reconstruction techniques based on tomography (e.g., NeRF, NeAT, and NeRP) have started showing unique capabilities in medical imaging. In this work, we present a novel Polychromatic neural representation (Polyner) to tackle the challenging problem of CT imaging when metallic implant…

2022

Node-Aligned Graph Convolutional Network for Whole-Slide Image Representation and Classification

CVPR 2022oral

The large-scale whole-slide images (WSIs) facilitate the learning-based computational pathology methods. However, the gigapixel size of WSIs makes it hard to train a conventional model directly. Current approaches typically adopt multiple-instance learning (MIL) to tackle this problem. Among them, M…

Cited by 71PDFcodeScholar
2021

PIANO: A Parametric Hand Bone Model from Magnetic Resonance Imaging

IJCAI 2021poster

Hand modeling is critical for immersive VR/AR, action understanding, or human healthcare. Existing parametric models account only for hand shape, pose, or texture, without modeling the anatomical attributes like bone, which is essential for realistic hand biomechanics analysis. In this paper, we pre…

2017

Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition With Convolutional Neural Networks

CVPR 2017poster

Scene flow describes the motion of 3D objects in real world and potentially could be the basis of a good feature for 3D action recognition. However, its use for action recognition, especially in the context of convolutional neural networks (ConvNets), has not been previously studied. In this paper,…

Cited by 176PDFScholar
2016

Learning structured dictionary based on inter-class similarity and representative margins

ICASSP 2016accepted

We consider the problem of learning a structured and discriminative dictionary based on sparse representation for classification task. The structure comprises class-shared and class-specific partitions which allows the separation of common and class-specific information in the data for classificatio…

Cited by 0SourceScholar