← Search

Cheng Zhuo

9 accepted papers

2026

Circuit-Think: A Multimodal Reasoning Framework for Automated Circuit-to-Netlist Translation with Trajectory-Guided Reinforcement Learning

AAAI 2026technical

Vision Language Models (VLMs) have shown strong performance in multimodal understanding, offering promise for the circuit-to-netlist translation task. However, the diverse component symbols and complex connections in circuit images challenge VLMs in understanding physical layouts and reasoning for e

Cited by 0SourcePDFScholar
2026

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography

ICML 2026poster

As semiconductor technology nodes continue to shrink, computational lithography has become critical to yield and performance. However, real-world lithography is a continuous, multi-stage physical process driven by implicit interventions, which cannot be captured by the existing static or stage-wise …

Cited by 0SourceScholar
2026

MaxMark: High-Capacity Diffusion-Native Watermarking via Robust and Invertible Latent Embedding

CVPR 2026

Diffusion-native watermarking provides a more secure and reliable way to trace images from latent diffusion models (LDMs) by embedding information directly into the generative process. However, existing methods suffer from a fundamental limitation: their embedding capacity is extremely small. We int

Cited by 0SourcecodeScholar
2026

SMD: Multi-view Safety-Critical Driving Video Generation in the Real-world Domain

ICML 2026poster

Safety-critical scenarios are essential for evaluating autonomous driving (AD) systems, yet they are rare in practice. Existing generators produce trajectories, simulations, or single-view videos—but they don’t meet what modern AD systems actually consume: realistic multi-view video. We present SMD,…

Cited by 4SourcecodeScholar
2026

SafeCompass: Dynamic Chain-of-Thought Steering via Inference-Time Safety Signals

ICML 2026poster

Large reasoning models (LRMs) achieve strong performance by explicitly generating chain-of-thought (CoT) reasoning, but this reasoning process can be manipulated by adversarial prompts. Inference-time CoT interventions offer a simple and lightweight approach to improving safety, yet existing methods…

Cited by 0SourceScholar
2026

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

ICML 2026poster

Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are largely incompatible with speculative inference: they either introduce additional computation or disrupt the draft–verify mechanism, negating acceleration be…

Cited by 0SourceScholar
2025

SilentStriker: Toward Stealthy Bit-Flip Attacks on Large Language Models

NeurIPS 2025poster

The rapid adoption of large language models (LLMs) in critical domains has spurred extensive research into their security issues. While input manipulation attacks (e.g., prompt injection) have been well-studied, Bit-Flip Attacks (BFAs)—which exploit hardware vulnerabilities to corrupt model paramete…

Cited by 0SourceScholar
2023

HDNet: Hierarchical Dynamic Network for Gait Recognition using Millimeter-wave radar

ICASSP 2023accepted

Gait recognition is widely used in diversified practical applications. Currently, the most prevalent approach is to recognize human gait from RGB images, owing to the progress of computer vision technologies. Nevertheless, the perception capability of RGB cameras deteriorates in rough circumstances,…

Cited by 0SourceScholar
2020

Cross-denoising Network against Corrupted Labels in Medical Image Segmentation with Domain Shift

IJCAI 2020poster

Deep convolutional neural networks (DCNNs) have contributed many breakthroughs in segmentation tasks, especially in the field of medical imaging. However, domain shift and corrupted annotations, which are two common problems in medical imaging, dramatically degrade the performance of DCNNs in practi…

Cited by 0SourcePDFScholar