← Search

Yuehai Wang

7 accepted papers

2026

LMM4-IC4K: A Large Multimodal Model Powered Integrated Circuit Footprint Geometry Understanding

ICML 2026poster

Printed-Circuit-board (PCB) footprint geometry labeling of integrated circuits (IC) is essential in defining the physical interface between components and the PCB layout, requiring precise visual perception. However, the unstructured nature of footprint drawings and abstract diagram annotations prev…

Cited by 0SourceScholar
2025

Do Less and Achieve More: Free Condition Video Outpainting with Diffusion Model

ICASSP 2025accepted

Video outpainting aims to extend the content of a video beyond its original spatial boundaries. Existing methods tend to condition the generation process on a single frame or caption, failing to address the challenge in long videos with multiple video clips. To address this, we extend the diffusion-…

Cited by 0SourceScholar
2025

ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models

COLING 2025main

Text-to-speech (TTS) has seen significant advancements in high-quality, expressive speech synthesis. However, achieving diverse and natural prosody in synthesized speech remains challenging. In this paper, we propose ProsodyFlow, an end-to-end TTS model that integrates large self-supervised speech m…

Cited by 0SourcePDFScholar
2024

Audio Deepfake Detection With Self-Supervised Wavlm And Multi-Fusion Attentive Classifier

ICASSP 2024accepted

With the rapid development of speech synthesis and voice conversion technologies, Audio Deepfake has become a serious threat to the Automatic Speaker Verification (ASV) system. Numerous countermeasures are proposed to detect this type of attack. In this paper, we report our efforts to combine the se…

Cited by 0SourceScholar
2020

Collaborative Distillation for Ultra-Resolution Universal Style Transfer

CVPR 2020poster

Universal style transfer methods typically leverage rich representations from deep Convolutional Neural Network (CNN) models (e.g., VGG-19) pre-trained on large collections of images. Despite the effectiveness, its application is heavily constrained by the large model size to handle ultra-resolution…

Cited by 134PDFcodeScholar
2016

Multiple scattering effects on the localization of two point scatterers

ICASSP 2016accepted

Multiple scattering effects are commonly ignored in the detection and estimation of scatterers in signal processing research, because the energy of the first-order scattering is much larger than that of higher-order components. Although multiple scattering can significantly increase the estimation p…

Cited by 0SourceScholar