← Search

Zhiyuan Peng

8 accepted papers

2026

UGround: Towards Unified Visual Grounding with Unrolled Transformers

ICML 2026poster

We present UGround, a **U**nified visual **Ground**ing paradigm that dynamically selects intermediate layers across **U**nrolled transformers as "mask as prompt'', diverging from the prevailing pipeline that leverages the fixed last hidden layer as "$\texttt{\}$ as prompt''. UGround addresses two pr…

Cited by 0SourceScholar
2025

Language-Queried Target Sound Extraction Without Parallel Training Data

ICASSP 2025accepted

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are labor-intensive. We introduce a parallel-data-free training scheme,…

Cited by 0SourceScholar
2025

SolEval: Benchmarking Large Language Models for Repository-level Solidity Smart Contract Generation

EMNLP 2025

Large language models (LLMs) have transformed code generation.However, most existing approaches focus on mainstream languages such as Python and Java, neglecting the Solidity language, the predominant programming language for Ethereum smart contracts.Due to the lack of adequate benchmarks for Solidi

2023

Covariance Regularization for Probabilistic Linear Discriminant Analysis

ICASSP 2023accepted

Probabilistic linear discriminant analysis (PLDA) is commonly used in speaker verification systems to score the similarity of speaker embeddings. Recent studies improved the performance of PLDA in domain-matched conditions by diagonalizing its covariance. We suspect such a brutal pruning approach co…

Cited by 0SourceScholar
2020

Mixture Factorized Auto-Encoder for Unsupervised Hierarchical Deep Factorization of Speech Signal

ICASSP 2020accepted

Speech signal is constituted and contributed by various informative factors, such as linguistic content and speaker characteristic. There have been notable recent studies attempting to factorize speech signal into these individual factors without requiring any annotation. These studies typically ass…

Cited by 0SourceScholar
2019

Adversarial Multi-task Deep Features and Unsupervised Back-end Adaptation for Language Recognition

ICASSP 2019accepted

This paper presents an investigation into speaker-invariant feature learning and domain adaptation for language recognition (LR) with short utterances. While following the conventional design of i-vector front-end and probabilistic linear discriminant analysis (PLDA) back-end, we propose to apply sp…

Cited by 0SourceScholar