← Search

Zheng Zhao

22 accepted papers

2026

Utilising Gradient-Based Proposals Within Sequential Monte Carlo Samplers for Training of Partial Bayesian Neural Networks

ICASSP 2026poster

Partial Bayesian neural networks (pBNNs) have been shown to perform competitively with fully Bayesian neural networks while only having a subset of the parameters be stochastic. Using sequential Monte Carlo (SMC) samplers as the inference method for pBNNs gives a non-parametric probabilistic estimat…

Cited by 0SourcePDFScholar
2026

Verifying Chain-of-Thought Reasoning via Its Computational Graph

ICLR 2026oral

Current Chain-of-Thought (CoT) verification methods predict reasoning correctness based on outputs (black-box) or activations (gray-box), but offer limited insight into \textit{why} a computation fails. We introduce a white-box method: \textbf{Circuit-based Reasoning Verification (CRV)}. We hypothes…

Cited by 0SourcecodeScholar
2025

Conditioning diffusion models by explicit forward-backward bridging

AISTATS 2025poster

Given an unconditional diffusion model targeting a joint model $\pi(x, y)$, using it to perform conditional simulation $\pi(x \mid y)$ is still largely an open question and is typically achieved by learning conditional drifts to the denoising SDE after the fact. In this work, we express \emph{exact}…

Cited by 0SourcecodeScholar
2025

HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual Perceiver

CVPR 2025poster

This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified segmentation methods, limitations in adaptation to both image and video scenarios, as…

2025

Iterative Multilingual Spectral Attribute Erasure

EMNLP 2025

Multilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between languages. However, existing methods for debiassing are unable to exploit this opportunity because they operate on individual

Cited by 0SourcePDFScholar
2025

PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants

ACL 2025finding

Large language models (LLMs) have advanced conversational AI assistants. However, systematically evaluating how well these assistants apply personalization—adapting to individual user preferences while completing tasks—remains challenging. Existing personalization benchmarks focus on chit-chat, non-…

2025

Solving Linear-Gaussian Bayesian Inverse Problems with Decoupled Diffusion Sequential Monte Carlo

ICML 2025poster

A recent line of research has exploited pre-trained generative diffusion models as priors for solving Bayesian inverse problems. We contribute to this research direction by designing a sequential Monte Carlo method for linear-Gaussian inverse problems which builds on ``decoupled diffusion", where th…

2025

What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations

ACL 2025long

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains. VISTA contains 18,599 recorded AI conference presentations paire…

2024

Are Large Language Model Temporally Grounded?

NAACL 2024long

Are Large Language Models (LLMs) temporally grounded? Since LLMs cannot perceive and interact with the environment, it is impossible to answer this question directly. Instead, we provide LLMs with textual narratives and probe them with respect to their common-sense knowledge of the structure and dur…

Cited by 19SourcePDFScholar
2024

Controlling Vision-Language Models for Multi-Task Image Restoration

ICLR 2024poster

Vision-language models such as CLIP have shown great impact on diverse downstream tasks for zero-shot or label-free predictions. However, when it comes to low-level vision such as image restoration their performance deteriorates dramatically due to corrupted inputs. In this paper, we present a degra…

2024

Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding

NAACL 2024long

Large language models (LLMs) tend to inadequately integrate input context during text generation, relying excessively on encoded prior knowledge in model parameters, potentially resulting in generated text with factual inconsistencies or contextually unfaithful content. LLMs utilize two primary know…

2024

Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models

EMNLP 2024main

Fine-tuning pre-trained large language models (LLMs) on a diverse array of tasks has become a common approach for building models that can solve various natural language processing (NLP) tasks. However, where and to what extent these models retain task-specific knowledge remains largely unexplored.…

2024

On Feynman-Kac training of partial Bayesian neural networks

AISTATS 2024poster

Recently, partial Bayesian neural networks (pBNNs), which only consider a subset of the parameters to be stochastic, were shown to perform competitively with full Bayesian neural networks. However, pBNNs are often multi-modal in the latent variable space and thus challenging to approximate with para…

2024

Spectral Editing of Activations for Large Language Model Alignment

NeurIPS 2024poster

Large language models (LLMs) often exhibit undesirable behaviours, such as generating untruthful or biased content. Editing their internal representations has been shown to be effective in mitigating such behaviours on top of the existing alignment methods. We propose a novel inference-time editing…

2023

A Joint Matrix Factorization Analysis of Multilingual Representations

EMNLP 2023long findings

We present an analysis tool based on joint matrix factorization for comparing latent representations of multilingual and monolingual models. An alternative to probing, this tool allows us to analyze multiple sets of representations in a joint manner. Using this tool, we study to what extent and how…

Cited by 0SourcecodeScholar
2023

HPFTN: Hierarchical Progressive Fusion Transformer Network for Video Denoising

ICASSP 2023accepted

This paper presents a simple yet effective approach to modeling space-time correspondences in the context of video denoising. Unlike most existing approaches, our method, namely HPFTN, can operate end-to-end on consecutive frames without motion estimation. To do so, the proposed hierarchical patch m…

Cited by 0SourceScholar
2023

Image Restoration with Mean-Reverting Stochastic Differential Equations

ICML 2023poster

This paper presents a stochastic differential equation (SDE) approach for general-purpose image restoration. The key construction consists in a mean-reverting SDE that transforms a high-quality image into a degraded counterpart as a mean state with fixed Gaussian noise. Then, by simulating the corre…

2023

PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India

EMNLP 2023long findings

This paper introduces PMIndiaSum, a multilingual and massively parallel summarization corpus focused on languages in India. Our corpus provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs. We detail our construction workflow…

Cited by 0SourcecodeScholar
2021

Efficient On-Chip Learning for Optical Neural Networks Through Power-Aware Sparse Zeroth-Order Optimization

AAAI 2021technical

Optical neural networks (ONNs) have demonstrated record-breaking potential in high-performance neuromorphic computing due to their ultra-high execution speed and low energy consumption. However, current learning protocols fail to provide scalable and efficient solutions to photonic circuit optimizat…

Cited by 33SourcePDFScholar
2020

State-Space Gaussian Process for Drift Estimation in Stochastic Differential Equations

ICASSP 2020accepted

This paper is concerned with the estimation of unknown drift functions of stochastic differential equations (SDEs) from observations of their sample paths. We propose to formulate this as a non-parametric Gaussian process regression problem and use an Ito-Taylor expansion for approximating the SDE.…

Cited by 0SourceScholar