← Search

David Zhang

14 accepted papers

2026

Prompt Tuning for CLIP on the Pretrained Manifold

ICML 2026poster

Prompt tuning introduces learnable prompt vectors that adapt pretrained vision-language models to downstream tasks in a parameter-efficient manner. However, under limited supervision, prompt tuning alters pretrained representations and drives downstream features away from the pretrained manifold tow…

Cited by 0SourceScholar
2026

Toward Training Superintelligent Software Agents through Self-Play SWE-RL

ICML 2026poster

While current software agents powered by large language models (LLMs) and reinforcement learning (RL) can boost programmer productivity, their reliance on human-curated training data and environments creates a fundamental barrier to superintelligence. In this paper, we present Self-play SWE-RL (SSR)…

Cited by 0SourceScholar
2025

Non-Markovian Discrete Diffusion with Causal Language Models

NeurIPS 2025poster

Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current s…

Cited by 0SourceScholar
2024

Cell2Sentence: Teaching Large Language Models the Language of Biology

ICML 2024poster

We introduce Cell2Sentence (C2S), a novel method to directly adapt large language models to a biological context, specifically single-cell transcriptomics. By transforming gene expression data into "cell sentences," C2S bridges the gap between natural language processing and biology. We demonstrate…

Cited by 19SourcePDFScholar
2024

Tools Identification By On-Board Adaptation of Vision-and-Language Models

AAAI 2024technical

A robotic workshop assistant has been a long-standing grand challenge for robotics, speech, computer vision, and artificial intelligence (AI) research. We revisit the goal of visual identification of tools from human queries in the current era of Large Vision-and-Language models (like GPT-4). We fin…

Cited by 1SourcePDFScholar
2023

Learning ASR Pathways: A Sparse Multilingual ASR Model

ICASSP 2023accepted

Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, language-agnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard important language-spe…

Cited by 0SourceScholar
2023

Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities

ICASSP 2023accepted

End-to-end multilingual ASR has become more appealing because of several reasons such as simplifying the training and deployment process and positive performance transfer from high-resource to low-resource languages. However, scaling up the number of languages, total hours, and number of unique toke…

Cited by 0SourceScholar
2022

Learning Modal-Invariant and Temporal-Memory for Video-Based Visible-Infrared Person Re-Identification

CVPR 2022poster

Thanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the "probe-to-gallery", almost all existing RGB-IR based cro…

Cited by 64PDFcodeScholar
2018

A Hybrid l1-l0 Layer Decomposition Model for Tone Mapping

CVPR 2018poster

Tone mapping aims to reproduce a standard dynamic range image from a high dynamic range image with visual information preserved. State-of-the-art tone mapping algorithms mostly decompose an image into a base layer and a detail layer, and process them accordingly. These methods may have problems of h…

Cited by 175SourcePDFScholar
2018

Learning Convolutional Networks for Content-Weighted Image Compression

CVPR 2018poster

Lossy image compression is generally formulated as a joint rate-distortion optimization problem to learn encoder, quantizer, and decoder. Due to the non-differentiable quantizer and discrete entropy estimation, it is very challenging to develop a convolutional network (CNN)-based image compression…

Cited by 490SourcePDFScholar
2017

Multi-Channel Weighted Nuclear Norm Minimization for Real Color Image Denoising

ICCV 2017poster

Most of the existing denoising algorithms are developed for grayscale images. It is not trivial to extend them for color image denoising since the noise statistics in R, G, and B channels can be very different for real noisy images. In this paper, we propose a multi-channel (MC) optimization model f…

Cited by 359PDFScholar
2016

Joint Learning of Single-Image and Cross-Image Representations for Person Re-Identification

CVPR 2016poster

Person re-identification has been usually solved as either the matching of single-image representation (SIR) or the classification of cross-image representation (CIR). In this work, we exploit the connection between these two categories of methods, and propose a joint learning framework to unify SIR…

Cited by 518PDFScholar
2015

Patch Group Based Nonlocal Self-Similarity Prior Learning for Image Denoising

ICCV 2015poster

Patch based image modeling has achieved a great success in low level vision such as image denoising. In particular, the use of image nonlocal self-similarity (NSS) prior, which refers to the fact that a local patch often has many nonlocal similar patches to it across the image, has significantly enh…

Cited by 472PDFScholar