← Search

David Li

8 accepted papers

2026

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

CVPR 2026

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-language modeling that extensively improves upon BabyVLM-V1 through a longitudinal,

Cited by 0SourcecodeScholar
2026

Diffusion & Adversarial Schrödinger Bridges via Iterative Proportional Markovian Fitting

ICLR 2026poster

The Iterative Markovian Fitting (IMF) procedure, which iteratively projects onto the space of Markov processes and the reciprocal class, successfully solves the Schrödinger Bridge (SB) problem. However, an efficient practical implementation requires a heuristic modification-alternating between fitti…

Cited by 0SourcecodeScholar
2026

GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

ICLR 2026poster

We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable knowledge-work tasks. GDPval covers the majority of Department of Labor O*NET Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are…

Cited by 0SourceScholar
2026

IDLM: Inverse-distilled Diffusion Language Models

ICML 2026poster

Diffusion Language Models (DLMs) have recently achieved strong results in text generation. However, their multi-step sampling leads to slow inference, limiting practical use. To address this, we extend Inverse Distillation, a technique originally developed to accelerate continuous diffusion models, …

Cited by 0SourceScholar
2026

One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation

ICML 2026poster

Diffusion models for super-resolution (SR) produce high-quality visual results but require expensive computational costs. Despite the development of several methods to accelerate diffusion-based SR models, some (e.g., SinSR) fail to produce realistic perceptual details, while others (e.g., OSEDiff) …

Cited by 0SourceScholar
2026

Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)

ICLR 2026oral

While achieving exceptional generative quality, modern diffusion, flow, and other matching models suffer from slow inference, as they require many steps of iterative generation. Recent distillation methods address this by training efficient one-step generators under the guidance of a pre-trained tea…

Cited by 0SourcecodeScholar
2025

Inverse Bridge Matching Distillation

ICML 2025poster

Learning diffusion bridge models is easy; making them fast and practical is an art. Diffusion bridge models (DBMs) are a promising extension of diffusion models for applications in image-to-image translation. However, like many modern diffusion and flow models, DBMs suffer from the problem of slow i…

Cited by 0SourcePDFScholar
2018

Cross-Lingual Phoneme Mapping for Language Robust Contextual Speech Recognition

ICASSP 2018accepted

Standard automatic speech recognition (ASR) systems are increasingly expected to recognize foreign entities, yet doing so while preserving accuracy on native words remains a challenge. We describe a novel approach for recognizing foreign words by injecting them with appropriate pronunciations into t…

Cited by 0SourceScholar