← Search

Zhiwei Deng

18 accepted papers

2026

Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification

CVPR 2026

Text-to-image (T2I) models are increasingly used for synthetic dataset generation, but generating synthetic training data to improve fine-grained classification performance remains challenging. Fine-tuning the T2I model with a few real examples can help generate more appropriate synthetic training d

Cited by 1SourcecodeScholar
2025

Dynamic Diffusion Schrödinger Bridge in Astrophysical Observational Inversions

NeurIPS 2025poster

We study Diffusion Schrödinger Bridge (DSB) models in the context of dynamical astrophysical systems, specifically tackling observational inverse prediction tasks within Giant Molecular Clouds (GMCs) for star formation. We introduce the Astro-DSB model, a variant of DSB with the pairwise domain assu…

Cited by 0SourcecodeScholar
2024

A Label is Worth A Thousand Images in Dataset Distillation

NeurIPS 2024poster

Data *quality* is a crucial factor in the performance of machine learning models, a principle that dataset distillation methods exploit by compressing training datasets into much smaller counterparts that maintain similar downstream performance. Understanding how and why data distillation methods wo…

2023

A Zero-Shot Language Agent for Computer Control with Structured Reflection

EMNLP 2023long findings

Large language models (LLMs) have shown increasing capacity at planning and executing a high-level goal in a live computer environment (e.g. MiniWoB++). To perform a task, recent works often require a model to learn from trace examples of the task via either supervised learning or few/many-shot prom…

Cited by 0SourceScholar
2023

Boundary Guided Learning-Free Semantic Control with Diffusion Models

NeurIPS 2023poster

Applying pre-trained generative denoising diffusion models (DDMs) for downstream tasks such as image semantic editing usually requires either fine-tuning DDMs or learning auxiliary editing networks in the existing literature. In this work, we present our BoundaryDiffusion method for efficient, effec…

2022

Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks

NeurIPS 2022accept

We propose an algorithm that compresses the critical information of a large dataset into compact addressable memories. These memories can then be recalled to quickly re-train a neural network and recover the performance (instead of storing and re-training on the full original dataset). Building upon…

2020

Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language Navigation

NeurIPS 2020poster

The ability to perform effective planning is crucial for building an instruction-following agent. When navigating through a new environment, an agent is challenged with (1) connecting the natural language instructions with its progressively growing knowledge of the world; and (2) performing long-ran…

Cited by 95SourcePDFScholar
2019

Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and Synthesis

ICCV 2019poster

We present a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations. Our dataset features synchronized body and finger motion as well as audio data. To the best of our knowledge, it represents the largest motion capture and audio dataset of natural conve…

Cited by 118PDFScholar
2018

Probabilistic Neural Programmed Networks for Scene Generation

NeurIPS 2018spotlight

In this paper we address the text to scene image generation problem. Generative models that capture the variability in complicated scenes containing rich semantics is a grand goal of image generation. Complicated scene images contain rich visual elements, compositional visual concepts, and complicat…

2018

Sparsely Aggregated Convolutional Networks

ECCV 2018poster

We explore a key architectural aspect of deep convolutional neural networks: the pattern of internal skip connections used to aggregate outputs of earlier layers for consumption by deeper layers. Such aggregation is critical to facilitate training of very deep networks in an end-to-end manner. This…

2017

Factorized Variational Autoencoders for Modeling Audience Reactions to Movies

CVPR 2017poster

Matrix and tensor factorization methods are often used for finding underlying low-dimensional patterns from noisy data. In this paper, we study non-linear tensor factoriza- tion methods based on deep variational autoencoders. Our approach is well-suited for settings where the relationship between th…

Cited by 68PDFScholar
2016

A Hierarchical Deep Temporal Model for Group Activity Recognition

CVPR 2016poster

In group activity recognition, the temporal dynamics of the whole activity can be inferred based on the dynamics of the individual people representing the activity. We build a deep model to capture these dynamics based on LSTM (long short-term memory) models. To make use of these observations, we pr…

Cited by 642PDFcodeScholar
2016

Learning Structured Inference Neural Networks With Label Relations

CVPR 2016poster

Images of scenes have various objects as well as abundant attributes, and diverse levels of visual categorization are possible. A natural image could be assigned with fine-grained labels that describe major components, coarse-grained labels that depict high level abstraction or a set of labels that…

Cited by 162PDFScholar
2016

Structure Inference Machines: Recurrent Neural Networks for Analyzing Relations in Group Activity Recognition

CVPR 2016poster

Rich semantic relations are important in a variety of visual recognition problems. As a concrete example, group activity recognition involves the interactions and relative spatial relations of a set of people in a scene. State of the art recognition methods center on deep learning approaches for tr…

Cited by 308PDFScholar