← Search

Nebojsa Jojic

22 accepted papers

2025

Fast constrained sampling in pre-trained diffusion models

NeurIPS 2025poster

Large denoising diffusion models, such as Stable Diffusion, have been trained on billions of image-caption pairs to perform text-conditioned image generation. As a byproduct of this training, these models have acquired general knowledge about image statistics, which can be useful for other inference…

Cited by 0SourcecodeScholar
2025

Make Some Noise: Towards LLM audio reasoning and generation using sound tokens

ICASSP 2025accepted

Integrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. Here, we introduce a novel approach that combines Variational Quantization with Conditional Flow Matching to convert audio…

Cited by 0SourceScholar
2025

fLSA: Learning Semantic Structures in Document Collections Using Foundation Models

EMNLP 2025

Humans can learn to solve new tasks by inducing high-level strategies from example solutions to similar problems and then adapting these strategies to solve unseen problems. Can we use large language models to induce such high-level structure from example documents or solutions? We introduce fLSA, a

Cited by 0SourcePDFScholar
2024

PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization

ICLR 2024poster

Expert-level prompts, carefully engineered by human experts who have a deep understanding of both large language models (LLMs) and domain knowledge, are the future of prompting and pivotal to harnessing the full power of advanced LLMs. Discovering such prompts with an automated process remains a sou…

2024

Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling

ICML 2024poster

We introduce Reprompting, an iterative sampling algorithm that automatically learns the Chain-of-Thought (CoT) recipes for a given task without human intervention. Through Gibbs sampling, Reprompting infers the CoT recipes that work consistently well for a set of training samples by iteratively samp…

Cited by 29SourcePDFScholar
2023

Evaluating Cognitive Maps and Planning in Large Language Models with CogEval

NeurIPS 2023poster

Recently an influx of studies claims emergent cognitive abilities in large language models (LLMs). Yet, most rely on anecdotes, overlook contamination of training sets, or lack systematic Evaluation involving multiple tasks, control conditions, multiple iterations, and statistical robustness tests.…

Cited by 64SourcePDFScholar
2023

ThinkSum: Probabilistic reasoning over sets using large language models

ACL 2023long

Large language models (LLMs) have a substantial capacity for high-level analogical reasoning: reproducing patterns in linear text that occur in their training data (zero-shot evaluation) or in the provided context (few-shot in-context learning). However, recent studies show that even the more advanc…

Cited by 31SourcePDFScholar
2022

Coherence boosting: When your pretrained language model is not paying enough attention

ACL 2022long

Long-range semantic coherence remains a challenge in automatic language generation and understanding. We demonstrate that large language models have insufficiently learned the effect of distant words on next-token prediction. We present coherence boosting, an inference procedure that increases a LM’…

2022

Diffusion Models as Plug-and-Play Priors

NeurIPS 2022accept

We consider the problem of inferring high-dimensional data $x$ in a model that consists of a prior $p(x)$ and an auxiliary differentiable constraint $c(x,y)$ on $x$ given some additional information $y$. In this paper, the prior is an independently trained denoising diffusion generative model. The a…

2022

Resolving label uncertainty with implicit posterior models

UAI 2022poster

We propose a method for jointly inferring labels across a collection of data samples, where each sample consists of an observation and a prior belief about the label. By implicitly assuming the existence of a generative model for which a differentiable predictor is the posterior, we derive a trainin…

2021

GPT Perdetry Test: Generating new meanings for new words

NAACL 2021long

Human innovation in language, such as inventing new words, is a challenge for pretrained language models. We assess the ability of one large model, GPT-3, to process new words and decide on their meaning. We create a set of nonce words and prompt GPT-3 to generate their dictionary definitions. We fi…

2021

Multi-Label Learning From Single Positive Labels

CVPR 2021poster

Predicting all applicable labels for a given image is known as multi-label classification. Compared to the standard multi-class case (where each image has only one label), it is considerably more challenging to annotate training data for multi-label classification. When the number of potential label…

Cited by 133PDFcodeScholar
2021

Studying word order through iterative shuffling

EMNLP 2021main

As neural language models approach human performance on NLP benchmark tasks, their advances are widely seen as evidence of an increasingly complex understanding of syntax. This view rests upon a hypothesis that has not yet been empirically tested: that word order encodes meaning essential to perform…

2020

FSNet: Compression of Deep Convolutional Neural Networks by Filter Summary

ICLR 2020poster

We present a novel method of compression of deep Convolutional Neural Networks (CNNs) by weight sharing through a new representation of convolutional filters. The proposed method reduces the number of parameters of each convolutional layer by learning a $1$D vector termed Filter Summary (FS). The co…

Cited by 21SourceScholar
2020

Local Context Normalization: Revisiting Local Normalization

CVPR 2020oral

Normalization layers have been shown to improve convergence in deep neural networks, and even add useful inductive biases. In many vision applications the local spatial context of the features is important, but most common normalization schemes including Group Normalization (GN), Instance Normalizat…

Cited by 36PDFcodeScholar
2020

Mining self-similarity: Label super-resolution with epitomic representations

ECCV 2020poster

We show that simple patch-based models, such as epitomes (Jojic et al., 2003), can have superior performance to the current state of the art in semantic segmentation and label super-resolution, which uses deep convolutional neural networks. We derive a new training algorithm for epitomes which allow…

2019

Label super-resolution networks

ICLR 2019poster

We present a deep learning-based method for super-resolving coarse (low-resolution) labels assigned to groups of image pixels into pixel-level (high-resolution) labels, given the joint distribution between those low- and high-resolution labels. This method involves a novel loss function that minimiz…

Cited by 37SourcePDFScholar
2019

Large Scale High-Resolution Land Cover Mapping With Multi-Resolution Data

CVPR 2019poster

In this paper we propose multi-resolution data fusion methods for deep learning-based high-resolution land cover mapping from aerial imagery. The land cover mapping problem, at country-level scales, is challenging for common deep learning methods due to the scarcity of high-resolution labels, as wel…

Cited by 124PDFcodeScholar
2018

WSNet: Compact and Efficient Networks Through Weight Sampling

ICML 2018oral

We present a new approach and a novel architecture, termed WSNet, for learning compact and efficient deep neural networks. Existing approaches conventionally learn full model parameters independently and then compress them via ad hoc processing such as model pruning or filter factorization. Alternat…

2017

ER3: A Unified Framework for Event Retrieval, Recognition and Recounting

CVPR 2017poster

We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames an…

Cited by 28PDFScholar
2017

Summarization and Classification of Wearable Camera Streams by Learning the Distributions Over Deep Features of Out-Of-Sample Image Sequences

ICCV 2017poster

A popular approach to training classifiers of new image classes is to use lower levels of a pre-trained feed-forward neural network and retrain only the top. Thus, most layers simply serve as highly nonlinear feature extractors. While these features were found useful for classifying a variety of sce…

Cited by 8PDFScholar
2016

Iterative Refinement of the Approximate Posterior for Directed Belief Networks

NeurIPS 2016poster

Variational methods that rely on a recognition network to approximate the posterior of directed graphical models offer better inference and learning than previous methods. Recent advances that exploit the capacity and flexibility in this approach have expanded what kinds of models can be trained. Ho…