← Search

Morris Alper

7 accepted papers

2025

ProtoSnap: Prototype Alignment For Cuneiform Signs

ICLR 2025poster

The cuneiform writing system served as the medium for transmitting knowledge in the ancient Near East for a period of over three thousand years. Cuneiform signs have a complex internal structure which is the subject of expert paleographic analysis, as variations in sign shapes bear witness to histor…

2025

WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild

NeurIPS 2025poster

Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean multi-view training data, beyond manually curated datasets with limited diversity, camera variation, or licensing issues.…

Cited by 0SourceScholar
2024

ICC : Quantifying Image Caption Concreteness for Multimodal Dataset Curation

ACL 2024findings

Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in removing mismatched text-image pairs, but permit semantically related but highly a…

2024

Mitigating Open-Vocabulary Caption Hallucinations

EMNLP 2024main

While recent years have seen rapid progress in image-conditioned text generation, image captioning still suffers from the fundamental issue of hallucinations, namely, the generation of spurious details that cannot be inferred from the given image. Existing methods largely use closed-vocabulary objec…

2023

Is BERT Blind? Exploring the Effect of Vision-and-Language Pretraining on Visual Language Understanding

CVPR 2023poster

Most humans use visual imagination to understand and reason about language, but models such as BERT reason about language using knowledge acquired during text-only pretraining. In this work, we investigate whether vision-and-language pretraining can improve performance on text-only tasks that involv…