← Search

Daniel Li

3 accepted papers

2025

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

CVPR 2025poster

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP, self-supervised visual features are not readily aligned with language, hindering their adoption in open-v…

Cited by 5SourcePDFScholar
2021

Sentence Boundary Augmentation for Neural Machine Translation Robustness

ICASSP 2021accepted

Neural Machine Translation (NMT) models have demonstrated strong state of the art performance on translation tasks where well-formed training and evaluation data are pro-vided, but they remain sensitive to inputs that include errors of various types. Specifically, in the context of long-form speech…

Cited by 0SourceScholar
2018

Adaptive Memory Networks

ICLR 2018workshop

Real-world Question Answering (QA) tasks consist of thousands of words that often represent many facts and entities. Existing models based on LSTMs require a large number of parameters to support external memory and do not generalize well for long sequence inputs. Memory networks attempt to address…

Cited by 0SourceScholar