← Search

Dante Everaert

4 accepted papers

2025

Towards Knowledge Checking in Retrieval-augmented Generation: A Representation Perspective

NAACL 2025long

Retrieval-Augmented Generation (RAG) systems have shown promise in enhancing the performance of Large Language Models (LLMs). However, these systems face challenges in effectively integrating external knowledge with the LLM’s internal knowledge, often leading to issues with misleading or unhelpful i…

2024

AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset

EMNLP 2024industry

Query Autocomplete (QAC) is a critical feature in modern search engines, facilitating user interaction by predicting search queries based on input prefixes. Despite its widespread adoption, the absence of large-scale, realistic datasets has hindered advancements in QAC system development. This paper…

Cited by 1SourcePDFScholar
2024

GIO: Gradient Information Optimization for Training Dataset Selection

ICLR 2024spotlight

It is often advantageous to train models on a subset of the available train examples, because the examples are of variable quality or because one would like to train with fewer examples, without sacrificing performance. We present Gradient Information Optimization (GIO), a scalable, task-agnostic ap…

2024

Retrieval Augmented Spelling Correction for E-Commerce Applications

EMNLP 2024industry

The rapid introduction of new brand names into everyday language poses a unique challenge for e-commerce spelling correction services, which must distinguish genuine misspellings from novel brand names that use unconventional spelling. We seek to address this challenge via Retrieval Augmented Genera…

Cited by 0SourcePDFScholar