← Search

Davor Golac

3 accepted papers

2025

CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories

NAACL 2025long

The increasing complexity of computer science research projects demands more effective tools for deploying code repositories. Large Language Models (LLMs), such as Anthropic Claude and Meta Llama, have demonstrated significant advancements across various fields of computer science research, includin…

2025

Learning LLM Preference over Intra-Dialogue Pairs: A Framework for Utterance-level Understandings

NAACL 2025industry

Large language models (LLMs) have demonstrated remarkable capabilities in handling complex dialogue tasks without requiring use case-specific fine-tuning. However, analyzing live dialogues in real-time necessitates low-latency processing systems, making it impractical to deploy models with billions…

Cited by 0SourcePDFScholar
2024

BPID: A Benchmark for Personal Identity Deduplication

EMNLP 2024industry

Data deduplication is a critical task in data management and mining, focused on consolidating duplicate records that refer to the same entity. Personally Identifiable Information (PII) is a critical class of data for deduplication across various industries. Consumer data, stored and generated throug…

Cited by 0SourcePDFScholar