← Search

Adian Liusie

7 accepted papers

2024

Efficient LLM Comparative Assessment: A Product of Experts Framework for Pairwise Comparisons

EMNLP 2024main

LLM-as-a-judge approaches are a practical and effective way of assessing a range of text tasks. However, when using pairwise comparisons to rank a set of candidates, the computational cost scales quadratically with the number of candidates, which has practical limitations. This paper introduces a Pr…

2024

Investigating the Emergent Audio Classification Ability of ASR Foundation Models

NAACL 2024long

Text and vision foundation models can perform many tasks in a zero-shot setting, a desirable property that enables these systems to be applied in general and low-resource settings. There has been far less work, however, on the zero-shot abilities of ASR foundation models, with these systems typicall…

2024

Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

EMNLP 2024main

Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems. Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to adversarial manipulation. This work presents the…

2024

Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models

ACL 2024findings

Large Language Models (LLMs) have demonstrated impressive zero-shot capabilities and versatility in NLP tasks, however they sometimes fail to maintain crucial invariances for specific tasks. One example is permutation sensitivity, where LLMs’ outputs may significantly vary depending on the order of…

Cited by 6SourcePDFScholar
2024

WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models

NAACL 2024findings

Watermarking generative-AI systems, such as LLMs, has gained considerable interest, driven by their enhanced capabilities across a wide range of tasks. Although current approaches have demonstrated that small, context-dependent shifts in the word distributions can be used to apply and detect waterma…

Cited by 16SourcePDFScholar
2023

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

EMNLP 2023long main

Generative Large Language Models (LLMs) such as GPT-3 are capable of generating highly fluent responses to a wide variety of user prompts. However, LLMs are known to hallucinate facts and make non-factual statements which can undermine trust in their output. Existing fact-checking approaches either…

Cited by 0SourcecodeScholar