← Search

Qiji Zhou

4 accepted papers

2026

Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model

AAAI 2026technical

Direct Preference Optimization (DPO) simplifies reinforcement learning from human feedback (RLHF) for large language models (LLMs) by directly training on offline preference data to align with human preferences. During DPO training, the reference model serves as a data weight adjuster. However, the

Cited by 0SourcePDFScholar
2025

ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrieval

EMNLP 2025

Many contemporary data-driven research efforts in the natural sciences, such as chemistry and materials science, require large-scale, high-performance entity recognition from scientific datasets. Large language models (LLMs) have increasingly been adopted to solve the entity recognition task, with t

Cited by 0SourcePDFScholar
2025

Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation

ACL 2025finding

Counterfactual reasoning is crucial for robust video understanding but remains underexplored in existing multimodal benchmarks. In this paper, we introduce **COVER** (**CO**unterfactual **V**id**E**o **R**easoning), a multidimensional multimodal benchmark that systematically evaluates MLLMs across t…

2023

LogiCoT: Logical Chain-of-Thought Instruction Tuning

EMNLP 2023long findings

Generative Pre-trained Transformer 4 (GPT-4) demonstrates impressive chain-of-thought reasoning ability. Recent work on self-instruction tuning, such as Alpaca, has focused on enhancing the general proficiency of models. These instructions enable the model to achieve performance comparable to GPT-3…

Cited by 0SourcecodeScholar