WritingCohereCoherepublished Oct 22, 2023seen 4w

Which Prompts Make The Difference Data Prioritization For Efficient Human Llm Evaluation 2023 10 22

Open original ↗

Captured source

source ↗

Which Prompts Make The Difference? Data Prioritization For Efficient Human LLM Evaluation North Mini Code. Cohere's first model for developers. Learn more

Oct 22, 2023 Which Prompts Make The Difference? Data Prioritization For Efficient Human LLM Evaluation Human evaluation of LLMs is critical, but comes at a high cost.

This research proposes prompt ranking methods to make pairwise human evaluation more efficient.

Read the paper

Authors

Meriem Boubdir, Edward Kim, Beyza Ermis, Marzieh Fadaee, Sara Hooker

Abstract

Human evaluation is increasingly critical for assessing large language models, capturing linguistic nuances, and reflecting user preferences more accurately than traditional automated metrics. However, the resource-intensive nature of this type of annotation process poses significant challenges. The key question driving our work: "is it feasible to minimize human-in-the-loop feedback by prioritizing data instances which most effectively distinguish between models?" We evaluate several metric-based methods and find that these metrics enhance the efficiency of human evaluations by minimizing the number of required annotations, thus saving time and cost, while ensuring a robust performance evaluation. We show that our method is effective across widely used model families, reducing instances of indecisive (or "tie") outcomes by up to 54% compared to a random sample when focusing on the top-20 percentile of prioritized instances. This potential reduction in required human effort positions our approach as a valuable strategy in future large language model evaluations.

Evaluation Efficiency Language Generative Models Scholars

Related works

Research SimMerge: Learning to Select Merge Operators from Similarity Signals

Read

Research Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation

Read

Research When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs

Read

Notability

notability 6.0/10

Substantive research post on LLM evaluation prioritization.

Cohere has a writing signal matching data demand, evals and quality.