A comprehensive survey of methods for evaluating abstractive summaries and detecting hallucinations. Covers four evaluation dimensions (fluency, coherence, relevance, consistency), then dives into reference-based metrics (ROUGE, METEOR, BERTScore, MoverScore), context-based reference-free metrics (ROUGE-C, G-Eval with GPT-4), entailment-based hallucination detection (SummaC, TrueTeacher), QA-based consistency metrics (QuestEval, QAFactEval), preference-based reward models, and sampling-based approaches (SelfCheckGPT). Key findings include that modern LLMs often outperform gold reference summaries, SOTA inconsistency detection achieves only 60–75% balanced accuracy, and a pragmatic iterative evaluation approach is recommended for production use.

14m read timeFrom eugeneyan.com
Post cover image
Table of contents
Fluency, Coherence, Relevance, ConsistencyReference-based metricsContext-based (reference-free) metricsPreference-based metricsSampling-based metricsReferences
3 Impressions