Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

From Calculation to Adjudication: Examining LLM Judges on Mathematical Reasoning Tasks

  • Andreas Stephan
  • , Dawei Zhu
  • , Matthias Aßenmacher
  • , Xiaoyu Shen
  • , Benjamin Roth

Veröffentlichungen: Beitrag in BuchBeitrag in KonferenzbandPeer Reviewed

Abstract

To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typically evaluated by measuring the correlation with human judgments on generative tasks such as summarization or machine translation. In contrast, we study LLM judges on mathematical reasoning tasks. These tasks require multi-step reasoning, and the correctness of their solutions is verifiable, enabling a more objective evaluation. We perform a detailed performance analysis and find that easy samples are easy to judge, and difficult samples are difficult to judge. Our analysis uncovers a strong correlation between judgment performance and the candidate model task performance, indicating that judges tend to favor higher-quality models even if their answer is incorrect. As a consequence, we test whether we can predict the behavior of LLM judges using simple features such as part-of-speech tags and find that we can correctly predict 70%-75% of judgments. We conclude this study by analyzing practical use cases, showing that LLM judges consistently detect the on-average better model but largely fail if we use them to improve task performance.
OriginalspracheEnglisch
TitelThe 63rd Annual Meeting of the Association for Computational Linguistics
UntertitelProceedings of the GEM² Workshop
ErscheinungsortKerrville
VerlagACL Anthology
Seiten759-773
ISBN (Print)979-8-89176-261-9
PublikationsstatusVeröffentlicht - Juli 2025
Veranstaltung
The 63rd Annual Meeting of the Association for Computational Linguistics
-
Dauer: 27 Juli 20251 Aug. 2025

Konferenz

Konferenz
The 63rd Annual Meeting of the Association for Computational Linguistics
Zeitraum27/07/251/08/25

ÖFOS 2012

  • 102001 Artificial Intelligence

Zitationsweisen