Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles

Veröffentlichungen: Beitrag in BuchBeitrag in KonferenzbandPeer Reviewed

Abstract

Calibration, the alignment between model confidence and prediction accuracy, is critical for the reliable deployment of large language models (LLMs). Existing works neglect to measure the generalization of their methods to other prompt styles and different sizes of LLMs. To address this, we define a controlled experimental setting covering 12 LLMs and four prompt styles. We additionally investigate if incorporating the response agreement of multiple LLMs and an appropriate loss function can improve calibration performance. Concretely, we build Calib-n, a novel framework that trains an auxiliary model for confidence estimation that aggregates responses from multiple LLMs to capture inter-model agreement. To optimize calibration, we integrate focal and AUC surrogate losses alongside binary cross-entropy. Experiments across four datasets demonstrate that both response agreement and focal loss improve calibration from baselines. We find that few-shot prompts are the most effective for auxiliary model-based methods, and auxiliary models demonstrate robust calibration performance across accuracy variations, outperforming LLMs’ internal probabilities and verbalized confidences. These insights deepen the understanding of influence factors in LLM calibration, supporting their reliable deployment in diverse applications.
OriginalspracheEnglisch
TitelProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Redakteure*innenWanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
ErscheinungsortKerrville
VerlagACL Anthology
Seiten3740 - 3761
Seitenumfang22
ISBN (elektronisch)9798891762510
ISBN (Print)979-8-89176-251-0
DOIs
PublikationsstatusVeröffentlicht - Juli 2025
Veranstaltung
The 63rd Annual Meeting of the Association for Computational Linguistics
-
Dauer: 27 Juli 20251 Aug. 2025

Konferenz

Konferenz
The 63rd Annual Meeting of the Association for Computational Linguistics
Zeitraum27/07/251/08/25

Fördermittel

This research has been funded by the Vienna Science and Technology Fund (WWTF)[10.47379/VRG19008] “Knowledge infused Deep Learning for Natural Language Processing”, and co-funded by the European Union.

ÖFOS 2012

  • 102001 Artificial Intelligence

Fingerprint

Untersuchen Sie die Forschungsthemen von „Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles“. Zusammen bilden sie einen einzigartigen Fingerprint.

Zitationsweisen