Skip to yearly menu bar Skip to main content


When the Grader Is Fooled: Measuring How Often LLM Judges Accept Plausible-Wrong Answers Across Six Model Families

Veerendra Kumar Sunkavalli

Abstract

Chat is not available.