SarcBench: A Bilingual Benchmark for Contextual Sarcasm Understanding, Response, and Generation
Abstract
Sarcasm understanding and response are important for large language models (LLMs) to engage naturally in real-world conversations. However, existing sarcasm benchmarks primarily focus on detection or related subtasks in isolation, lacking diagnostic power and failing to reflect the complexity of real conversational interactions. To address this gap, we introduce SarcBench, a bilingual benchmark with 30,083 aligned samples. It defines three interconnected tasks, Intent Recognition, Sarcasm Response, and Sarcasm Generation, enabling diagnostic evaluation of sarcasm understanding and conversational behavior. We evaluate 12 contemporary language models on SarcBench and find that even the top-performing model achieves only 61.21%. We also identify a clear understanding–action gap: models are consistently better at recovering intended meanings than at producing socially appropriate responses or generating sarcasm that faithfully realizes a specified rhetorical mechanism. In addition, performance varies substantially across models and task types, and controlled sarcasm generation often collapses rhetorical diversity into a narrow set of default templates. Taken together, these findings indicate that current models remain limited in handling sarcasm as an interactive social behavior rather than merely a recognition task.