Is Generative Content Ready to Ship? A Production-Scale Playbook
Abstract
LLMs make it easy to prototype generative product experiences, but not to decide whether they are ready to ship. The bottleneck has shifted from generation to evaluation. In a large-scale lodging marketplace, we wanted to use query-aware listing summaries to reduce search effort. The prototype came together quickly, but figuring out how to validate it was a different story. Each observed shortcoming suggested another LLM judge, and the resulting pile of metrics gave little guidance about whether an iteration had actually made the product better overall. That hurdle is our starting point. We needed a practical way to evaluate. Treating the query, listing, and summary as three overlapping information sets yields pre- cision and recall. When those saturate, a pairwise utility judge rewards marginal information gain; when utility pushes summaries too long or too complex, length and readability become launch gates. We calibrate rubric-based LLM judges against expert labels. To compare iterations, we track how far each metric is from its launch target, avoiding hand-tuned weights between metrics. Offline success still does not establish user impact. User-level A/B tests mea- sure behavior, but traffic is scarce and conventional tests run for weeks and still have high minimum detectable effect (MDE). We therefore adapt interleaving to generated content on repeated product surfaces: randomize the content variant at the deterministic (user,item) level while holding ranking and layout fixed, so a session can expose both variants. In production, interleaving results reach significance for top-of-funnel metrics on a single day of data, while preserving booking guardrails and surfacing segment- level effects. A severity judge completes our evaluation loop by separating minor imprecisions from significant hallucinations.