HireWatch: Evaluating LLM Compliance with U.S. Employment Law Under Contextual Pressure
Abstract
The deployment of large language models (LLMs) in hiring pipelines is growing, yet existing evaluations generally overlook an important aspect of this development: the compliance of LLMs with applicable legal rules. To fill this gap, we introduce HireWatch, a benchmark that assesses whether LLM hiring assistants comply with U.S. employment discrimination law (e.g., the Civil Rights Act of 1964). We construct a 200-item bank consisting of interview questions that are either lawful or unlawful under U.S. federal statute, regulation, and enforcement guidance. We evaluate the propensity of 18 proprietary and open-source LLMs to select lawful vs. unlawful questions from this bank across ~100,000 trials under a novel seven-tier contextual escalation ladder, crossing occupational details with organizational pressure. Our findings reveal that LLMs' compliance with employment discrimination law is brittle. While LLMs ordinarily comply with legal rules under different job contexts, they readily violate the law when subject to institutional pressure (e.g., an unlawful manager instruction or company policy). Meanwhile, although the inclusion of explicit legal reminders in prompts can improve legal compliance, it does not altogether eliminate violations of employment discrimination law.