STRATA: Automated Construction of Professional Tasks and Rubrics for Agent Evaluation
Fenil Bardoliya ⋅ Shivali Dalmia ⋅ Abhishek Mukherji
Abstract
Evaluating Agentic AI systems on professional tasks is bounded not by model access but by the cost of building the benchmark. Benchmarks with professional deliverables are handcrafted: experts source each task with gold deliverables and curate rubrics for grading them. GDPval, for instance, spans 9 domains, 44 occupations, and 1320 tasks (30 per occupation, only 220 are open-sourced). We present STRATA: Scalable Tiered Rubrics And TAsks construction, an autonomous, multi-agentic system for constructing professional benchmarks for a given domain. Phase one, TaskGEN curates a grounded task along with publicly-sourced references and deliverables with a mere brief of domain and occupation. Phase two, RubricGEN crafts three-tiered rubrics: Domain, Occupation, Task for the evaluation. Gating enforces the trust: no item is tuned or generated against the same eval model, every factual value is verbatim from a verified public-source, fixed config set, and any record failing a gate after repairs is discarded. Outputs surviving those gates are inspectable - making scaling an integral property of the system. On the Financial Domain, STRATA produces 102 records across seven occupations at a median construction cost of \\$11.80 and 32 minutes per record. Two Subject Matter Experts rate the task rigor 4.5, 5.0; representativeness 4.0 and 4.5; appropriateness 4.0 and 3.7 and agree within one-point on 72\% of criteria (Cohen's $\kappa=0.50$). The automated construction with Human-in-the-Loop at scale brings the coverage within reach.
Chat is not available.
Successful Page Load