Component Accuracy Does Not Compose: An Exact Account of What a Staged Discovery Agent Actually Yields
Abstract
An agentic discovery system is a funnel: a generator proposes, a retrieval or reasoning step ranks, a cheap in-silico model filters, an assay confirms. Each part is reported with its own number - a hit rate, an AUROC, a rank correlation - and the system's promise is read off from those numbers as if they multiplied. They do not. We compute exactly what such a funnel yields, using the classical theory of multi-stage selection under a truncated multivariate normal, and four things follow that component reporting hides. First, component errors are shared in practice - the same embedding, the same force field, the same literature prior - and assuming them independent overstates the end-to-end yield in 118 of 135 configurations, by a median 5.5% and by up to 50%, with a mild shared-error correlation of 0.3 already worth 13.3%. In the remaining 17 it understates, and the two cases are separated by a rule: correlated errors help only when a much weaker stage runs first, which is its own warning about reading component numbers in isolation. Second, and more usefully: the answer to "which component should we improve?" is decided by the funnel's geometry, not by which component is worse. Over 400 pairs of component accuracies, improving the stage that makes the deeper cut is right 84% of the time while improving the less accurate stage - the natural reading of component metrics - is right 48%, no better than a coin flip, and at worst forgoes 102% of the improvement available. Where the first stage is loose the later stage wins in 240 of 240 cases regardless of the two accuracies. Third, the funnel structure itself has a price: against a single joint decision on the same scores at the same total selection rate, staging costs a median 4.5% and up to 37% of the gain, and the stage order always matters (180 of 180 configurations, by up to 165%). Fourth, a small but exact negative: in 9 of 60 sweeps the end-to-end yield is maximised at an interior component accuracy, so making a component better makes the system worse - by at most 2.5%, which we report as the small effect it is. The computation is exact to 6.7e-16 against the closed-form single-stage gain and agrees with a 200,000-candidate simulated library.