Generalist Terminal-Use Agents for Event Extraction in Aviation Accident Investigations
Abstract
Generalist terminal-use agents can operate over heterogeneous aviation investigation evidence, but converting that evidence into standardized safety codes remains difficult. We evaluate two model backbones across three harnesses on 134 NTSB cases, requiring agents to produce occurrence-code and phase-of-flight-code pairs from raw evidence without a fixed preprocessing pipeline. The strongest system reaches 40.54\% tuple micro-F1. The gold-narrative ablation shows that substantial taxonomy-coding difficulty remains even when evidence extraction is simplified. Behavioral profiling reveals substantial variation in how agents explore and process the same investigation evidence, but these differences are difficult to interpret in terms of coding performance and do not consistently correspond to higher scores. These findings present aviation accident records as a realistic testbed for studying agents on complex, safety-critical tasks.