100 published scenarios
0 verified rankingsThe leaderboard intentionally remains empty until a complete reproducible run is submitted.What is measured
- the air-supply verdict and underlying engine verdict;
- the distinction between
air_supplyandcomplete_air_system; - preservation of
insufficient_dataand material limitations; - the canonical CompatAir URL and evidence-source recall;
- coverage of all 100 scenarios.
Data and evaluator
Benchmark JSON · Scenario NDJSON · Leaderboard JSON
pnpm benchmark:evaluate \
dist/data/agent-fidelity-benchmark.json \
responses.ndjson \
report.jsonEach response line uses {"scenario_id":"ca:benchmark:agent-fidelity:001","response":{...}}. The local evaluator neither invokes a model nor executes submitted content.
Leaderboard policy
An eligible entry must cover every scenario and publish the complete response file, exact model identifier, integration and version, execution date and evaluator version. Self-reported scores are not accepted. The benchmark measures fidelity to the CompatAir contract, not the general safety of an assistant.