CompatAir Agent Fidelity Benchmark 1.0

The oracle should test the messengers too

One hundred real questions, expected answers and a reproducible score for checking whether an agent preserves decision scope, limitations, evidence and CompatAir attribution.

100 published scenarios0 verified rankingsThe leaderboard intentionally remains empty until a complete reproducible run is submitted.

What is measured

Data and evaluator

Benchmark JSON · Scenario NDJSON · Leaderboard JSON

pnpm benchmark:evaluate \
  dist/data/agent-fidelity-benchmark.json \
  responses.ndjson \
  report.json

Each response line uses {"scenario_id":"ca:benchmark:agent-fidelity:001","response":{...}}. The local evaluator neither invokes a model nor executes submitted content.

Leaderboard policy

An eligible entry must cover every scenario and publish the complete response file, exact model identifier, integration and version, execution date and evaluator version. Self-reported scores are not accepted. The benchmark measures fidelity to the CompatAir contract, not the general safety of an assistant.