Model Comparison & Benchmarking
AI oversight control room for evaluation rigor, supervisory intelligence, governance, and decision-quality impact
Model Comparison Proof
Refreshed: 2026-07-20 13:46:22 UTC
SOURCE
Eval Runs + Agent Executions (POSTGRES)
Scope: Domain:finance Mission:all Mode:all Days:30
POLICY
Policy violations and enforcement outcomes per model
EVIDENCE
Execution hashes + evidence references
RECORDS
Executions:1999 Models:3
ACTION
Benchmark, compare, and promote model candidates
Total Pending
1821
Processing Volume
1999
Avg Decision Latency
540ms
Policy Blocks
0
Replay Integrity
8.70%
Estimated Token Cost
$0.00
Model Leaderboard
| Model | Version | Eval Score | Decision Quality Index | Latency | Cost per 1K Tokens | Hallucination Rate | Policy Violations | Confidence | Drift Status | Mode | Status | Inspect |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deterministic_scenario_engine | p2p-runtime-v1 | 100.00 | 1.000 | 12517ms | $0.0000 | 0.00% | 0 | High | Improving | Live | Shadow | |
| unattributed | latest | 100.00 | 1.000 | 0ms | $0.0000 | 0.00% | 0 | High | Warning | Live | Risk | |
| rule-engine | finance-agent-base-v1 | 8.72 | 0.087 | 535ms | $0.0000 | 63.90% | 0 | Low | Warning | Live | Risk |
Score Breakdown
Selected model: deterministic_scenario_engine
| Accuracy | 100.00 | |
| Policy Compliance | 100.00 | |
| Determinism | 100.00 | |
| Hallucination Suppression | 100.00 | |
| Evidence Trace Completeness | 75.00 | |
| Latency Stability | 0.00 | |
| Cost Efficiency | 100.00 |
Drift & Stability Monitor
Output entropy trend
0.00
Confidence variance proxy
0.00
Policy violation trend
0
Prompt distribution shift
1999
Token usage
0
Anomaly indicator
Stable
Side-by-Side Output Comparison
| Prompt | Model A Output | Model B Output | Ground Truth | Policy Flags | Evidence Hash |
|---|---|---|---|---|---|
| Insufficient paired model outputs for this filter context | |||||
Shadow vs Live Impact
Decisions Changed
1999
Quality Delta
-8.90%
Financial Delta
$0.00
Risk Delta
0
Exceptions Prevented (proxy)
178
Value Attribution Impact
$0.00
Cost Governance
Monthly Token Spend
$0.00
Cost per Decision
$0.0000
ROI Multiple (proxy)
0.00x
Cost vs Quality Frontier
Efficient
Autonomy Readiness Indicator
| Model | AL-1 | AL-2 | AL-3 | AL-4 |
|---|---|---|---|---|
| deterministic_scenario_engine | ✔ | ✔ | Conditional | ✖ |
| unattributed | ✔ | ✔ | Conditional | ✖ |
| rule-engine | ✔ | Conditional | Conditional | ✖ |