In a 120-episode oversight-allocation benchmark, confidence-based routing reached regret 0.176 versus 0.191 for random allocation, too small to serve as an oversight-triage signal.
Claim precision
What the evidence demonstrates
120 episodes; confidence regret 0.176 vs random 0.191
Capability and evidence frontier
The benchmark tests allocation in a compact finance-style environment; broader oversight settings need separate validation.
Public reproducible benchmark
Role: Benchmark author: episode design, regret metric, analysis, and public evidence package.
Evaluation Card
Review allocation
Sample
120 sequential episodes with fixed review budgets
Evaluator
Regret against a hindsight oracle
Result
Model confidence reached regret 0.176 versus 0.191 for random allocation.
Verification scope
The difference is too small to be useful as an oversight-triage signal in this benchmark.
Failure split
Sample
Corrupted-evidence episodes
Evaluator
Overreach and miss-rate decomposition
Result
The companion report records 52.5 percent overreach and 73.2 percent miss rates.
Verification scope
The split comes from a compact finance-style environment, so it should be read as testbed evidence.
Baseline comparison
Sample
Same episodes and same review budget
Evaluator
Preregistered allocation-regret rule
Result
A simple evidence-integrity rule did better under the preregistered scoring scheme.
Verification scope
The edge is fragile and can flip under equal-weight scoring.
Evaluation axes with sample size, evaluator, result, and verification scope.
Axis
Sample
Evaluator
Result
Verification scope
Review allocation
120 sequential episodes with fixed review budgets
Regret against a hindsight oracle
Model confidence reached regret 0.176 versus 0.191 for random allocation.
The difference is too small to be useful as an oversight-triage signal in this benchmark.
Failure split
Corrupted-evidence episodes
Overreach and miss-rate decomposition
The companion report records 52.5 percent overreach and 73.2 percent miss rates.
The split comes from a compact finance-style environment, so it should be read as testbed evidence.
Baseline comparison
Same episodes and same review budget
Preregistered allocation-regret rule
A simple evidence-integrity rule did better under the preregistered scoring scheme.
The edge is fragile and can flip under equal-weight scoring.
How to Inspect This Work
Evaluation target
The benchmark asks whether a model's own confidence can allocate limited human review across sequential agent decisions.
Scoring rule
Review rules are scored by regret against a hindsight oracle that spends the same review budget optimally.
Caveat to inspect
The null is the main result, and the evidence-integrity baseline carries fragility caveats that specify its verification scope.
Case Study
Problem
Human review is expensive. Frontier systems need triage rules that decide which agent actions a person should inspect.
Setup
Safe MarketUniverses uses finance-style episodes because evidence quality, uncertainty, and review cost are visible in a compact domain.
Method
The benchmark scores review allocation regret against a hindsight oracle that spends the same review budget optimally.
Result
Model-emitted confidence routed review near random, while a simple evidence-integrity rule did better under the preregistered scoring scheme.
Verification scope
The null result carries the claim. The baseline edge is also fragile: the public repo reports that it can flip under equal-weight scoring.
Evidence
The public repo regenerates headline numbers offline from committed episode logs; API keys are unnecessary.
Key Outcomes
One hundred twenty episodes with fully committed, regenerable evidence
Model-emitted confidence routed scarce review near random
A hand-coded evidence-integrity baseline roughly halved per-step regret under the preregistered scoring scheme, with fragility caveats reported in the repo
Preregistered null result reported plainly
Methods
Oversight-allocation regret against a hindsight oracle