Study kit proposed
Can a claims professional reconstruct the decision?
Insurers cite opacity as a reason to exclude AI. This study tests the narrowest claim that answers it: that claims professionals reconstruct an automated denial faster and more accurately from a receipt than from ordinary logs. It has not been run. If receipts do not help, that result is published and the claim is withdrawn.
Design
- Three synthetic insurance cases, each in two conditions: plain logs or receipt packet.
- Each participant sees every case once, counterbalanced across conditions.
- Primary measure: accuracy on eight reconstruction questions, with partial credit for each required fact named. Secondary: time. Confidence is reported separately.
- Analysis, stopping rule and publication commitment are fixed before recruitment.
Reconstruction questions
- What action was taken or proposed, on what object?
- Under whose authority, and does that authority trace to someone other than the system itself?
- Which evidence was relied on, and was it about this claim?
- Was the evidence current when the decision took effect?
- Did review complete before the decision was committed?
- Which claimant routes (review, correction, appeal) remained available afterwards?
- Did the decision take effect before it was warranted?
- What would have had to change for the decision to be warranted?
Materials
Protocol · Preregistration fields (draft, not frozen) · Scoring sheet · Answer key
| Case | Log condition | Receipt condition |
|---|---|---|
| INS-001 Denial after completed human review, appeal routes kept | INS-001-log.txt | |
| INS-002 Denial committed in minutes, before review, appeal route closed | INS-002-log.txt | |
| INS-003 Denial already sent, authority self-asserted | INS-003-log.txt |
Receipt packets are computed in your browser with a fixed record time, so every participant receives identical bytes. The answer key is generated from the gate's checks, so it cannot drift from what a receipt shows; its SHA-256 is recorded in the preregistration. It has not yet been reviewed independently, which the protocol requires before any run.