Trust
Data and accuracy
What TenderReader knows today, how extraction is tested, and what is still deliberately unpublished.
Source
| Source | Count | Freshness |
|---|---|---|
| Find a Tender notices | 859 | 2026-07-21 |
| BOAMP notices | 14149 | 2026-07-17 |
| Contracts Finder awards | 4130 | 2026-07-18 |
Dated evaluation results
- Corpus
- broad-real-public-attachments
- Engine
- claude-cli
- Model
- claude-opus-4-8
- Generated
- 2026-06-15
- Measurement scope
- Measured on a local real-public sample snapshot; production and customer-pack accuracy are Not measured.
- Cases
- 12
- Fields
- 40/50
- Overall accuracy
- Not measured
- Labels
- provisional_pending_human_review
- Proof tier
- Locally tested
- Measured in local or repository fixtures; not customer-proven production accuracy.
- Human-confirmed labels
- 0
- Source artifact
- .planning/master-plan/2026-06-14_2032/AI_EVAL_RUN_FINAL.md
- Citation faithfulness
- 99.1%
- Observed source-span support.
- Unsupported-claim rate
- 0.9%
- Must stay below the public gate.
- Critical false-negative rate
- 45.5%
- Packs with a missed critical field.
- Calibration ECE
- 0.193
- Observed emitted fields only.
- Schema fail rate
- 0.0%
- Invalid JSON/schema outputs.
| Field class | Correct | Accuracy | Recall | Bar | Status |
|---|---|---|---|---|---|
| Deadlines | 19/26 | 84.4% | 73.1% | 82.4% | Local sample metric: clears snapshot gate; production not measured |
| Values | 7/8 | 93.3% | 87.5% | 91.3% | Local sample metric: clears snapshot gate; production not measured |
| Eligibility | 9/11 | 69.2% | 81.8% | 67.2% | Local sample metric: clears snapshot gate; production not measured |
| Lots | 5/5 | 100.0% | 100.0% | 98.0% | Local sample metric: clears snapshot gate; production not measured |
Overall production accuracy: Not measured. Language split: not yet measured in the real snapshot; labels are provisional and still need human review.
True production accuracy is Not measured. These are provisional count-based real-eval field metrics from public tender attachments. Development validation (local Claude CLI, not production): submission-deadline extraction scored 473/473 = 100% on a 560-pack real UK+FR public corpus (95% CI lower bound 0.992, clearing the >=99% deadlines-sacred bar); contract value 96.3%. Production accuracy remains Not measured pending a production provider path.
Provenance
Notices and awards come from public feeds: Find a Tender, Contracts Finder, and BOAMP. Product surfaces keep the source name and the OGL/Etalab attribution with the data.
Evaluation method
This page reads the versioned real snapshot committed in the repo: public tender attachments, per-class metrics, citation faithfulness, calibration, and critical failures.
The labels remain provisional. They are not a public accuracy claim for customer packs or a production model provider.
Accuracy numbers
The numbers below come from the real snapshot JSON. Unmeasured dimensions stay rendered as Not measured.
The gates protect regressions, but they do not turn provisional metrics into outcome guarantees.
Data handling boundaries
Customer packs, analyses, profiles, and preferences are kept while the account is active. Preferences now includes JSON export and a deletion request; billing audit records may be retained where legally required.
Uploaded documents stay inside the authenticated account workspace. A configured extraction provider may receive documents only for processing; the public metrics above come from separate evaluation snapshots, not customer packs.
- No public accuracy metric is calculated from customer packs.
- Deletion is request/review based, not an instant destructive action.
- No SOC 2, ISO 27001, or independent-audit certification is claimed here.