TenderReader
Menu

Founding Pro: £39/mo for 12 months. 0/50 founding Pro places claimed. 50 remain.

See the offer

Source

Indexed public sources and freshness
Source Count Freshness
Find a Tender notices 859 2026-07-21
BOAMP notices 14149 2026-07-17
Contracts Finder awards 4130 2026-07-18

Dated evaluation results

Corpus
broad-real-public-attachments
Engine
claude-cli
Model
claude-opus-4-8
Generated
2026-06-15
Measurement scope
Measured on a local real-public sample snapshot; production and customer-pack accuracy are Not measured.
Cases
12
Fields
40/50
Overall accuracy
Not measured
Labels
provisional_pending_human_review
Proof tier
Locally tested
Measured in local or repository fixtures; not customer-proven production accuracy.
Human-confirmed labels
0
Source artifact
.planning/master-plan/2026-06-14_2032/AI_EVAL_RUN_FINAL.md
Citation faithfulness
99.1%
Observed source-span support.
Unsupported-claim rate
0.9%
Must stay below the public gate.
Critical false-negative rate
45.5%
Packs with a missed critical field.
Calibration ECE
0.193
Observed emitted fields only.
Schema fail rate
0.0%
Invalid JSON/schema outputs.
Results by field class
Field class Correct Accuracy Recall Bar Status
Deadlines 19/26 84.4% 73.1% 82.4% Local sample metric: clears snapshot gate; production not measured
Values 7/8 93.3% 87.5% 91.3% Local sample metric: clears snapshot gate; production not measured
Eligibility 9/11 69.2% 81.8% 67.2% Local sample metric: clears snapshot gate; production not measured
Lots 5/5 100.0% 100.0% 98.0% Local sample metric: clears snapshot gate; production not measured

Overall production accuracy: Not measured. Language split: not yet measured in the real snapshot; labels are provisional and still need human review.

True production accuracy is Not measured. These are provisional count-based real-eval field metrics from public tender attachments. Development validation (local Claude CLI, not production): submission-deadline extraction scored 473/473 = 100% on a 560-pack real UK+FR public corpus (95% CI lower bound 0.992, clearing the >=99% deadlines-sacred bar); contract value 96.3%. Production accuracy remains Not measured pending a production provider path.

Provenance

Notices and awards come from public feeds: Find a Tender, Contracts Finder, and BOAMP. Product surfaces keep the source name and the OGL/Etalab attribution with the data.

Evaluation method

This page reads the versioned real snapshot committed in the repo: public tender attachments, per-class metrics, citation faithfulness, calibration, and critical failures.

The labels remain provisional. They are not a public accuracy claim for customer packs or a production model provider.

Accuracy numbers

The numbers below come from the real snapshot JSON. Unmeasured dimensions stay rendered as Not measured.

The gates protect regressions, but they do not turn provisional metrics into outcome guarantees.

Data handling boundaries

Customer packs, analyses, profiles, and preferences are kept while the account is active. Preferences now includes JSON export and a deletion request; billing audit records may be retained where legally required.

Uploaded documents stay inside the authenticated account workspace. A configured extraction provider may receive documents only for processing; the public metrics above come from separate evaluation snapshots, not customer packs.

  • No public accuracy metric is calculated from customer packs.
  • Deletion is request/review based, not an instant destructive action.
  • No SOC 2, ISO 27001, or independent-audit certification is claimed here.