Skip to content

Validation results ​

The tagger was developed iteratively against a stable oldest-first batch. These two reviewed Batch 1 snapshots record the point at which the classification and evidence methodology was accepted for application. They document this batch's automation outcome, not a universal model accuracy rate.

Batch 1 ​

Batch 1 — 20260927T030125Z ​

Batch 1 — 20260927T030125Z records 20 selected and validated bookmarks: 19 high-confidence results, one medium-confidence result, no failed classifications, and no blocked high-confidence results.

The medium-confidence result had no preserved page content, and controlled direct retrieval received HTTP 403. The remaining metadata supported a provisional recommendation for the STEM / Infrastructure / Networking collection, but the evidence limitation prevented full verification.

Batch 1 — 20260927T031250Z ​

Batch 1 — 20260927T031250Z records 20 selected bookmarks: 19 applied results, one skipped medium-confidence result, no blocked high-confidence results, and no failed classifications or writes.

The evidence-limited bookmark remained unchanged for manual handling. The later Batch 1 snapshot records fresh inference rather than replayed results from the earlier snapshot.

Demonstrated behavior ​

The batch demonstrated that high-confidence results based on complete one-pass, complete multi-pass, and controlled direct-fetch evidence can participate in normal apply. The per-request context limit did not require silent document truncation because multi-pass processing covered oversized documents.

Degraded evidence still produced a useful provisional classification without forcing high confidence. Explicit apply mutated the 19 applicable high-confidence results, while the remaining exception stayed untouched for manual handling. No classification or write failed in Batch 1.

Built with VitePress.