← All case studies

Health, Safety & Environmental Compliance

Classifying More Than 100,000 Health, Safety and Environmental Records for a Regulatory Handover

An engineering contractor on a power-generation megaproject used Docwize to classify a decade-plus archive of more than 100,000 health, safety and environmental documents into a controlled taxonomy of more than 200 types and subtypes, preparing the record for formal handover to the facility owner under occupational health and safety law.

100,000+ documents classified

Health, safety and environmental records processed through the extraction pipeline in the first phase.

200+ types and subtypes in a controlled taxonomy

Maintained as Docwize lookup data and injected into the extraction step at run time.

70%+ reduction in the manual review queue

Achieved through nine or more prompt and taxonomy revisions, each measured against a quality dashboard.

Close to 600,000 documents identified programme-wide

Inventory reconciled by content hash, with page counts recorded before extraction.

The challenge

The contractor held health and safety and environmental documentation spanning more than a decade of construction and operational activity. As the project moved toward a legally required handover of that record to the facility owner, the archive told a very different story from a clean, audit-ready set: more than 100,000 files sat in an unstructured folder hierarchy with almost no usable metadata. Fewer than one document in four hundred carried a recorded document type, and none carried a facility reference. Close to 600,000 documents had been identified across the wider two-facility programme. Manually reviewing and classifying a set of that size against a handover deadline was not a viable option.

The approach

Docwize built a purpose-specific extraction pipeline rather than applying generic OCR and tagging. The taxonomy and the contractor list lived as lookup grids in Docwize and were injected into the extraction step at run time, so refining the classification scheme never meant changing code. Every run was measured, and the pipeline was iterated until the manual review queue converged.

  • Reconcile the file inventory against the import, identifying duplicates by file content hash and recording page counts so that only unique files entered the pipeline
  • Build a controlled taxonomy of more than 200 document type and subtype categories mapped to the handover sections, and a canonical list of more than 100 contractor names with aliases
  • Render each document's first page as an image, pair it with its OCR text, and run AI extraction against the taxonomy to capture discipline, type, subtype, date and its source, reference number, revision, description, contractor and author — with multi-document bundles flagged for splitting
  • Validate every result against a structured schema with automatic retry, record a confidence level per field and a review reason where needed, and route low-confidence results to a manual review queue in a dataview
  • Track review-queue size, confidence distribution and error patterns per run in a quality dashboard, refine the extraction prompt and taxonomy across nine or more revisions, and write validated metadata back to each document's custom fields for filing by handover section

What the analysis revealed

Metadata was effectively absent

With fewer than one in four hundred documents typed and none tagged to a facility, the folder structure offered no reliable way to confirm what existed or what it was. Manual triage alone could not have made the archive handover-ready in time.

Iteration measurably improved results

Across successive production batches, the share of documents flagged for manual review fell by more than 70% as the extraction prompt and taxonomy were refined — turning an open-ended AI exercise into a converging, measured process.

Taxonomy discipline mattered as much as the model

At close to 600,000 documents across the programme, the controlled taxonomy and canonical contractor list did as much to keep classification consistent as the extraction technology itself.

Outcome

More than 100,000 documents were classified in the first phase, each carrying validated document type, subtype, date, reference, contractor and author, and filed by handover section. Duplicates were identified by content rather than file name, and bundled scans were flagged for splitting. The contractor moved from a folder structure with no reliable way to say what it held to a searchable, classified record it can defend at handover — with a continuously improving pipeline in place to carry the remaining volume through to completion.

Facing a regulatory handover or audit of an archive nobody has classified?

Docwize combines AI extraction with controlled taxonomies, measured review queues and structured metadata written back to your documents. Get in touch to discuss your archive.