BUILD RECORD
PagePurger
Medical-legal review is billed by the page. The money is in the pages nobody has to read. A specialized pipeline that flags and extracts only the non-routine pages in massive discovery.
OVERVIEW
Fewer pages to review, and the reason each removed page came out.
Retired
A production pipeline for medical-legal record review. Every page got one relevance call, and fingerprints caught the duplicates.
Every duplicate and unrelated page pulled out was $3 nobody billed. Each came back watermarked, its reason on a spreadsheet row.
8,000+ pages in one case file, read by hand at about $3 each.
aws batch gpu bedrock perceptual hashing sftp
Architecture, pipeline, classification, multi-tenant platform, operations.
Each tenant ran separately; review sets went back over SFTP.
THE METHOD
How a case file became a review set
It pulled blank separator sheets, the same fax scanned four times, and rescans a hair off.
Took the case file
Records arrived over SFTP into the tenant's pipeline.
Judged every page against the case
Each page got one relevance call, Claude Haiku on Bedrock.
Collapsed the duplicates
A file fingerprint found exact copies. Perceptual hashing found near-identical rescans.
Handed back the review set
It returned review pages, watermarked exclusions, the decision spreadsheet, and the invoice.
LIMITS STAY ATTACHED
What it refused to do
Nothing was deleted. Every excluded page came back watermarked with its reason, for the lawyer who needs it later.
A long job checkpointed and resumed where it stopped.
Tenants stayed separate. No case or client appears here.
RECORD
Where it went
Retired. Built and run solo.
aiPDF Engine, the same pipeline without the case rules.
NEXT
THE ARCHITECTURE
A RELATED PROBLEM
Is page-by-page review setting the cost of the work?
This build is retired, but the review problem remains. If a system can remove obvious waste and show the reason for every removal, it is worth a conversation.