Source platforms
Structurally different inputs
A seven-month paid engagement covering collection, standardization, quality recovery, validation, revisions, and client handoff.
Standardized three source structures into 163,229 consolidated records, investigated a material duplication issue, and delivered 141,228 analysis-ready records with controls and a documented handover.
View résuméThe work required repeatable collection and a common schema without hiding source differences. A later quality review exposed a 43.6% duplication issue in one 139,554-row output, requiring transparent disclosure, root-cause review, recovery, renewed validation, and stronger recurrence controls.
Source A, B, and C entered one repeatable Python workflow, but their identities were retained until validation. The quality incident was disclosed and recovered rather than hidden inside a consolidated total.
Create a common schema while retaining source lineage.
Normalize currencies, role categories, experience levels, and rate types before comparison.
Disclose and quantify the quality incident rather than hide it.
Rebuild around validated unique records and add recurrence controls.
Document limitations and handoff requirements for downstream use.
Structurally different inputs
Accounting total, not unique
After cleaning and alignment
Across seven months
Detect → disclose → investigate → rebuild → validate → add recurrence controls.
Python, pandas, Scrapy, Playwright, and Beautiful Soup supported collection and preparation. The workflow was 60–75% standardized or automated; a full refresh moved from roughly five days to one or two, and manual cross-platform comparison reduced an estimated 50–70%.
The client built a pricing GPT using Isaac’s delivered data foundation; Isaac did not build or own the GPT. A seven-page technical handover documented workflow, outputs, caveats, limitations, and use boundaries.
Seven-month paid client engagement. The client and raw delivery materials remain private; the case uses approved aggregate outcomes and original diagrams. Consolidated records are not presented as unique records.
Client identity, raw records, restricted source names, URLs, correspondence, private delivery files, payment detail, and adapted client screenshots are omitted.