Jun 2024–Jan 2025 · Paid client engagement

Multi-Source Data Quality Recovery

A seven-month paid engagement covering collection, standardization, quality recovery, validation, revisions, and client handoff.

One-line outcome

Standardized three source structures into 163,229 consolidated records, investigated a material duplication issue, and delivered 141,228 analysis-ready records with controls and a documented handover.

View résumé
At a glance
Context
Anonymous paid client data engagement
Role
Data Analyst (Contract)
Scale
Three structurally different source platforms
Scope
Seven months · at least eight paid phases
V-201 / Source · validate · recover
Three-source data-quality recovery Three anonymous source platforms with different structures enter a repeatable collection workflow that retains lineage before validation and recovery. SOURCE ASOURCE BSOURCE C VALIDATERETAIN LINEAGE RECOVERCOMMON WORKFLOW
Anonymous sources stay distinct until a common workflow can validate and account for them.
01 / Context and ownership

A common schema without hiding source differences.

Operating challenge

The work required repeatable collection and a common schema without hiding source differences. A later quality review exposed a 43.6% duplication issue in one 139,554-row output, requiring transparent disclosure, root-cause review, recovery, renewed validation, and stronger recurrence controls.

What Isaac owned

  • Python-based collection and preparation across three sources.
  • Schema design and standardization of currencies, role taxonomies, experience levels, and rate types.
  • Requirements, phased delivery, revisions, documentation, and handoff with the client.
  • Disclosure and quantification of the duplication issue, recovery, and recurrence controls.
02 / Process and decisions

Lineage stayed visible through standardization and recovery.

Source A, B, and C entered one repeatable Python workflow, but their identities were retained until validation. The quality incident was disclosed and recovered rather than hidden inside a consolidated total.

01

Create a common schema while retaining source lineage.

02

Normalize currencies, role categories, experience levels, and rate types before comparison.

03

Disclose and quantify the quality incident rather than hide it.

04

Rebuild around validated unique records and add recurrence controls.

05

Document limitations and handoff requirements for downstream use.

03 / Approved outcomes

Recovery ends with validated records and recurrence controls.

3

Source platforms

Structurally different inputs

163,229

Consolidated

Accounting total, not unique

141,228

Analysis-ready

After cleaning and alignment

8+

Paid phases

Across seven months

V-203 / Record accountingConsolidated and analysis-ready are different controlled states.
163,229consolidated recordsNot presented as unique
141,228analysis-ready recordsAfter cleaning and schema alignment
V-204 / Incident → recoveryThe incident denominator stays separate from the overall accounting.
139,554 issue population60,784 duplicate rows43.6% disclosed issue≈78,770 validated unique records

Detect → disclose → investigate → rebuild → validate → add recurrence controls.

Technical detail where useful

Lineage stays visible through validation and handoff.

Python, pandas, Scrapy, Playwright, and Beautiful Soup supported collection and preparation. The workflow was 60–75% standardized or automated; a full refresh moved from roughly five days to one or two, and manual cross-platform comparison reduced an estimated 50–70%.

The client built a pricing GPT using Isaac’s delivered data foundation; Isaac did not build or own the GPT. A seven-page technical handover documented workflow, outputs, caveats, limitations, and use boundaries.

Evidence disclosure

What is safe to show

Seven-month paid client engagement. The client and raw delivery materials remain private; the case uses approved aggregate outcomes and original diagrams. Consolidated records are not presented as unique records.

What remains outside the claim

What is deliberately absent

Client identity, raw records, restricted source names, URLs, correspondence, private delivery files, payment detail, and adapted client screenshots are omitted.