Client project · Jun 2024 to Jan 2025

Multi-source data quality

A seven-month contract to collect, clean, and standardize data from three online platforms.

3
source platforms
140,000+
analysis-ready records
7 months
client engagement
The challenge

Turn three different sources into one dataset that could be trusted.

Each platform used different fields, labels, currencies, and rate formats. The client needed one consistent dataset that could support analysis and future updates.

What I did

Collected, standardized, checked, and documented the data.

  • Collected and prepared data from three platforms using Python.
  • Standardized job roles, experience levels, currencies, and rate types.
  • Found and fixed a serious duplicate-data problem, then added checks to reduce the risk of recurrence.
  • Managed requirements, revisions, delivery notes, and the final documentation directly with the client.
The method

Make the sources comparable, then validate the handoff.

I aligned fields and labels across the three sources, checked for repeated records, and documented the output. The checks mattered because a consistent-looking dataset could still contain errors.

Data preparation process: collect three sources, standardize fields, check repeated records, and document the handoff.
An illustration of the data preparation method. Client records and delivery files remain confidential. Open the image for a larger view.
Result

The client received one consistent dataset with more than 140,000 analysis-ready records and a documented workflow for future updates.

The client later used the delivered data foundation to build a freelance pricing model.

Tools

Python, pandas, Scrapy, Playwright, and Beautiful Soup.

Confidential

Raw datasets, source URLs, delivery files, and supporting documentation are confidential.

Contact

Interested in this kind of work?

Copy my email address, then send a note about the role, project, or reporting problem you are working on.

Copy email