Multi-source data quality
A seven-month contract to collect, clean, and standardize data from three online platforms.
- 3
- source platforms
- 140,000+
- analysis-ready records
- 7 months
- client engagement
Turn three different sources into one dataset that could be trusted.
Each platform used different fields, labels, currencies, and rate formats. The client needed one consistent dataset that could support analysis and future updates.
Collected, standardized, checked, and documented the data.
- Collected and prepared data from three platforms using Python.
- Standardized job roles, experience levels, currencies, and rate types.
- Found and fixed a serious duplicate-data problem, then added checks to reduce the risk of recurrence.
- Managed requirements, revisions, delivery notes, and the final documentation directly with the client.
Make the sources comparable, then validate the handoff.
I aligned fields and labels across the three sources, checked for repeated records, and documented the output. The checks mattered because a consistent-looking dataset could still contain errors.
The client received one consistent dataset with more than 140,000 analysis-ready records and a documented workflow for future updates.
The client later used the delivered data foundation to build a freelance pricing model.
Python, pandas, Scrapy, Playwright, and Beautiful Soup.
Raw datasets, source URLs, delivery files, and supporting documentation are confidential.
Interested in this kind of work?
Copy my email address, then send a note about the role, project, or reporting problem you are working on.
Copy email