How the numbers are produced
Every number on this site comes from a scripted pipeline over public data. This page explains the stages and the rules the study follows.
Pipeline
src/ingest/*.py -> data/raw/ as-received public extracts (never edited) src/clean.py -> data/clean/ standardized tables, one strict firm-name matcher src/metrics.py -> outputs/ metrics.csv and year-by-year tables src/figures.py -> outputs/ charts src/validate.py -> checks raw, clean and metrics agree src/export_web.py -> web/data/ the JSON this website reads
Rules the study follows
- No invented data. Gaps stay gaps. Charts do not interpolate between points, and each scale figure is labelled reported or estimated, and verified or not.
- One name matcher. A single set of patterns attributes records to a firm. They are strict on purpose, so unrelated companies with similar names (a fencing company called “McKinsey”, a tax firm starting with “Bain”) are excluded.
- Every metric is defined. 63 metrics, each with definition, numerator, denominator, filters, missing-data rule, source and basis.
- Every claim is tested. 26 written claims are registered with their source and limitation, and an automated test fails if the data stops supporting one.
- The website holds no second copy of the numbers. It reads generated JSON, so a statistic cannot drift from the analysis.
Data quality checks
278 automated checks run on every build: expected firm names only, no duplicate records, missing values, valid date ranges and categories, source and claim references, and an independent recomputation of headline metrics directly from the raw files rather than through the metrics script.
Reading the measures
- H-1B and PERM filings measure US hiring and sponsorship activity, not headcount or revenue. Firms differ in how many positions they put on one application.
- Federal prime contracts cover direct awards in the US registry. Subcontract work is reported separately.
- UK, Canada and Australia contracts cover published awards above each registry’s threshold. UK framework ceilings are excluded from totals.
- SEC mentions show who writes about a firm, not where it earns revenue.
- Sitemap counts measure what a firm publishes, not what it earns.
- Scale figures other than BCG’s are secondary-source estimates; verification status is listed per data point on the Data & Sources page.
What was not done
- No significance tests or causal analysis. Findings are descriptive.
- No industry-revenue measure, because no public source exists.
- McKinsey’s website could not be read by automated requests, so McKinsey is absent from sitemap-based counts. Only sitemaps named in robots.txt were read, and no pages were crawled.
- The Power BI dashboard and slide deck planned in the original brief have not been built.
Full method text lives in the repository: docs/methodology.md. To rerun everything, see Reproducibility.