Hiring-side brief

What a data or ML seat actually costs you

Every number below is tied to a live page we opened. If a popular figure has no primary page in this pack, it is named as a kill, not printed as a finding. Accessed 2 September 2026.

Not in this brief

The seat is six figures before benefits, compute, or manager time

The median annual wage for data scientists was $120,230 in May 2025.
The lowest 10 percent earned less than $67,240, and the highest 10 percent earned more than $199,130.
Data scientists held about 275,600 jobs in 2025.

BLS Occupational Outlook Handbook, Data Scientists, SOC 15-2051.
https://www.bls.gov/ooh/math/data-scientists.htm

That wage is not fully loaded employer cost. Benefits, GPU, and the hiring manager's calendar sit on top and are not in the OOH median.

O*NET republishes BLS wages for the same SOC and shows the same 120230 / 67240 / 199130 figures. The O*NET wages page wording is "workers on average earn." The BLS OOH wording is median. This brief uses the BLS wording.

The occupation is growing fast, and many openings are replacements

Employment of data scientists is projected to grow 35 percent from 2025 to 2035, much faster than the average for all occupations.

Quick Facts employment change, 2025 to 2035: 95,400.

About 24,800 openings for data scientists are projected each year, on average, over the decade. Many of those openings are expected to result from the need to replace workers who transfer to different occupations or exit the labor force, such as to retire.

Table: SOC 15-2051. Projected employment 2035: 371,000.

Same BLS OOH page.
https://www.bls.gov/ooh/math/data-scientists.htm

The cheap win is the model. The bill is the system around it

developing and deploying ML systems is relatively fast and cheap, but maintaining them over time is difficult and expensive.
We refer to this here as the CACE principle: Changing Anything Changes Everything.

Figure 1 caption: only a small fraction of real-world ML systems is composed of the ML code; surrounding infrastructure is vast. Also: glue code around black-box packages; undeclared consumers; data dependencies.

Sculley et al., Hidden Technical Debt in Machine Learning Systems, NIPS 2015.
https://papers.nips.cc/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf

A NeurIPS 2015 reviewer wrote that technical debt in ML systems carries significant monetary costs. That is reviewer language, not a measured rate.

When AI work fails, the method we can show is interviews

To investigate why AI projects fail, we interviewed 65 experienced data scientists and engineers. Participants had at least five years of experience building AI/ML models in industry or academia.
Too often, trained AI models are deployed that have been optimized for the wrong metrics or do not fit into the overall business workflow and context.

RAND RR-A2680-1, The Root Causes of Failure for Artificial Intelligence Projects.
https://www.rand.org/content/dam/rand/pubs/research_reports/RRA2600/RRA2680-1/RAND_RRA2680-1.pdf

Use n=65 and the workflow/metrics finding. Do not print 80 percent as RAND.

Cost of a wrong hire: no DOL primary in this pack

A 2017 SHRM news page quotes Jorgen Sundberg (Link Humans) that recruiting, hiring, and onboarding can cost as much as 240000 dollars, with extra cost if the person is a poor fit. Named third-party estimate on a SHRM page. Not DOL. Not a SHRM survey sample size.

https://www.shrm.org/topics-tools/news/employee-relations/cost-bad-hire-can-astronomical

The wage is already six figures. A hire who never ships still burns that wage plus the Sculley maintenance bill.

Title versus operating problem

Staffing-firm blogs argue that "data scientist" and "ML engineer" are different jobs. Those pages are vendor copy with no sample size. Operator language from public job ads can be used as a word bank, not as a census. This brief does not invent a mismatch rate.

What production-ready means here

From Sculley: the notebook is the small box. Production means data stays valid, predictions arrive, failures recover, undeclared consumers do not silently depend on a half-built pipeline. RAND's interviewees describe models optimized for the wrong metric or dropped into a workflow they do not fit. That is the hiring-manager problem. It is not a proof that Einstein Data Lab has a bench.

Wanted: people who have shipped under load, know the business question, can use current AI tooling, and have worked inside large messy systems. That is the pitch. It is not a verified Einstein roster.