The Data Factory Operating System
Capacity Model
The same arithmetic, expressed in people. Once the pipeline is running, workforce health says whether the plan is holding: who's performing, where the pool is thin, and whether the recruiting funnel can refill it in time. The two workforces are scored separately: the collection side owns the yield curve, the annotation side owns the RTF curve.
Explore the live model yourself, the title above opens the source sheet.
The model
DFOS Staffing Model
The workforce side of the same arithmetic: how many annotators and contributors the pilot needs, in which roles, over how many months. The two workforces are scored separately: the collection side owns the yield curve, the annotation side owns the RTF curve.
Capacity Model
| Annotation workforce | Value |
|---|---|
| Target certified pool | 100 |
| Current certified pool | 0 |
| Candidates to source | 387.0 |
| Sourcing rate needed (hc/wk) | 60.2 |
| Weeks till productive | 6.4 |
| Contributor workforce | Value |
|---|---|
| Pilot delivered hours (demand) | 200 |
| Recorded Hours required | 285.7 |
| Output rate (hrs/mo) | 286 |
| Collection window (months) | 1 |
| Candidates to source | 9 |
| Sourcing rate needed (hc/wk) | 2.1 |
| Weeks till productive | 4.3 |
What it means
This turns pilot demand into a hiring plan tied to a calendar: the tighter the timeline, the more people have to run in parallel to hit it.
Once the pipeline is running, the same model becomes a diagnostic. The two workforces are scored separately because they fail differently: a thin annotation pool caps throughput on the hard-tail classes the RTF Model is funding on purpose, while a thin contribution pool shows up as scarce raw material rather than a quality problem. Whichever one the model flags shows whether the fix is recruiting or coaching, and whether the recruiting funnel can catch up before the calendar slips.