The Data Factory Operating System
Pilot → Scale, in Seven Checkpoints
Both models are answering the same two questions at every scale: what does a delivered audio-hour cost, and how many people does it take to deliver it. The seven checkpoints below trace that answer as it actually gets built: pick which of 70-plus candidate classes are worth a 100-hour pilot, price what the pilot costs down to the dollar, run it against five tripwires that decide graduate, keep piloting, or kill, then re-run the same math at 1,000 hours once the shape has earned it. From there cost turns into a hiring plan: convert labor-hours into headcount, then staff against it, contributors first, since nothing exists for an annotator to label until a contributor has recorded it, and annotators recruiting from day one anyway, because certification takes longer to clear than sourcing does.
- 1
Shape the class mix
RTF Model · 1. Pilot Class SelectionWhich of 70-plus candidate capability classes are worth funding first?
- 100 hrs
- Pilot budget
- 12
- Classes funded
- 8.3
- Avg hrs / class
- 1.1× avg
- Mix difficulty
We can use the Pilot Class Selection tab to select varied classes and mold a rough shape of the data we'd like to build. Once the shape and number of classes have been determined we input the amount of hours required for the pilot. For the worked example we're assuming 100 hours is the minimum threshold for quality output. From here we can see how many hours are required for each class along with the relative difficulty. Our value-add comes from targeting long-tail classes, so we expect a higher-than-average difficulty for annotation.
- 2
Pay the tuition
RTF Model · 2. AssumptionsWhat does the pilot actually cost, and where does that money actually go?
- $154,833
- Pilot cost
- $512.50
- Contributor $ / produced hr
- $392.08
- Annotator $ / produced hr
- $3,168.33
- $ / delivered hr
The business model encompasses both audio contribution and annotation. A varied mix of Wizard-of-Oz (WoZ) sessions, scripted speech, and synthetic TTS is used for the contribution side. Again, focusing on the long-tail requires a larger contribution of WoZ sessions, trading off cost for depth. For the annotation side, we're assuming a base Real-Time-Factor (RTF) of 13.5 labor-hours per audio-hour, which jumps to 15.3 when adjusting for the relative difficulty we calculated from the data shape. Accounting for QA, reworks, and fixed operational costs, we land on a pilot figure of $3,168.33.
- 3
Graduate or kill
RTF Model · 3. Pilot Portfolio (in-flight)Now that the pilot is running, which classes earned the scale-up?
- 2–3 days
- Fastest tripwire
- 3–6 wks
- Slowest tripwire
- $2,090 / pt
- Cheapest SCALE
- $759,430 / pt
- Priciest in-flight
We use five tripwires while the pilot is running to determine whether classes within the data shape should be scaled, need to be piloted further, or should be cut altogether. The tripwires include:
- Elicitation yield — does the scenario even produce the class?
- Inter-annotator kappa — is it labelable?
- Count floor — do we have enough to judge lift yet?
- Marginal lift per +100 labels — does the model actually learn from it?
- Cost per point of lift vs. median — is it gold per dollar?
Ultimately this acts as the gate that further refines the shape of the data and optimizes value output.
- 4
Scale the answer
RTF Model · 2. AssumptionsSame arithmetic, ten times the volume — where does the cost curve land?
- $1,548.33
- $ / delivered hr
- $1,548,332
- Program total
- $103.22
- $ / interaction
- 12.4
- Blended RTF
Once the shape has been refined we can model cost at scale, pushing from 100 hours to 1,000 hours. We can reasonably expect improvements in both RTF and yield as our systems become more refined and annotators earn tenure. In this case, we have RTF improving by 10% every time volume doubles. The final cost per delivered hour and program cost accounts for these gains and gives the team a target to measure progress against.
- 5
Staff the program
RTF Model · 4. Capacity Plan1,000 hours costs a fixed amount of labor — how does that convert into headcount, for the pool that has to ramp first?
- 22,232
- Total labor-hrs
- 496.3
- Annotator-months
- 15 contrib. / 56 annot.
- 3-mo plan
- 4 contrib. / 4 annot.
- 12-mo plan
The Capacity Plan tab takes the 1,000-hour target and converts it into labor-hours, then headcount, for both pools. Annotation labor-hours come from produced hours times blended RTF, adjusted for redundancy and adjudication, with a QA pass added on top. Dividing by productive hours per person per month gives us the total in annotator-months, which we can then spread across a shorter or longer calendar. For the worked example, we size the contributor pool against the same timeline, since their output sets the pace annotation can actually run at.
- 6
Recruit contributors first
Capacity Model · 2. Funnel & Hiring PlanNothing exists for an annotator to label until a contributor has recorded it — who fills the WoZ sessions?
- 200
- Delivered-hrs demand
- 285.7
- Recorded hrs required
- 9
- Candidates needed
- 4.3
- Weeks to productive
We use the Funnel & Hiring Plan tab to size contributor recruiting against the pilot's raw-audio demand. Sourcing runs through five stages, from an initial response through profile verification to actually producing sessions, and each stage loses some share of candidates along the way. Because our worked example leans on the long-tail, we source by condition (i.e. acoustic environment and fluency profile) rather than as one undifferentiated pool, so a scarce profile like accented speech gets recruited separately from the rest. Contributors have to be producing before there's anything to annotate, so this pool has to start first.
- 7
Certify annotators next
Capacity Model · 2. Funnel & Hiring PlanOnce contributors are producing, can annotation capacity keep pace — and can recruiting actually deliver it in time?
- 20 certified
- Target pool
- 26% cum.
- Funnel yield
- 78
- Candidates needed
- 6.4
- Weeks to productive
The same Funnel & Hiring Plan tab handles annotator certification, which runs candidates through screening, a calibration batch, and a gold-set attempt before anyone is certified. Each stage carries its own pass rate, so the yields compound down to a fraction of who we started with. The target pool size comes straight from the Capacity Plan's headcount for whichever timeline we picked in the step above, so changing the calendar there moves this number too. Certification takes longer than contributor sourcing, so we start recruiting annotators on day one rather than waiting for contributors to already be producing.