The Data Factory Operating System
RTF Model
The RTF Model acts as a lever: pull the class mix one way and cost, headcount, and recommended pilot scope move with it. That's what makes it useful before a pilot starts and worth revisiting mid-pilot, if actuals start drifting from what was priced.
Explore the live model yourself, the title above opens the source sheet.
The model
DFOS Pilot RTF Model
The cost-and-capacity model the pilot is scoped from. Everything is denominated in one delivered audio hour and driven by RTF. Class mix goes in; blended RTF, cost per hour, headcount, and the recommended pilot scope come out.
Pilot Economics
| Metric | Value |
|---|---|
| Recommended pilot hours | 100 |
| Classes funded | 12 |
| Estimated pilot cost | $154,833 |
| Average hours per class | 8.3 |
| Estimated pilot RTF | 15.3 |
| Unit Economics at Scale | Value |
|---|---|
| All-in cost per delivered audio hour | $1,548.33 |
| Total program cost | $1,548,332 |
| Cost per interaction | $103.22 |
| Cost per tool-call label | $17.20 |
| Blended RTF at scale | 12.4 |
| Staffing to Deliver 1,000 hrs | Value |
|---|---|
| Annotators required @ 12 months | 42 |
| Annotators required @ 3 months | 166 |
| Contributors required @ 12 months | 4 |
| Contributors required @ 3 months | 15 |
DFOS Pilot Class Selection
Ranks capability classes on value ÷ difficulty and funds the top N within a fixed pilot budget, producing the class mix that feeds the RTF model above.
Pilot Class Selection — which tool-call capabilities to fund
| Pilot budget (audio hrs) | 100 | |
| Classes to fund (top N) | 12 | |
| Floor per funded class (hrs) | 4 | |
| Incedent Elicitation | 3 | |
| Capacity Check | VALID | |
| Estimated Pilot Cost | $154,833 | |
| Hour-weighted difficulty of funded mix | 1.1 | |
| Base annotation RTF | 13.5 | |
| Estimated Pilot RTF | 15.3 | |
| Class ID | Class name | VOID | Voice dependency (1-5) | Traffic freq | Severity (1-5) | Model err rate | Value score | Difficulty (x avg) | Rig produce (1-5) | Instances / hr | Hour weight | Priority | Priority (Difficulty weighted) | Rank | Fund? | Pilot hrs | Est. cost | Instances delivered | Grounding | Production demand signal | Description | Example tool-call shape | Est. tier | Domain | Domain / Category | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| C-03 | filled_pause_hesitation | FALSE | 5 | 45.0% | 3 | 20% | 0.27 | 1.02 | 5 | 20.3 | 0.33 | 6.8 | 6.6 | 1 | YES | 7.2 | $11,159 | 146 | Shriberg 1994 | um / uh; drives endpointing and barge-in tuning | — | Head | SPOKEN-CONVERSATION PHENOMENA | Disfluency | ||
| C-01 | dialogue_acts_core | FALSE | 3 | 80.0% | 2 | 12% | 0.19 | 0.77 | 5 | 36.0 | 0.10 | 2.9 | 3.7 | 2 | YES | 5.0 | $7,772 | 181 | ISO 24617-2; SWBD-DAMSL (42 tags) | Backbone layer on all data | inform / request / confirm / acknowledge layer on every turn | — | Head | SPOKEN-CONVERSATION PHENOMENA | Dialogue structure | |
| C-17 | side_speech_addressee | FALSE | 5 | 17.0% | 4 | 34% | 0.23 | 1.41 | 4 | 7.7 | 0.43 | 4.6 | 3.3 | 3 | YES | 8.2 | $12,708 | 63 | Device-directed-speech detection literature | Background speech not addressed to the agent | suppress action | Tail | SPOKEN-CONVERSATION PHENOMENA | Robustness | ||
| C-09 | alphanumeric_sequence_capture | FALSE | 5 | 9.0% | 4 | 32% | 0.12 | 0.90 | 5 | 4.1 | 0.79 | 2.9 | 3.2 | 4 | YES | 11.8 | $18,200 | 48 | Practical; ATIS-era spelling tasks | #1 voice-specific failure mode for tool args | Order IDs, VINs, emails, confirmation codes spoken/spelled aloud | arg for any lookup call | Head | SPOKEN-CONVERSATION PHENOMENA | ASR-hard entities | |
| C-19 | multi_party | FALSE | 5 | 21.0% | 3 | 36% | 0.23 | 1.54 | 4 | 9.5 | 0.31 | 4.5 | 2.9 | 5 | YES | 7.1 | $10,930 | 67 | AMI / ICSI meeting corpora | Second speaker joins the call | diarization-dependent | Tail | SPOKEN-CONVERSATION PHENOMENA | Robustness | ||
| C-10 | spoken_entity_normalization | FALSE | 5 | 11.0% | 4 | 26% | 0.11 | 1.02 | 5 | 5.0 | 0.57 | 2.9 | 2.8 | 6 | YES | 9.6 | $14,801 | 47 | SLU slot-normalization practice | twenty-two grand', 'next Friday' → canonical values | 22000; ISO date | Head | SPOKEN-CONVERSATION PHENOMENA | ASR-hard entities | ||
| C-04 | user_correction_post_action | FALSE | 5 | 6.0% | 5 | 38% | 0.11 | 1.54 | 5 | 2.7 | 0.69 | 2.9 | 1.9 | 7 | YES | 10.7 | $16,609 | 29 | Schegloff, Jefferson & Sacks 1977 | Highest-value tail class for agent training | Other-initiated repair after agent acts — undo/redo the tool call | chain: cancel → re-execute | Tail | SPOKEN-CONVERSATION PHENOMENA | Repair | |
| B-05 | missing_arg_clarification | FALSE | 3 | 8.0% | 5 | 22% | 0.09 | 0.77 | 5 | 3.6 | 0.48 | 1.3 | 1.7 | 8 | YES | 8.7 | $13,430 | 31 | API-Bank; BFCL | Hallucinated args cause wrong writes — critical | Required argument absent — ask, never hallucinate | clarify → call | Head | TOOL-CALL STRUCTURAL PATTERNS | Structural | |
| C-05 | barge_in_interruption | FALSE | 5 | 8.0% | 4 | 30% | 0.10 | 1.41 | 5 | 3.6 | 0.47 | 2.4 | 1.7 | 9 | YES | 8.6 | $13,378 | 31 | Spoken-dialogue-systems literature | Duplex / streaming agents live or die here | User talks over agent TTS; requires state rollback | — | Tail | SPOKEN-CONVERSATION PHENOMENA | Repair | |
| B-08 | write_confirmation_gate | FALSE | 3 | 8.0% | 5 | 18% | 0.07 | 0.64 | 5 | 3.6 | 0.47 | 1.1 | 1.7 | 10 | YES | 8.6 | $13,317 | 31 | τ-bench | The core safety behavior for voice agents | Read silently; confirm before irreversible writes | confirm → write | Head | TOOL-CALL STRUCTURAL PATTERNS | Structural | |
| C-15 | asr_error_recovery | FALSE | 5 | 7.0% | 3 | 33% | 0.07 | 1.15 | 5 | 3.2 | 0.48 | 1.7 | 1.5 | 11 | YES | 8.7 | $13,461 | 27 | SLU noise-robustness literature | Misrecognition → confirm / repair loop | confirm(slot) | Tail | SPOKEN-CONVERSATION PHENOMENA | Robustness | ||
| B-11 | cross_turn_slot_accumulation | FALSE | 3 | 16.0% | 4 | 20% | 0.13 | 1.41 | 5 | 7.2 | 0.19 | 1.9 | 1.4 | 12 | YES | 5.9 | $9,067 | 42 | MultiWOZ DST; SpokenWOZ | Arguments gathered across many turns (dialogue state tracking) | state → call | Head | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| C-11 | accent_dialect_coverage | TRUE | 5 | 25.0% | 4 | 30% | 0.30 | 0.64 | 3 | 11.3 | 0.00 | 4.5 | 7.0 | - | VOID | 0.0 | $0 | 0 | Common Voice methodology | Fairness + WER tail | L1/L2 and regional variation; conjoins with all intents | — | Mid | SPOKEN-CONVERSATION PHENOMENA | Speaker variation | |
| C-02 | self_repair_same_turn | FALSE | 5 | 6.0% | 4 | 35% | 0.08 | 1.28 | 4 | 2.7 | 0.00 | 1.7 | 1.3 | 14 | Queue | 0.0 | $0 | 0 | Shriberg 1994; Switchboard | Slot extraction breaks exactly here | Reparandum → interregnum → repair ('the red— the blue one') | affects slot extraction | Tail | SPOKEN-CONVERSATION PHENOMENA | Disfluency | |
| C-06 | backchannel | FALSE | 4 | 12.0% | 2 | 18% | 0.04 | 0.51 | 4 | 5.4 | 0.00 | 0.7 | 1.4 | 13 | - | 0.0 | $0 | 0 | SWBD-DAMSL | False-trigger prevention | uh-huh' that is not a turn — must NOT trigger action | suppress | Mid | SPOKEN-CONVERSATION PHENOMENA | Dialogue structure | |
| A-24 | identity_verification | FALSE | 4 | 7.0% | 5 | 16% | 0.06 | 0.90 | 5 | 3.2 | 0.00 | 1.1 | 1.2 | 15 | - | 0.0 | $0 | 0 | Regulated-vertical practice | Gate for every account-touching flow | KBA step-up before any write action | verify_identity(answers) | Head | DOMAIN × INTENT CLASSES | Account & identity | |
| C-20 | prosodic_meaning | FALSE | 5 | 5.0% | 4 | 45% | 0.09 | 1.15 | 3 | 2.3 | 0.00 | 1.4 | 1.2 | 16 | - | 0.0 | $0 | 0 | Prosody literature | Text-identical, label-different pairs | Stress / intonation changes meaning ('I did NOT order that') | affects intent label | Tail | SPOKEN-CONVERSATION PHENOMENA | Prosody | |
| B-10 | tool_error_recovery | FALSE | 3 | 5.0% | 4 | 30% | 0.06 | 0.90 | 5 | 2.3 | 0.00 | 0.9 | 1.0 | 17 | - | 0.0 | $0 | 0 | ToolBench | API error / empty result — retry, reformulate, inform user | retry / fallback | Tail | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| B-06 | no_call_needed | FALSE | 2 | 15.0% | 3 | 14% | 0.06 | 0.64 | 5 | 6.8 | 0.00 | 0.6 | 1.0 | 18 | - | 0.0 | $0 | 0 | BFCL relevance detection | Spurious calls cost latency + errors | Answerable without tools; suppress spurious calls | no call | Head | TOOL-CALL STRUCTURAL PATTERNS | Structural | |
| C-18 | grounding_confirmation_strategy | FALSE | 3 | 22.0% | 2 | 15% | 0.07 | 1.02 | 5 | 9.9 | 0.00 | 1.0 | 1.0 | 19 | - | 0.0 | $0 | 0 | Clark & Schaefer 1989 | Explicit vs implicit confirmation strategy labels | — | Mid | SPOKEN-CONVERSATION PHENOMENA | Dialogue structure | ||
| B-12 | mid_flow_arg_revision | FALSE | 4 | 4.0% | 4 | 32% | 0.05 | 1.28 | 5 | 1.8 | 0.00 | 1.0 | 0.8 | 20 | - | 0.0 | $0 | 0 | MultiWOZ; SpokenWOZ | User changes an earlier slot before execution | state update → call | Tail | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| A-28 | credential_reset_deflection | FALSE | 4 | 3.0% | 5 | 15% | 0.02 | 0.64 | 5 | 1.4 | 0.00 | 0.5 | 0.7 | 21 | - | 0.0 | $0 | 0 | Security practice | Security-mandatory class | Never handle credentials by voice; route to secure reset | send_reset_link(acct) | Mid | DOMAIN × INTENT CLASSES | Account & identity | |
| B-02 | tool_selection_among_many | FALSE | 1 | 30.0% | 3 | 10% | 0.09 | 0.64 | 5 | 13.5 | 0.00 | 0.5 | 0.7 | 21 | - | 0.0 | $0 | 0 | BFCL multiple; ToolBench | Choose the correct tool from several plausible ones | 1 of N tools | Head | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| A-15 | payment_intent_deflection | FALSE | 4 | 3.0% | 5 | 18% | 0.03 | 0.77 | 5 | 1.4 | 0.00 | 0.5 | 0.7 | 23 | - | 0.0 | $0 | 0 | PCI-DSS | Compliance-mandatory class | Detect payment intent; route to secure channel — never capture card data by voice | handoff(secure_payment) | Mid | DOMAIN × INTENT CLASSES | Billing & payments | |
| C-16 | emotional_escalation | FALSE | 4 | 3.0% | 5 | 28% | 0.04 | 1.02 | 4 | 1.4 | 0.00 | 0.7 | 0.7 | 24 | - | 0.0 | $0 | 0 | Call-center emotion literature (IEMOCAP-adjacent) | Frustration / anger → de-escalate or handoff | transfer_to_agent | Tail | SPOKEN-CONVERSATION PHENOMENA | Affect | ||
| C-08 | anaphora_coreference | FALSE | 2 | 14.0% | 3 | 18% | 0.08 | 1.15 | 5 | 6.3 | 0.00 | 0.8 | 0.7 | 25 | - | 0.0 | $0 | 0 | MultiWOZ DST | book it', 'the earlier one' — resolve to entity from context | resolve → call | Head | SPOKEN-CONVERSATION PHENOMENA | Ambiguity | ||
| A-25 | update_contact_info | FALSE | 4 | 4.0% | 4 | 18% | 0.03 | 0.90 | 5 | 1.8 | 0.00 | 0.6 | 0.6 | 26 | - | 0.0 | $0 | 0 | Write with read-back confirmation | update_contact(field, value) | Mid | DOMAIN × INTENT CLASSES | Account & identity | |||
| C-07 | underspecified_request | FALSE | 2 | 7.0% | 3 | 24% | 0.05 | 0.90 | 5 | 3.2 | 0.00 | 0.5 | 0.6 | 27 | - | 0.0 | $0 | 0 | SGD; clarification-question literature | Ambiguous request → clarifying question before any call | ask_clarify | Tail | SPOKEN-CONVERSATION PHENOMENA | Ambiguity | ||
| B-14 | result_grounded_followup | FALSE | 2 | 10.0% | 3 | 18% | 0.05 | 1.02 | 5 | 4.5 | 0.00 | 0.5 | 0.5 | 28 | - | 0.0 | $0 | 0 | SGD | User refers to returned results ('the second one') | resolve ref → call | Head | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| B-13 | multi_intent_decomposition | FALSE | 3 | 5.0% | 4 | 30% | 0.06 | 1.41 | 4 | 2.3 | 0.00 | 0.7 | 0.5 | 29 | - | 0.0 | $0 | 0 | BFCL parallel-multiple | One utterance decomposes into 2+ intents / calls | split → calls | Tail | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| B-04 | sequential_dependent_chain | FALSE | 2 | 5.0% | 5 | 35% | 0.09 | 1.66 | 4 | 2.3 | 0.00 | 0.7 | 0.4 | 30 | - | 0.0 | $0 | 0 | ComplexFuncBench; ToolBench | Chain breaks = task failure | Output of call N feeds call N+1 | chained calls | Tail | TOOL-CALL STRUCTURAL PATTERNS | Structural | |
| C-14 | safety_privacy_refusal | FALSE | 3 | 3.0% | 5 | 16% | 0.02 | 0.90 | 5 | 1.4 | 0.00 | 0.4 | 0.4 | 31 | - | 0.0 | $0 | 0 | Policy | Compliance-mandatory | Requests the agent must refuse (PCI / PII / policy) | refuse + explain | Mid | SPOKEN-CONVERSATION PHENOMENA | Scope & safety | |
| A-04 | cancel_order | FALSE | 2 | 6.0% | 4 | 10% | 0.02 | 0.64 | 5 | 2.7 | 0.00 | 0.2 | 0.4 | 32 | - | 0.0 | $0 | 0 | τ-bench retail; ABCD | Cancel with policy eligibility check | cancel_order(order_id, reason) | Head | DOMAIN × INTENT CLASSES | Commerce & orders | ||
| A-34 | escalate_to_human | FALSE | 2 | 4.0% | 5 | 12% | 0.02 | 0.64 | 5 | 1.8 | 0.00 | 0.2 | 0.4 | 32 | - | 0.0 | $0 | 0 | Universal | Failure here is churn; high severity | Structured handoff with context summary | transfer_to_agent(context) | Head | DOMAIN × INTENT CLASSES | Support & troubleshooting | |
| B-01 | single_simple_call | FALSE | 1 | 45.0% | 2 | 4% | 0.04 | 0.51 | 5 | 20.3 | 0.00 | 0.2 | 0.4 | 34 | - | 0.0 | $0 | 0 | BFCL simple | Baseline competence | One tool, all arguments present in one utterance | single call | Head | TOOL-CALL STRUCTURAL PATTERNS | Structural | |
| B-09 | policy_constrained_action | FALSE | 1 | 5.0% | 5 | 28% | 0.07 | 1.02 | 5 | 2.3 | 0.00 | 0.4 | 0.3 | 35 | - | 0.0 | $0 | 0 | τ-bench | Compliance | Action valid only under policy conditions | check_policy → act | Tail | TOOL-CALL STRUCTURAL PATTERNS | Structural | |
| A-39 | send_message_email | FALSE | 4 | 3.0% | 4 | 16% | 0.02 | 1.15 | 5 | 1.4 | 0.00 | 0.4 | 0.3 | 36 | - | 0.0 | $0 | 0 | MASSIVE | Irreversible-send confirmation gate | Compose; confirm before irreversible send | send_message(to, body) | Mid | DOMAIN × INTENT CLASSES | Assistant & device | |
| A-05 | initiate_return | FALSE | 2 | 7.0% | 3 | 12% | 0.03 | 0.77 | 5 | 3.2 | 0.00 | 0.3 | 0.3 | 37 | - | 0.0 | $0 | 0 | ABCD | Eligibility check then RMA creation | create_return(order_id, items, reason) | Head | DOMAIN × INTENT CLASSES | Commerce & orders | ||
| C-12 | code_switching | FALSE | 5 | 4.0% | 3 | 38% | 0.05 | 1.41 | 2 | 1.8 | 0.00 | 0.5 | 0.3 | 38 | - | 0.0 | $0 | 0 | Bangor Miami corpus; CS literature | Mid-utterance language switch | — | Tail | SPOKEN-CONVERSATION PHENOMENA | Speaker variation | ||
| A-17 | book_appointment | FALSE | 2 | 8.0% | 3 | 12% | 0.03 | 0.90 | 5 | 3.6 | 0.00 | 0.3 | 0.3 | 39 | - | 0.0 | $0 | 0 | MultiWOZ; SGD | Slot search then booking | find_slots(...); book_slot(id) | Head | DOMAIN × INTENT CLASSES | Scheduling & booking | ||
| A-21 | book_travel | FALSE | 3 | 3.0% | 4 | 22% | 0.03 | 1.28 | 5 | 1.4 | 0.00 | 0.4 | 0.3 | 40 | - | 0.0 | $0 | 0 | ATIS; MultiWOZ; τ-bench airline | Multi-slot flight / hotel / train booking | book_flight(origin, dest, date, class) | Head | DOMAIN × INTENT CLASSES | Scheduling & booking | ||
| A-08 | update_shipping_address | FALSE | 3 | 3.0% | 4 | 15% | 0.02 | 0.90 | 5 | 1.4 | 0.00 | 0.3 | 0.3 | 41 | - | 0.0 | $0 | 0 | τ-bench write-gates | Write action; confirm before commit | update_order_address(order_id, addr) | Mid | DOMAIN × INTENT CLASSES | Commerce & orders | ||
| B-07 | near_miss_distractor_tool | FALSE | 1 | 4.0% | 4 | 32% | 0.05 | 0.90 | 5 | 1.8 | 0.00 | 0.3 | 0.3 | 42 | - | 0.0 | $0 | 0 | BFCL | Plausible wrong tool present in the schema | distractor present | Tail | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| C-13 | out_of_scope_deflection | FALSE | 2 | 4.0% | 3 | 15% | 0.02 | 0.51 | 4 | 1.8 | 0.00 | 0.1 | 0.3 | 43 | - | 0.0 | $0 | 0 | CLINC150 OOS | Not in taxonomy — decline gracefully, never force-classify | no_call + deflect | Head | SPOKEN-CONVERSATION PHENOMENA | Scope & safety | ||
| A-14 | subscription_change | FALSE | 2 | 5.0% | 4 | 12% | 0.02 | 0.90 | 5 | 2.3 | 0.00 | 0.2 | 0.3 | 44 | - | 0.0 | $0 | 0 | SaaS / telecom demand | Upgrade / downgrade / cancel with retention offer | update_subscription(plan_id) | Head | DOMAIN × INTENT CLASSES | Billing & payments | ||
| A-45 | trade_in_appraisal | FALSE | 4 | 2.0% | 3 | 24% | 0.01 | 1.15 | 5 | 0.9 | 0.00 | 0.3 | 0.3 | 45 | - | 0.0 | $0 | 0 | Conjunction with C-09 (VIN capture) | VIN-keyed valuation flow; alphanumeric-heavy | get_appraisal(vin) | Mid | DOMAIN × INTENT CLASSES | Automotive retail (example vertical) | ||
| A-03 | modify_order_items | FALSE | 2 | 4.0% | 4 | 14% | 0.02 | 0.90 | 5 | 1.8 | 0.00 | 0.2 | 0.2 | 46 | - | 0.0 | $0 | 0 | τ-bench retail | High save-rate value | Change items/quantity pre-fulfillment; time-window policy | modify_order(order_id, changes) | Mid | DOMAIN × INTENT CLASSES | Commerce & orders | |
| A-18 | reschedule_appointment | FALSE | 2 | 5.0% | 3 | 18% | 0.03 | 1.15 | 5 | 2.3 | 0.00 | 0.3 | 0.2 | 47 | - | 0.0 | $0 | 0 | MultiWOZ | Cancel + rebook dependent chain preserving constraints | chain: cancel → book | Head | DOMAIN × INTENT CLASSES | Scheduling & booking | ||
| A-12 | dispute_charge | FALSE | 2 | 3.0% | 4 | 14% | 0.02 | 0.77 | 5 | 1.4 | 0.00 | 0.2 | 0.2 | 48 | - | 0.0 | $0 | 0 | Contact-center demand | Regulated timelines (Reg E / Reg Z) | Open dispute case with transaction reference | create_dispute(txn_id, reason) | Mid | DOMAIN × INTENT CLASSES | Billing & payments | |
| A-01 | check_order_status | FALSE | 2 | 13.0% | 2 | 5% | 0.01 | 0.64 | 5 | 5.9 | 0.00 | 0.1 | 0.2 | 49 | - | 0.0 | $0 | 0 | ABCD; τ-bench retail; SGD | Top-3 contact-center intent by volume in retail | Look up order state by ID or phone number | get_order(order_id) | Head | DOMAIN × INTENT CLASSES | Commerce & orders | |
| A-30 | create_ticket | FALSE | 2 | 6.0% | 3 | 10% | 0.02 | 0.90 | 5 | 2.7 | 0.00 | 0.2 | 0.2 | 50 | - | 0.0 | $0 | 0 | Ticket creation with summary + priority | create_ticket(summary, priority) | Head | DOMAIN × INTENT CLASSES | Support & troubleshooting | |||
| A-37 | smart_home_control | FALSE | 3 | 5.0% | 2 | 10% | 0.01 | 0.77 | 5 | 2.3 | 0.00 | 0.2 | 0.2 | 51 | - | 0.0 | $0 | 0 | SLURP iot domain | Device on/off/setpoint by device + room | iot.set(device, state) | Head | DOMAIN × INTENT CLASSES | Assistant & device | ||
| B-03 | parallel_independent_calls | FALSE | 2 | 3.0% | 3 | 30% | 0.03 | 1.15 | 4 | 1.4 | 0.00 | 0.2 | 0.2 | 52 | - | 0.0 | $0 | 0 | BFCL parallel | Check both my orders' — simultaneous independent calls | 2+ parallel calls | Tail | TOOL-CALL STRUCTURAL PATTERNS | Structural | ||
| A-19 | cancel_appointment | FALSE | 2 | 4.0% | 3 | 10% | 0.01 | 0.64 | 5 | 1.8 | 0.00 | 0.1 | 0.2 | 53 | - | 0.0 | $0 | 0 | MultiWOZ | Cancellation with fee/policy check | cancel_booking(id) | Head | DOMAIN × INTENT CLASSES | Scheduling & booking | ||
| A-40 | navigation_transport | FALSE | 3 | 4.0% | 2 | 12% | 0.01 | 0.77 | 5 | 1.8 | 0.00 | 0.1 | 0.2 | 54 | - | 0.0 | $0 | 0 | MASSIVE transport | Directions, transit queries | navigate(destination) | Head | DOMAIN × INTENT CLASSES | Assistant & device | ||
| A-29 | guided_diagnosis | FALSE | 2 | 5.0% | 3 | 20% | 0.03 | 1.41 | 4 | 2.3 | 0.00 | 0.2 | 0.2 | 55 | - | 0.0 | $0 | 0 | SGD; telecom demand | Multi-turn decision tree with device lookups | get_device_status(); run_diagnostic() | Mid | DOMAIN × INTENT CLASSES | Support & troubleshooting | ||
| A-16 | promo_apply | FALSE | 3 | 3.0% | 2 | 12% | 0.01 | 0.64 | 5 | 1.4 | 0.00 | 0.1 | 0.2 | 56 | - | 0.0 | $0 | 0 | Validate and apply code; explain failures | apply_promo(code) | Mid | DOMAIN × INTENT CLASSES | Billing & payments | |||
| A-13 | payment_plan_setup | FALSE | 2 | 2.0% | 5 | 20% | 0.02 | 1.02 | 4 | 0.9 | 0.00 | 0.2 | 0.2 | 57 | - | 0.0 | $0 | 0 | Collections demand | Eligibility-gated payment arrangement | create_payment_plan(acct, terms) | Tail | DOMAIN × INTENT CLASSES | Billing & payments | ||
| A-07 | refund_status | FALSE | 2 | 6.0% | 2 | 6% | 0.01 | 0.51 | 5 | 2.7 | 0.00 | 0.1 | 0.1 | 58 | - | 0.0 | $0 | 0 | ABCD | Refund state lookup and timeline explanation | get_refund_status(refund_id) | Head | DOMAIN × INTENT CLASSES | Commerce & orders | ||
| A-02 | track_shipment | FALSE | 2 | 9.0% | 2 | 5% | 0.01 | 0.64 | 5 | 4.1 | 0.00 | 0.1 | 0.1 | 59 | - | 0.0 | $0 | 0 | SGD; ABCD | WISMO ('where is my order') dominates retail call volume | Carrier tracking lookup + ETA explanation | track_shipment(tracking_no) | Head | DOMAIN × INTENT CLASSES | Commerce & orders | |
| A-11 | bill_explanation | FALSE | 1 | 15.0% | 2 | 6% | 0.02 | 0.64 | 5 | 6.8 | 0.00 | 0.1 | 0.1 | 59 | - | 0.0 | $0 | 0 | HarperValleyBank; contact-center | #1 call driver in telecom / utilities | Explain line-item or unexpected charge | get_invoice(acct, period) | Head | DOMAIN × INTENT CLASSES | Billing & payments | |
| A-36 | timer_alarm_reminder | FALSE | 3 | 5.0% | 2 | 6% | 0.01 | 0.64 | 5 | 2.3 | 0.00 | 0.1 | 0.1 | 59 | - | 0.0 | $0 | 0 | SLURP | Set / cancel / query | set_alarm(time) | Head | DOMAIN × INTENT CLASSES | Assistant & device | ||
| A-32 | warranty_check | FALSE | 3 | 3.0% | 2 | 12% | 0.01 | 0.77 | 5 | 1.4 | 0.00 | 0.1 | 0.1 | 62 | - | 0.0 | $0 | 0 | Serial-keyed warranty lookup | check_warranty(serial) | Mid | DOMAIN × INTENT CLASSES | Support & troubleshooting | |||
| A-09 | product_search_availability | FALSE | 2 | 5.0% | 2 | 10% | 0.01 | 0.77 | 5 | 2.3 | 0.00 | 0.1 | 0.1 | 63 | - | 0.0 | $0 | 0 | SGD | Filtered inventory / availability search | search_products(filters) | Head | DOMAIN × INTENT CLASSES | Commerce & orders | ||
| A-06 | exchange_item | FALSE | 2 | 2.0% | 4 | 28% | 0.02 | 1.41 | 4 | 0.9 | 0.00 | 0.2 | 0.1 | 64 | - | 0.0 | $0 | 0 | τ-bench retail | Chain failures are costly and common | Return + reorder as a dependent chain | chain: create_return → place_order | Tail | DOMAIN × INTENT CLASSES | Commerce & orders | |
| A-43 | financing_prequal_status | FALSE | 2 | 2.0% | 5 | 16% | 0.02 | 1.02 | 4 | 0.9 | 0.00 | 0.1 | 0.1 | 65 | - | 0.0 | $0 | 0 | Regulated disclosures | Application status; compliance-sensitive language | get_application(app_id) | Mid | DOMAIN × INTENT CLASSES | Automotive retail (example vertical) | ||
| A-33 | schedule_technician | FALSE | 2 | 3.0% | 3 | 20% | 0.02 | 1.28 | 4 | 1.4 | 0.00 | 0.1 | 0.1 | 66 | - | 0.0 | $0 | 0 | Booking chain with parts + geography constraints | chain: check_parts → book_visit | Mid | DOMAIN × INTENT CLASSES | Support & troubleshooting | |||
| A-35 | media_control | FALSE | 3 | 6.0% | 1 | 6% | 0.00 | 0.51 | 5 | 2.7 | 0.00 | 0.1 | 0.1 | 67 | - | 0.0 | $0 | 0 | SLURP; MASSIVE | Play / pause / skip / volume | media.play(query) | Head | DOMAIN × INTENT CLASSES | Assistant & device | ||
| A-44 | delivery_pickup_scheduling | FALSE | 2 | 2.0% | 3 | 14% | 0.01 | 0.90 | 5 | 0.9 | 0.00 | 0.1 | 0.1 | 68 | - | 0.0 | $0 | 0 | Slot booking with logistics constraints | schedule_delivery(vin, slot) | Mid | DOMAIN × INTENT CLASSES | Automotive retail (example vertical) | |||
| A-42 | vehicle_search | FALSE | 2 | 3.0% | 2 | 10% | 0.01 | 0.77 | 5 | 1.4 | 0.00 | 0.1 | 0.1 | 69 | - | 0.0 | $0 | 0 | SGD-analog schema | Inventory filter by make / model / price / features | search_inventory(filters) | Head | DOMAIN × INTENT CLASSES | Automotive retail (example vertical) | ||
| A-38 | list_management | FALSE | 3 | 3.0% | 1 | 8% | 0.00 | 0.51 | 5 | 1.4 | 0.00 | 0.0 | 0.1 | 70 | - | 0.0 | $0 | 0 | MASSIVE | Add / remove / read list items | list.add(item) | Head | DOMAIN × INTENT CLASSES | Assistant & device | ||
| A-26 | account_close | FALSE | 2 | 1.0% | 5 | 14% | 0.01 | 0.90 | 4 | 0.5 | 0.00 | 0.1 | 0.1 | 71 | - | 0.0 | $0 | 0 | High severity | Irreversible; retention flow + confirmation gate | close_account(acct_id) | Tail | DOMAIN × INTENT CLASSES | Account & identity | ||
| A-27 | preference_update | FALSE | 1 | 2.0% | 3 | 8% | 0.00 | 0.51 | 5 | 0.9 | 0.00 | 0.0 | 0.0 | 72 | - | 0.0 | $0 | 0 | Notification / marketing consent toggles | set_preference(key, value) | Mid | DOMAIN × INTENT CLASSES | Account & identity | |||
| A-22 | modify_reservation_policy | FALSE | 1 | 1.0% | 4 | 26% | 0.01 | 1.15 | 4 | 0.5 | 0.00 | 0.0 | 0.0 | 73 | - | 0.0 | $0 | 0 | τ-bench airline | Policy violations have direct dollar cost | Changes gated by fare rules and change fees | modify_booking(id, changes) | Tail | DOMAIN × INTENT CLASSES | Scheduling & booking | |
| A-46 | registration_title_status | FALSE | 1 | 1.0% | 3 | 28% | 0.01 | 1.15 | 4 | 0.5 | 0.00 | 0.0 | 0.0 | 74 | - | 0.0 | $0 | 0 | State-dependent rules; long tail of edge cases | get_title_status(vin, state) | Tail | DOMAIN × INTENT CLASSES | Automotive retail (example vertical) | |||
| A-20 | availability_query | FALSE | 1 | 5.0% | 1 | 5% | 0.00 | 0.51 | 5 | 2.3 | 0.00 | 0.0 | 0.0 | 75 | - | 0.0 | $0 | 0 | MultiWOZ | Read-only slot lookup | find_slots(filters) | Head | DOMAIN × INTENT CLASSES | Scheduling & booking | ||
| A-31 | ticket_status | FALSE | 1 | 5.0% | 1 | 5% | 0.00 | 0.51 | 5 | 2.3 | 0.00 | 0.0 | 0.0 | 75 | - | 0.0 | $0 | 0 | Status lookup and next-step explanation | get_ticket(ticket_id) | Head | DOMAIN × INTENT CLASSES | Support & troubleshooting | |||
| A-41 | info_query_weather_news | FALSE | 1 | 5.0% | 1 | 4% | 0.00 | 0.51 | 5 | 2.3 | 0.00 | 0.0 | 0.0 | 77 | - | 0.0 | $0 | 0 | SLURP; MASSIVE | Tool-backed informational answers | weather.get(location) | Head | DOMAIN × INTENT CLASSES | Assistant & device | ||
| A-10 | price_match_request | FALSE | 1 | 1.0% | 2 | 16% | 0.00 | 0.77 | 4 | 0.5 | 0.00 | 0.0 | 0.0 | 78 | - | 0.0 | $0 | 0 | Contact-center demand | Policy lookup + case creation | create_case(type=price_match) | Tail | DOMAIN × INTENT CLASSES | Commerce & orders | ||
| A-23 | waitlist_join | FALSE | 1 | 1.0% | 1 | 10% | 0.00 | 0.51 | 5 | 0.5 | 0.00 | 0.0 | 0.0 | 79 | - | 0.0 | $0 | 0 | Join waitlist with notify preference | join_waitlist(id) | Tail | DOMAIN × INTENT CLASSES | Scheduling & booking | |||
What it means
The point of this model isn't the price of a single delivered audio hour, it's using that price to test whether a capability class is worth funding before committing to it at full scale. RTF is treated as a learning curve rather than a constant, so the pilot's blended cost looks worse than where the program eventually settles, and difficulty multiplies that cost instead of sitting in its own line item, so a hard-tail class shows up as expensive here before it shows up as a blown budget in production.
That's what makes it a go/no-go tool rather than a dashboard to check in on: run a class mix through the sheet, read the blended cost it returns, and decide pilot scope before committing headcount. That decision is what the Capacity Model turns into a hiring plan.
Read-outs
The sheet is the model; these are views onto it, showing whether production is holding to what the model assumes, and the unit it prices.
DFOS Production Dashboard
Open full screenIn-flight dashboard to measure quality and trace fidelity. Allows the team to measure progress and develop operationally tangible action plans.