# GTM workflow evaluation kit

Growth Cab — version 1.0, 8 September 2026.

Use this protocol to compare an enrichment provider, an account-to-draft workflow or a replacement model on the same defined job. This is an original evaluation template, not a benchmark result. The companion CSV contains synthetic cases and no real contact data.

## 1. Freeze the job

Record the decision being evaluated, the owner, the input snapshot, the provider/model configuration, the instruction version and the observation date. Choose one job: accepted account records, accepted research briefs or accepted drafts. Do not mix these output types into a single acceptance rate.

Write the field contract before the run. For each required field specify the meaning, acceptable evidence, freshness requirement and action when missing. Record exclusions and who may approve the next action. A completed record is not permission to send a message.

## 2. Choose the cases before seeing results

Use representative cases from the actual market when authorized. Add difficult cases deliberately: ambiguous identity, outdated source, missing required field, conflicting ownership, exclusion, duplicate import and uncertain action outcome. Keep the selection method and sample size visible. A handpicked sample is useful for finding failures but does not estimate population accuracy.

Use the same input snapshot for each candidate. Preserve unknown values and unavailable sources. The synthetic CSV is a starter test of control behavior; passing it does not establish provider coverage, message quality or sales performance.

## 3. Define acceptance and expected stops

A normal case advances only when the field contract and evidence checks pass. Missing or conflicting evidence must remain visible. An excluded contact stops. A duplicate must not create a second action. A timeout with an uncertain outcome requires a destination check before retrying.

Define whether a case should advance, wait, stop or require review before running it. An expected stop is correct behavior. Keep that control result separate from how many commercially usable records the workflow produces.

## 4. Record comparable costs

Capture provider charges, verification charges, model charges, review minutes and repair minutes. If internal labor is included, state the hourly rate as an assumption. Allocate shared setup cost explicitly or show it separately; do not hide it in a different candidate's total.

Cost per accepted output = attributed operating cost / accepted outputs. If accepted outputs are zero, report the measure as undefined, alongside the cost incurred. Do not divide by one to manufacture a figure.

First-pass acceptance = outputs accepted without repair / outputs submitted for review. Keep excluded inputs and tasks that correctly stopped out of this denominator, and report their counts separately. If there were no reviewed outputs, acceptance is undefined.

Control correctness = cases with the expected control decision / evaluated control cases. Report the case count and failures. Twelve synthetic successes prove only that those twelve cases passed under the tested conditions.

## 5. Use a review ledger

For each case record case_id, input_version, instruction_version, candidate, observed_at, expected_decision, actual_decision, evidence_reference, first_pass_accepted, repair_minutes, cost, reviewer and unresolved_gap. Keep private evidence in your authorized system; do not put prospect data in a public copy of this kit.

Record candidate failures alongside successes. Repeat imports and simulated timeouts in a test environment without external sends. Inspect the resulting records, not only the workflow's success message. Keep an independent review of factual claims wherever possible.

## 6. Decide and preserve the working version

Set acceptance thresholds from the cost of failure in your process. This kit supplies no universal accuracy or savings threshold. A candidate that is faster but loses exclusions or duplicates actions fails an essential control even when its average cost is lower.

Document accepted use, required human review, unsupported cases and rollback conditions. Preserve the previous configuration and representative test cases. A new model, changed source schema or edited prompt should trigger the relevant checks again.

## Companion files and reading

- Synthetic cases: https://www.growthcab.com/assets/gtm-workflow-evaluation-cases.csv
- Enrichment field selection: https://www.growthcab.com/b2b-data-enrichment-tools
- Stack architecture: https://www.growthcab.com/how-to-build-an-ai-gtm-stack
- Regression checks: https://www.growthcab.com/llm-regression-testing-workflow

This kit is provided for practical evaluation. It contains no measured vendor rankings, deliverability guarantee, sales forecast or legal assessment.
