Growth Cab Apply to GC
Blog/AI OPERATIONS
AI OPERATIONS · September 1, 2026 · 6 MIN READ

How to Evaluate AI Models on the Work That Pays You

A practical guide on how to evaluate AI models with representative work, acceptance criteria, review time, latency and cost per usable result.

Federico DonatoneBy Federico Donatone · Founder, Growth Cab
How to Evaluate AI Models on the Work That Pays You

A week ago I ranked fifteen AI models after using them on real work. Fable 5 sat alone at the top. GPT 5.6 Sol won the orchestrator role. Several cheaper models fell because the job still cost more once retries and review entered the picture. The post drew 430 reactions and 122 comments, which told me people wanted the method behind the argument.

The ranking was a field report from my workload inside Growth Cab. It was never a universal table for every company. A model that wins at coding can lose at research, structured extraction or tool use. The useful question is how to evaluate AI models for the job you need completed, under the constraints your team and customers will actually experience.

Federico Donatonein
Federico Donatone
Founder, Growth Cab · This article started as a LinkedIn post

“I ranked 15 AI models. Here are the 5 takes that will start fights:”

430REACTIONS
122COMMENTS
Read the original post →

How to Evaluate AI Models With a Real Work Scorecard

Start by naming one job in operational language. Research an account and produce a sourced brief. Turn a call transcript into a follow-up that preserves every commitment. Modify a landing page and prove the live result. A category such as writing or reasoning is too broad. The job needs a visible finish line that two reviewers can recognize.

Define acceptance before you run the first model. Write the required fields, factual standards, tone, source rules, permission boundaries and maximum review time. Include a few failures that make the output unusable. If the criteria arrive after you read the answers, the most persuasive response will quietly redefine the test in its own favor.

Score five things: usable quality, completion rate, human review time, end-to-end latency and total cost per accepted result. Token price belongs in the last number rather than at the top of the table. A cheap answer that needs three attempts and fifteen minutes of repair can be the most expensive answer in the batch.

Build the Test From Work You Already Paid to Learn

Use twenty to thirty cases from the actual workflow. Include common inputs, awkward inputs, sparse context and examples that previously failed. Keep the original files and expected outcomes. Synthetic prompts make setup fast, but production cases carry the vocabulary, missing data and strange edge conditions that expose whether the system is ready for a real operator.

Freeze the surrounding conditions. Keep the system instructions, prompt, tools, retrieval sources, output schema and effort setting constant across candidates. Record the exact model route and date. If each model receives a different harness, the comparison measures your prompt changes alongside the model. You may still tune later, after the clean baseline reveals where tuning matters.

Run each case more than once when the task has meaningful variation. One brilliant response can be luck, and one strange response can be noise. Repeated runs expose stability. For a workflow that feeds sales, finance or client delivery, predictable acceptable output often creates more value than a higher peak surrounded by volatile failures.

Measure Cost Per Accepted Result

Model pricing makes a weak proxy for operating cost. Add input and output charges, tool calls, failed attempts, waiting time and human review. Then divide the total by outputs the team would actually use. This is how a lower token price can produce a higher cost per job, which was the reason several popular models landed low in my tier list.

Review time deserves its own timer. A response can look polished while hiding one invented source, a missed exclusion or a broken handoff. Count the minutes between first output and acceptance. If a person has to reconstruct the work to trust it, the model produced a draft. It did not complete the job the scorecard was designed to measure.

Latency also has to reach the final artifact. A fast model that waits on a slow tool, retries a browser action or leaves a file in the wrong place is part of a slow system. Start the clock when the operator submits the job and stop it when the usable result reaches the place where the next person or process consumes it.

Test the Workflow Around the Model

The model rarely works alone. It receives a prompt, retrieves context, calls tools, writes files and hands results into another system. Test that route end to end. Check whether sources survived, fields reached the CRM, permissions stayed narrow and the public page or client artifact reflects the change. Fluent prose cannot prove that the underlying job happened.

This is where orchestration becomes a separate capability. GPT 5.6 Sol ranked highly for me because it coordinated work well across systems. Another model may produce stronger prose inside a single response. Those are different jobs. The scorecard should let both win where their evidence is strongest instead of forcing one permanent champion across the company.

Route by failure cost as well as average quality. A lightweight model can handle reversible classification or formatting when its acceptance rate is high. A stronger route may be worth the price for ambiguous research or a change that touches production. Keep a human gate for sending, publishing, paying and any action that is difficult to reverse.

AI FRONTIER
Get one useful AI play every Thursday
The AI changes that matter, Federico's direct read and one practical play, plus the free 10-page Operator Pack.
Free · under five minutes · unsubscribe anytime

Once a model is live, selection becomes maintenance. The companion LLM regression testing workflow explains how to freeze a baseline, catch silent quality drops and preserve every confirmed failure as a permanent test case.

Where This AI Model Evaluation Framework Breaks

A small test set can reward the wrong specialist. Twenty cases may represent this month's work while missing the new customer segment arriving next quarter. Review the set whenever the workflow, data or risk changes. Keep the old cases for continuity, then add new cases before the scorecard hardens into another public benchmark that your real work has outgrown.

Human judgment can drift too. Two reviewers may disagree about tone, completeness or acceptable evidence. Use explicit rubrics, compare disagreements and resolve critical examples before scoring the full batch. An evaluator model can reduce review load after calibration, but its output remains another judgment to inspect rather than an invisible authority over the decision.

The framework also fails when the team evaluates novelty instead of durability. A new model creates attention, and the first impressive demo feels like proof. Run it for seven working days. Track every correction, retry and broken handoff. The best model is the one that keeps producing accepted work after the surprise has disappeared and the boring inputs arrive.

Run a Seven-Day Model Trial

On day one, freeze twenty representative jobs and the acceptance rules. During days two through four, run the same batch through every serious candidate. On day five, inspect failures and reviewer disagreements. On day six, calculate acceptance rate, review minutes, latency and cost per accepted result. On day seven, choose a route for each job and document the boundary.

Keep the raw outputs and the cases that decided the result. A tier list is useful when it compresses evidence. It becomes dangerous when the letters survive and the evidence disappears. My ranking will change because the models and my work will change. The evaluation method should remain stable enough to explain every movement on the board.

Every Thursday, AI Frontier gives you one signal, my read on it, and one practical play from the AI and go to market systems we run inside Growth Cab, all in under five minutes. The original tier list and debate are on LinkedIn. If your current favorite loses the seven-day test, I want to know which failure knocked it out.

Want a GTM engine that runs like this?

Growth Cab is the #1 GTM & sales advisory in the US & Europe. We build the outbound, LinkedIn, and closing systems behind these playbooks for founders selling high-ACV deals.

Apply to GC ← All articles
AI FRONTIER

Turn this week's AI noise
into one useful move

Every Thursday: the signal, Federico's direct view and one practical play. Join free and get the 10-page AI Frontier Operator Pack.

Free · under five minutes · unsubscribe anytime