Same scenario. 25× cheaper.
We gave the same multi-step scenario to several leading models and measured whether they got it right, how long each run took and what it cost.
One scenario, 20 runs per model.
| Model | Success rate | Avg. time per run | Avg. cost per run |
|---|---|---|---|
95% | 1.8 min | $0.73 | |
65% | 88s | $0.98 | |
90% | 3.1 min | $0.029 |
Each model ran the scenario several times at the reasoning effort shown. The figures are averages across runs, last run Sep 24, 2026.
What the same work costs on each model.
Average model cost of one full run in US dollars,
including every tool call and saved note.
Cost per run
Five messages, start to finish.
Bars share one linear scale from 0. Shorter is cheaper.
Your monthly usage
How many workflows like this does your team run each month?
saved by running these workflows on GLM 5.3 Flash instead of Opus 5.5, $8,354 a year.
GLM 5.3 Flash first, Opus 5.5 when it fails: $101/mo, 86% less than Opus 5.5 alone.
Close on quality, far cheaper.
Success rate
Share of runs that passed.
Line marks 80%.
Time per run
From the first message to the final summary.
Bar is the average. Tick is the slowest 10% of runs (p90).
Efficiency
Successful runs per $1 spent.
Success rate divided by cost per run.
Success rate against cost
Large dots are each model's success rate and average cost per run. Small dots are its individual runs, placed by what each run cost: filled if the run passed, hollow if it failed. Up and to the left is better.
A founder, two weeks away, one unfinished project.
A founder comes back to their AI chief of staff after two weeks away and picks up a project they'd only started discussing. Over five messages, the assistant has to remember, plan, save, adapt and summarize.
- 1
Remember
Find the earlier conversation about the project without being told what it was, and report the current size of the team.
- 2
Plan
Turn a rough idea into a clear plan in phases.
- 3
Save
Store the plan and two hiring needs as notes, tagged so they are easy to find later.
- 4
Adapt
When the founder reorders the phases, fix the saved plan instead of starting a new one.
- 5
Summarize
Pull everything together into one accurate, up-to-date summary.
Would a capable human assistant have done this?
A run passes a step only if the assistant
does what a capable human assistant would.
Finds the right history
It recalls the earlier project conversation instead of asking the founder to repeat themselves.
Uses accurate facts
It reports the team size correctly from what it already knows.
Follows through
It saves every note it was asked to save, with the right tags.
Edits instead of duplicating
When the plan changes, the original is updated and exactly one plan remains. Saving a second copy fails the step.
Gets the change right
The updated plan has the new phase order and explains why it changed.
Leaves nothing out
The final summary covers the original conversation, the corrected plan and both hiring needs. Showing the old order or dropping a detail fails the step.
Some checks look for specific facts in the assistant's reply. Others are graded by an independent AI reviewer that reads the reply and the notes the assistant saved. A run counts as a success only if it passes.
Pick GLM 5.3 Flash in any
Vybe agent, app or chat.
GLM 5.3 Flash sits in the model picker next to Opus 5.5 and GPT-6 Astra. Switch per agent, or route everyday work to it and keep a frontier model for the hard cases.