MODEL BENCHMARK

Same scenario. 25× cheaper.

We gave the same multi-step scenario to several leading models and measured whether they got it right, how long each run took and what it cost.

RESULTS

One scenario, 20 runs per model.

Opus 5.5GPT-6 AstraGLM 5.3 Flash
ModelSuccess rateAvg. time per runAvg. cost per run
Opus 5.5Medium effort
95%
1.8 min
$0.73
GPT-6 AstraMedium effort
65%
88s
$0.98
GLM 5.3 FlashMax effort
90%
3.1 min
$0.029

Each model ran the scenario several times at the reasoning effort shown. The figures are averages across runs, last run Sep 24, 2026.

GLM 5.3 Flash success rate90%Opus 5.5 95% · GPT-6 Astra 65%
Cheaper than Opus 5.525×$0.029 vs $0.73 per run
Avg. time per run, GLM 5.3 Flash3.1 minOpus 5.5 1.8 min · GPT-6 Astra 88s
Successful runs per $1 vs Opus 5.524×31 vs 1.3 runs per $1
COST

What the same work costs on each model.

Average model cost of one full run in US dollars,
including every tool call and saved note.

Cost per run

Five messages, start to finish.

Opus 5.5
$0.73per runbaseline
GPT-6 Astra
$0.98per run1.4× Opus 5.5 cost
GLM 5.3 Flash
4% of the Opus 5.5 bill
$0.029per run25× cheaper than Opus 5.5

Bars share one linear scale from 0. Shorter is cheaper.

Your monthly usage

How many workflows like this does your team run each month?

1002505001k2.5k5k10k
Opus 5.5$725/mo
GPT-6 Astra$984/mo
GLM 5.3 Flash$29/mo
$696/mo

saved by running these workflows on GLM 5.3 Flash instead of Opus 5.5, $8,354 a year.

GLM 5.3 Flash first, Opus 5.5 when it fails: $101/mo, 86% less than Opus 5.5 alone.

QUALITY, SPEED, EFFICIENCY

Close on quality, far cheaper.

Success rate

Share of runs that passed.

Opus 5.5
95%
GPT-6 Astra
65%
GLM 5.3 Flash
90%

Line marks 80%.

Time per run

From the first message to the final summary.

Opus 5.5
1.8 min
GPT-6 Astra
88s
GLM 5.3 Flash
3.1 min

Bar is the average. Tick is the slowest 10% of runs (p90).

Efficiency

Successful runs per $1 spent.

Opus 5.5
1.3
GPT-6 Astra
0.7
GLM 5.3 Flash
31

Success rate divided by cost per run.

Success rate against cost

Large dots are each model's success rate and average cost per run. Small dots are its individual runs, placed by what each run cost: filled if the run passed, hollow if it failed. Up and to the left is better.

50%60%70%80%90%100%$0.020$0.050$0.10$0.20$0.50$1.00Cost per run, USD (log scale)Success rate80%Opus 5.5 run: passed, $0.75Opus 5.5 run: passed, $0.76Opus 5.5 run: failed, $0.63Opus 5.5 run: passed, $0.74Opus 5.5 run: passed, $0.76Opus 5.5 run: passed, $0.86Opus 5.5 run: passed, $0.71Opus 5.5 run: passed, $0.71Opus 5.5 run: passed, $0.67Opus 5.5 run: passed, $0.66Opus 5.5 run: passed, $0.71Opus 5.5 run: passed, $0.77Opus 5.5 run: passed, $0.72Opus 5.5 run: passed, $0.69Opus 5.5 run: passed, $0.74Opus 5.5 run: passed, $0.69Opus 5.5 run: passed, $0.77Opus 5.5 run: passed, $0.67Opus 5.5 run: passed, $0.77Opus 5.5 run: passed, $0.74GPT-6 Astra run: failed, $0.91GPT-6 Astra run: passed, $1.02GPT-6 Astra run: failed, $0.88GPT-6 Astra run: failed, $0.95GPT-6 Astra run: passed, $1.06GPT-6 Astra run: passed, $1.04GPT-6 Astra run: passed, $1.02GPT-6 Astra run: passed, $0.91GPT-6 Astra run: failed, $1.12GPT-6 Astra run: failed, $0.90GPT-6 Astra run: passed, $1.14GPT-6 Astra run: passed, $1.00GPT-6 Astra run: passed, $0.91GPT-6 Astra run: passed, $0.96GPT-6 Astra run: failed, $0.89GPT-6 Astra run: passed, $0.91GPT-6 Astra run: failed, $0.92GPT-6 Astra run: passed, $1.07GPT-6 Astra run: passed, $1.07GPT-6 Astra run: passed, $1.01GLM 5.3 Flash run: passed, $0.029GLM 5.3 Flash run: passed, $0.026GLM 5.3 Flash run: passed, $0.027GLM 5.3 Flash run: passed, $0.031GLM 5.3 Flash run: passed, $0.037GLM 5.3 Flash run: passed, $0.030GLM 5.3 Flash run: passed, $0.024GLM 5.3 Flash run: passed, $0.039GLM 5.3 Flash run: passed, $0.027GLM 5.3 Flash run: passed, $0.028GLM 5.3 Flash run: passed, $0.027GLM 5.3 Flash run: passed, $0.022GLM 5.3 Flash run: passed, $0.025GLM 5.3 Flash run: passed, $0.026GLM 5.3 Flash run: failed, $0.028GLM 5.3 Flash run: passed, $0.022GLM 5.3 Flash run: passed, $0.028GLM 5.3 Flash run: failed, $0.049GLM 5.3 Flash run: passed, $0.027GLM 5.3 Flash run: passed, $0.025Opus 5.5: 95% passed, $0.73 per runOpus 5.595% · $0.73 per runGPT-6 Astra: 65% passed, $0.98 per runGPT-6 Astra65% · $0.98 per runGLM 5.3 Flash: 90% passed, $0.029 per runGLM 5.3 Flash90% · $0.029 per run
WHAT THE MODELS WERE TESTED ON

A founder, two weeks away, one unfinished project.

AIChief of staff agent in Vybe · 5 messages

A founder comes back to their AI chief of staff after two weeks away and picks up a project they'd only started discussing. Over five messages, the assistant has to remember, plan, save, adapt and summarize.

  1. 1

    Remember

    Find the earlier conversation about the project without being told what it was, and report the current size of the team.

  2. 2

    Plan

    Turn a rough idea into a clear plan in phases.

  3. 3

    Save

    Store the plan and two hiring needs as notes, tagged so they are easy to find later.

  4. 4

    Adapt

    When the founder reorders the phases, fix the saved plan instead of starting a new one.

  5. 5

    Summarize

    Pull everything together into one accurate, up-to-date summary.

HOW WE SCORED IT

Would a capable human assistant have done this?

A run passes a step only if the assistant
does what a capable human assistant would.

Finds the right history

It recalls the earlier project conversation instead of asking the founder to repeat themselves.

Uses accurate facts

It reports the team size correctly from what it already knows.

Follows through

It saves every note it was asked to save, with the right tags.

Edits instead of duplicating

When the plan changes, the original is updated and exactly one plan remains. Saving a second copy fails the step.

Gets the change right

The updated plan has the new phase order and explains why it changed.

Leaves nothing out

The final summary covers the original conversation, the corrected plan and both hiring needs. Showing the old order or dropping a detail fails the step.

Some checks look for specific facts in the assistant's reply. Others are graded by an independent AI reviewer that reads the reply and the notes the assistant saved. A run counts as a success only if it passes.

AVAILABLE NOW

Pick GLM 5.3 Flash in any
Vybe agent, app or chat.

GLM 5.3 Flash sits in the model picker next to Opus 5.5 and GPT-6 Astra. Switch per agent, or route everyday work to it and keep a frontier model for the hard cases.