Why pay more for the same work?

We gave the same real work tasks to GLM 5.3 Flash, Opus 5.5 and GPT-6 Astra. Read each answer and compare score, time and cost.

Team
5 tasks × 3 models

Same request, every model. Compare the answers yourself.

The requesteasyGoogle Drive

“From the Search Console export "gsc_export_90d CSV" (in the "06 Spreadsheets" folder of the "Lumora Eval Data" folder in my Google Drive), find queries gaining and losing clicks, pages with high impressions but low CTR, and propose fixes.”

GLM 5.3 FlashMax effortLow-cost

$0.012
2.2 min · 98.2% score40× cheaper than GPT-6 Astra

Snapshot (90d vs prior 90d):

  • Total clicks across the 8 queries: 3,130 vs 2,670 → +460 (+17.2%) overall
  • Biggest non-brand win is agent evaluation framework; the biggest loss is your branded query
  • Two huge generic terms are bleeding impressions with near-zero CTR — that's the real money left on the table

Queries gaining clicks:

  • agent evaluation framework: 640 vs 310 → +330 (+106.5%) · 14,800 impr · 4.3% CTR · pos 5.1 — your hottest query, nearly page 1
  • ai agent platform: 1,180 vs 790 → +390 (+49.4%) · 32,400 impr · 3.6% CTR · pos 4.2 — top driver, healthy CTR for pos 4
  • support ticket automation: 330 vs 290 → +40 (+13.8%) · pos 7.1 · 2.9% CTR — climbing steadily
  • crm for ai agents: 410 vs 380 → +30 (+7.9%) · pos 6.0 — solid niche hold
  • sso saml providers: 120 vs 110 → +10 (+9.1%) — growing, but weak CTR (see below)

Queries losing clicks:

  • lumora reviews (branded): 90 vs 340 → −250 (−73.5%) · pos 8.4 · CTR flat at 4.3% — clicks collapsed while CTR held, which means the ranking itself slipped or a review aggregator (G2, Capterra) took the top SERP slots. Check this one first — brand demand is leaking
  • customer support automation: 210 vs 260 → −50 (−19.2%) · pos 14.0
  • workflow automation software: 150 vs 190 → −40 (−21.1%) · pos 16.2 — both sliding slowly down page 2

High impressions, very low CTR (the striking-distance goldmine):

  • customer support automation: 41,200 impr (largest in the file) at 0.5% CTR, pos 14.0
  • workflow automation software: 38,900 impr at 0.4% CTR, pos 16.2
  • sso saml providers: 15,600 impr at 0.8% CTR, pos 12.5
  • Pattern: all three sit on page 2 for commercial-intent category terms where searchers expect comparison/list pages. If Lumora's matching page is a feature or blog post, that's the CTR killer

Proposed fixes (ranked):

  1. Rewrite/publish dedicated comparison pages for "customer support automation" and "workflow automation software" — head-to-head vs. top competitors, pricing table, and "best of" structure. These two queries alone = 80K impressions; even reaching pos 8 at 2% CTR is ~1,600 extra clicks per 90 days
  2. Internal-link pass: point 5–10 links from your strongest pages (the "ai agent platform" content) at those two category pages to lift them past position 10
  3. Diagnose "lumora reviews": pull the live SERP — if G2/Capterra own the top 3, feed them fresh reviews and add a "Reviews" landing page on your own domain targeting the query
  4. "sso saml providers": either commit to a "SSO/SAML providers comparison" article (listicle intent) or accept it as low-fit and don't chase it
  5. Defend the winners: "agent evaluation framework" is one CTR-tweak away from page 1 — refresh the title tag/meta to push 4.3% → 6%+

Want me to draft the two comparison-page briefs, or set this up as a recurring weekly SEO report so deltas come from a tracked ledger instead of a one-off export?

Opus 5.5Medium effort

$0.23
56s · 100% score2.1× cheaper than GPT-6 Astra

Single match. Downloading as CSV.The site gained 460 clicks this period, from 2,670 to 3,130 (+17.2%). But about 58% of impressions come from three queries that sit on page 2 of Google and get almost no clicks. That's where the upside is.

One gap in the file: it has only 8 rows and no page URLs or previous-period impressions and rankings. So I couldn't rank actual pages. The low-CTR section below goes by query, and I can't tell whether the drops come from worse rankings or fewer searches.

Queries gaining clicks

QueryClicks (prev → now)ChangePosition
ai agent platform790 → 1,180+390 (+49.4%)4.2
agent evaluation framework310 → 640+330 (+106.5%)5.1
support ticket automation290 → 330+40 (+13.8%)7.1
crm for ai agents380 → 410+30 (+7.9%)6.0
sso saml providers110 → 120+10 (+9.1%)12.5

Queries losing clicks

QueryClicks (prev → now)ChangePosition
lumora reviews340 → 90−250 (−73.5%)8.4
customer support automation260 → 210−50 (−19.2%)14.0
workflow automation software190 → 150−40 (−21.1%)16.2

High impressions, low CTR

QueryImpressionsCTRPosition
customer support automation41,2000.5%14.0
workflow automation software38,9000.4%16.2
sso saml providers15,6000.8%12.5

Together these have 95,700 impressions (57.7% of the total) but only 480 clicks (15.3%). This is a ranking problem, not a snippet problem. Rewriting titles won't fix a page sitting at position 14 to 16. Your queries at positions 5 to 7 get about 4.2–4.3% CTR. If these three reached that range, that's roughly +3,300 clicks per 90 days, about double today's total.

Proposed fixes, ranked

  1. Get "customer support automation" onto page 1 (biggest win). Rebuild the page to cover the topic fully: use cases, how it works, ROI, comparisons, FAQ. Add internal links from the "ai agent platform" and "support ticket automation" pages, which already rank well. Also check whether the two support pages compete with each other: one is losing clicks while the other gains. Pick one target page per query.
  2. Do the same for "workflow automation software." This is a crowded category term, so aim for a comparison page (like "best workflow automation software for AI teams") rather than a generic feature page, and build a few links to it.
  3. Fix "lumora reviews" now. When people search your brand name plus "reviews" and your own result sits at position 8.4, review sites and forums are taking those clicks. Search it and see who's above you. Then:
    • publish a customer stories page with review schema
    • fill out and push your G2 and Capterra profiles
    • respond to any negative threads that rank Lost branded clicks are prospects deciding right now.
  4. Check that "sso saml providers" fits your product before investing. It's a technical-buyer query that may not match an AI agent platform. If it doesn't lead to signups, deprioritize it. If it does, add a focused comparison section and an FAQ with schema to try for a featured snippet.
  5. Protect the winners. "ai agent platform" and "agent evaluation framework" drove 720 of the new clicks. Refresh those pages, add FAQ schema and internal links to your conversion pages, and test titles to push CTR from 3.6–4.3% toward 6%+. At positions 4 to 5, that's realistic.

Next step: export the same 90-day window from Search Console with page and query columns, plus previous-period impressions and position. That would let me name the exact pages to fix and confirm whether the "lumora reviews" drop is a ranking loss or fewer searches. Want me to set that up as a recurring weekly SEO report?

GPT-6 AstraMedium effort

$0.50
70s · 100% scorereference price

Clicks rose 17.2%, but “lumora reviews” is the biggest concern: down 73.5%. The strongest gains came from “ai agent platform” and “agent evaluation framework.”

Data limitation: The export contains 8 queries—not page URLs. I can identify query-level opportunities, but cannot reliably name underperforming pages. It compares current versus previous 90-day clicks without exact dates or prior impressions/positions.

Queries gaining and losing clicks

QueryPrevious → current clicksChange
ai agent platform790 → 1,180+390 · +49.4%
agent evaluation framework310 → 640+330 · +106.5%
support ticket automation290 → 330+40 · +13.8%
crm for ai agents380 → 410+30 · +7.9%
sso saml providers110 → 120+10 · +9.1%
workflow automation software190 → 150−40 · −21.1%
customer support automation260 → 210−50 · −19.2%
lumora reviews340 → 90−250 · −73.5%

Across these rows: 3,130 clicks, up from 2,670. These are export totals, not necessarily sitewide totals.

High impressions, low CTR

Using 15,000+ impressions and CTR below 1% as a screening threshold:

QueryImpressionsCTR¹Avg. positionRecommended approach
customer support automation41,2000.5%14.0Improve relevance and ranking before focusing on snippet copy
workflow automation software38,9000.4%16.2Sharpen the landing page’s intent and differentiation
sso saml providers15,6000.8%12.5Check whether the query attracts the right audience

¹CTR as reported in the export.

These queries account for 95,700 impressions—57.7% of the listed total. Their average positions are outside the top 10, so low CTR alone does not establish a title or description problem.

Prioritized fixes

1. Diagnose the “lumora reviews” decline first.
It represents 250 of the 340 gross clicks lost across declining queries. Compare its landing URLs, impressions, positions, device and country splits across both periods. Check indexing, canonical changes and competing review results. If the relevant page is outdated, refresh it with verifiable customer evidence and clear product-fit information. The export cannot distinguish weaker demand from lost rankings or CTR.

2. Improve the pages serving the two broad automation queries.
Once URLs are mapped, assess intent alignment:

  • Customer support automation: strengthen concrete workflows, integrations, implementation guidance and substantiated results.
  • Workflow automation software: clarify supported workflows, capabilities and differentiation; add relevant examples or comparisons.

Support those pages with contextual internal links. Investigate overlapping URLs before consolidating anything. At positions 14–16, title changes alone are unlikely to solve the problem.

3. Protect and extend the two biggest winners.
“AI agent platform” and “agent evaluation framework” added 720 clicks combined, at positions 4.2 and 5.1. Keep what is working; test targeted improvements to query-aligned titles, answer sections, examples and internal links rather than broad rewrites.

4. Validate the SSO opportunity before investing.
“SSO SAML providers” may signal someone shopping for an identity provider—not an AI platform that supports SSO. If that is a mismatch, prioritize qualified SSO-integration queries rather than chasing impressions.

Next step: Export query × page data for both 90-day periods, including clicks, impressions, CTR and position. That will let us assign fixes to actual URLs and separate ranking losses from demand or snippet issues.

One example answer per model. Cost, time and score are averages over 5 runs.
GLM 5.3 Flash average score97.5%Opus 5.5 97.9% · GPT-6 Astra 100%
Cheaper than GPT-6 Astra41×$17 vs $709 per 1,000 tasks
Time to finish a task, GLM 5.3 Flash2.0 minOpus 5.5 70s · GPT-6 Astra 1.6 min

Same work, very different bill.

What 1,000 tasks cost on each model, in US dollars, at each provider's public prices.

Cost per 1,000 tasks

All teams

GLM 5.3 Flash
2% of the GPT-6 Astra bill
$17$0.017 per task41× cheaper
Opus 5.5
$234$0.23 per task33% of the GPT-6 Astra price
GPT-6 Astra
$709$0.71 per taskreference

Shorter bar, smaller bill.

What your team would spend

How many tasks does your team hand to AI each month?

1k2.5k5k10k25k50k100k
GLM 5.3 Flash$173/mo
Opus 5.5$2,337/mo
GPT-6 Astra$7,090/mo
$6,917/mo

saved every month by using GLM 5.3 Flash instead of GPT-6 Astra. That's $83,006 a year.

Quality vs. price

Each big dot is a model. Small dots are single tasks: hover to see which one, click for the details. Top left is where you want to be: high score and cheap.

70%80%90%100%$10$20$50$100$200$500$1,000$2,000Cost per 1,000 tasks, US dollarsAverage score80%Opus 5.5: Search Console Query Analysis, 100% score, $232 per 1,000 tasksGLM 5.3 Flash: Search Console Query Analysis, 98.2% score, $12 per 1,000 tasksGPT-6 Astra: Search Console Query Analysis, 100% score, $497 per 1,000 tasksOpus 5.5: Resume Screening, 100% score, $340 per 1,000 tasksGLM 5.3 Flash: Resume Screening, 95.3% score, $28 per 1,000 tasksGPT-6 Astra: Resume Screening, 100% score, $1,280 per 1,000 tasksOpus 5.5: Meddicc Call Coaching, 92.6% score, $220 per 1,000 tasksGLM 5.3 Flash: Meddicc Call Coaching, 95.8% score, $16 per 1,000 tasksGPT-6 Astra: Meddicc Call Coaching, 100% score, $345 per 1,000 tasksOpus 5.5: Content Cannibalization, 98% score, $187 per 1,000 tasksGPT-6 Astra: Content Cannibalization, 100% score, $1,050 per 1,000 tasksGLM 5.3 Flash: Content Cannibalization, 98% score, $14 per 1,000 tasksOpus 5.5: Job Description From Notes, 98.7% score, $190 per 1,000 tasksGLM 5.3 Flash: Job Description From Notes, 100% score, $16 per 1,000 tasksGPT-6 Astra: Job Description From Notes, 100% score, $374 per 1,000 tasksGLM 5.3 Flash97.5% · $17Opus 5.597.9% · $234GPT-6 Astra100% · $709

Where the low-cost model shines,
and where it struggles.

Click a team to see the whole page for that team, or open its task list.

TeamAverage scoreCost per 1,000 tasksYou save
GLM 5.3 FlashOpus 5.5GPT-6 AstraGLM 5.3 FlashGPT-6 Astra
Google Drive95.8%92.6%100%$16$34595%
Google Drive98.1%99%100%$13$77398%
Google Drive97.6%99.3%100%$22$82797%

The tasks

5 tasks taken from real work in Vybe, across sales, marketing and recruiting, from easy to hard. Every model ran each task several times.

Grading

Each task has its own checklist: the facts a good answer must include, the instructions it must follow, and the mistakes it must avoid, like inventing a detail or changing a file it was only asked to read. Every run starts from a fresh agent with the same data and tools.

A mix of deterministic checks and a separate AI judge reviews each item one at a time, reading the answer and the steps the model took to get there. A check passes only when it is clearly met. The score is the share of checks passed, averaged over the runs. Every check counts the same, and cost is not part of the score for the scope of this comparison.

Price and time

Prices are each provider's public rates, including every step the AI takes along the way. Time is measured from the request to the final answer, in Vybe.

Use the low-cost model in any Vybe agent, app or chat

2 min setup · No credit card required