🔍 Read the full analysis: My September 2026 AI Stack, From First Build To Final Decision on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s September 29 AI stack assigns Claude Opus 5.5 to building and GPT-6.1 Sol to detailed review, with other models reserved for specific tasks. The source compares Artificial Analysis Intelligence Index v4.3.x scores and estimated task costs; those figures may not predict results on other workloads.
Meyer says six models fall within about 20 index points of one another, while the reported cost per task spans roughly 100 times. In his table, Opus 5.5 scores 58 at its top setting and costs $5.98 per task. GPT-6.1 Sol at xhigh scores 51 and costs $0.39. The report lists GPT-6 Luna at $0.07 per task, with a score of 37, and positions it for classification, extraction and routing.
The recommended division of work follows those comparisons. Meyer uses Opus 5.5 at high for features, APIs and refactors, and xhigh for more demanding architecture or migration work. He assigns Sol at high or xhigh to file-specific investigation and review, while Sonnet 5.5, Astra, Fable and Luna are alternatives for narrower tasks. He says Astra or Fable may be useful as a second opinion when Sol and Opus disagree.
The report also compares effort settings, which change both scores and costs. Opus at high scores 54 for a reported $1.82 per task; xhigh scores 56 at $3.46; max reaches 58 at $5.98. Meyer says Sonnet 5.5 at max costs $7.60 per task for a score of 56, and reports that it generated about 193,000 output tokens per task on the index. These are index-specific measurements, not a guarantee of equivalent costs in an individual deployment.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
How the Cost Curve Changes Model Choice
The comparison shifts the practical question from picking the highest-scoring model to deciding which model meets a team’s quality threshold at a manageable cost. If the reported costs transfer to a particular workflow, using a lower-cost model for routine review could make additional checks affordable. Meyer argues that a different model family reviewing code can provide a useful second perspective, while acknowledging that review quality depends on the requirements and evidence supplied.
That distinction matters for buyers because the most capable model in a benchmark may not be the most economical choice for every task. The source also cautions that model-token prices alone do not determine the cost of completed work: extra waiting or human review can erase savings. Its example on that point is described as illustrative, rather than measured evidence.
Benchmark Scores Behind the Stack
The report uses the Artificial Analysis Intelligence Index v4.3.x, a general capability index, alongside task-cost and output-token figures. It lists Opus 5.5 as released September 22, Sonnet 5.5 on September 28, and GPT-6.1 Sol on September 29. Fable 5.1 and Astra are listed as September releases; Luna is dated September 22.
Meyer’s table puts Astra at 53 points at its top setting and $3.26 per task, and Fable 5.1 at 53 points and $7.63. Sol xhigh is listed at 51 points and $0.39. The report says Sol’s high setting took 57 seconds to produce its first token and xhigh took 69 seconds, making those settings slower to respond. Meyer advises readers to shadow-test models before switching a workflow.
“The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.”
— Thorsten Meyer, in the September 29 report
Limits of the Index and Cost Estimates
The report does not establish how the six models perform on a reader’s own codebase, documents or evaluation criteria. Meyer says the index measures general capability and recommends shadow testing before a switch. The source does not provide details here on the index’s full methodology, sample variation or how its per-task cost estimates map to each provider’s billing for a specific user.
It also says one index point is within the noise, and that Artificial Analysis had not yet published low or max effort results for GPT-6.1 Sol at the time of writing. The reported high and xhigh first-token delays may also limit Sol’s fit for interactive work. The figures therefore support Meyer’s stated workflow, but do not settle which model is best for other teams.
Test Models Against Your Workload
Meyer’s immediate recommendation is to shadow-test candidate models on the work a team actually performs before changing its defaults. Teams following his approach would compare output quality, latency and total review effort, not benchmark score or token price alone. The source does not give a date for further benchmark updates or specify a future release milestone.
For GPT-6.1 Sol, the next relevant evidence would include additional effort-level results and tests on real development and review tasks. Until then, Meyer’s stack remains a dated recommendation based on the index and cost figures available on September 29, 2026.
Key Questions
What is the main recommendation in Meyer’s September 2026 stack?
He uses Claude Opus 5.5 for building and GPT-6.1 Sol for detailed investigation and review, with other models assigned to narrower tasks.
Why does Meyer use GPT-6.1 Sol for review?
The report lists Sol at high or xhigh at $0.32 to $0.39 per task, far below the listed costs for several higher-scoring alternatives. Meyer says that makes routine review more affordable; the figures are index estimates and may not match every deployment.
Does the index show which model is best for every team?
No. Meyer describes the Artificial Analysis Intelligence Index as a map of general capability, not a verdict on a particular workload, and recommends testing models against a team’s own tasks.
What limits does the report identify for GPT-6.1 Sol?
At high and xhigh, the reported time to first token is 57 to 69 seconds. Sol xhigh also scores below Opus 5.5 at xhigh in the listed comparison, and the source says low and max results were not yet published.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
