🔍 Read the full analysis: OpenAI’s Cost-Effective GPT‑6 Sol And Luna: Benchmark Scores Stay Flat on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI has introduced GPT‑6 Sol and Luna models that maintain previous performance benchmarks while reducing costs by 50%. The models focus on affordability, potentially expanding automation and AI deployment options, despite no immediate performance improvements.
OpenAI has released new versions of its GPT‑6 family, GPT‑6 Sol and GPT‑6 Luna, which maintain benchmark scores but are priced at roughly half their predecessors, GPT‑5.6. This shift underscores a focus on cost efficiency rather than performance gains, making AI more accessible for a broader range of applications.
Both models were launched on September 22, 2026, with OpenAI emphasizing improvements in caching and inference that enable the models to be served at lower costs. GPT‑6 Sol’s input cost per 1 million tokens is $2.00, and output is $10.00, a 50% reduction from GPT‑5.6, while Luna’s costs are $0.10 and $0.50 respectively, roughly 50-60% less than its predecessor. Despite the price cuts, benchmark scores on the Artificial Analysis Intelligence Index show that Sol scores 48, and Luna scores 37, both significantly above their respective medians, but no performance improvements are claimed or observed in these scores.
Independent evaluations indicate that the cost per task has halved, with artificial analysis noting that the models’ scores on various benchmarks remain flat or show minor regressions, especially in knowledge and presentation quality. The models also demonstrate reduced hallucination rates, with Sol’s hallucination rate dropping from 92% to 60%, and Luna’s from 93% to 77%. However, this reduction partly results from the models declining to answer more questions, which can impact usability depending on the use case.
GPT‑6 Sol and Luna: half the price, about the same intelligence
OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.
Per 1M input / output tokens. Cached input reads keep the 90% discount.
Cost per task, halved
Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.
The effort dial moves cost more than the model choice
| Model and effort | Intelligence Index | Cost per task |
|---|---|---|
| GPT‑6 Sol (max) | 48 | $1.06 |
| GPT‑6 Sol (low) | 34 | $0.13 |
| GPT‑6 Luna (max) | 37 | $0.07 |
| GPT‑6 Luna (low) | 21 | $0.0045 |
| GPT‑6 Luna (non‑reasoning) | 18 | $0.01 |
Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.
What got better, and what got worse
Better
- Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
- Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
- OpenAI reports about half as many factual mistakes for Sol as its predecessor
- Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing
Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.
Worse
- GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
- AA‑Briefcase v1.1: Luna down ~45 Elo
- Coding Agent Index: Luna 41, down 2 points
- Both models write more output tokens per task than their predecessors
Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.
What to do about it
Implications of Cost-Effective AI Deployment
The release of GPT‑6 Sol and Luna at half the previous price points could significantly expand AI adoption in business workflows, automation, and research tasks. Companies can now consider integrating powerful language models without the previous cost barriers, potentially increasing AI-driven productivity and innovation. However, the unchanged benchmark scores suggest that these models are not designed to push the performance frontier but to democratize access, which could shift market dynamics and competitive strategies among AI providers.
For users, the key takeaway is that cost savings may come with trade-offs in presentation quality and knowledge accuracy, especially in complex tasks. The reduced hallucination rates are promising, but the models’ lower performance on certain knowledge benchmarks indicates that they may require careful testing before deployment in high-stakes environments.
As an affiliate, we earn on qualifying purchases.
Background on OpenAI’s Model Pricing and Performance
Prior to the September 2026 release, OpenAI’s GPT‑5.6 models represented the state of the art in performance but came with high costs, limiting widespread adoption. The company’s focus on improving inference efficiency and caching mechanisms has now enabled a new generation of models that deliver similar performance at a substantially lower price. This move aligns with industry trends toward making large language models more affordable and accessible, especially for enterprise applications.
The launch follows OpenAI’s strategy of offering a tiered model family, with Astra positioned at the top for high-quality, high-cost tasks. The new Sol and Luna models are positioned as more affordable options, aiming to broaden the user base and enable more tasks to be automated at lower costs.
cost-effective AI chatbot software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Long-Term Performance
It is still unclear how these models will perform in real-world, high-stakes applications over time, especially regarding their ability to maintain accuracy and handle complex tasks. The observed regressions in some knowledge benchmarks suggest potential limitations that could affect certain use cases. Additionally, the impact of reduced presentation quality on workflow productivity remains to be fully assessed.
Further testing and user feedback are needed to determine whether these models can replace more expensive options in various operational contexts.
AI development tools for small businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps in OpenAI’s Model Strategy
OpenAI is likely to continue refining caching and inference techniques to further reduce costs and improve performance. They may also release updated versions that address the regressions observed in knowledge tasks and presentation quality. Industry analysts expect that user testing and real-world deployments over the coming months will provide clearer insights into the models’ practical effectiveness and limitations.
Additionally, competitors may respond with their own cost-efficient models, intensifying market competition and innovation in affordable AI solutions.
affordable AI text generation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main advantage of GPT‑6 Sol and Luna?
The main advantage is their significantly lower cost—about half that of previous models—while maintaining comparable benchmark scores, making AI more accessible for automation and large-scale deployment.
Do these models perform better than previous versions?
No, their benchmark scores are roughly the same or slightly lower in some areas. The key improvement is in cost efficiency, not raw performance.
How do the models reduce hallucinations?
They achieve lower hallucination rates partly by declining to answer more questions, which reduces false or hallucinated responses but may also limit their usefulness in some contexts.
Will these models replace higher-end models like Astra?
They are designed as more affordable options, not replacements for top-tier models. Astra remains the choice for tasks requiring the highest quality results.
What are the potential risks of deploying these models?
Potential risks include reduced accuracy in knowledge-heavy tasks and lower presentation quality, which could impact decision-making or customer-facing applications if not carefully tested.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
