|

GPT-6 Sol And Luna Debut With 50% Price Cuts, Stronger Caching, But Independent Tests Show Mixed Results

GPT-6 Sol And Luna Debut With 50% Price Cuts, Stronger Caching, But Independent Tests Show Mixed Results
GPT-6 Sol And Luna Debut With 50% Price Cuts, Stronger Caching, But Independent Tests Show Mixed Results

OpenAI has formally launched two new fashions, GPT-6 Sol and GPT-6 Luna, increasing the GPT-6 household alongside the flagship GPT-6 Astra launched earlier this month. According to the corporate, the brand new fashions ship near-Astra-level efficiency in skilled work, factuality, coding, and laptop use, whereas reducing API costs by 50% in comparison with the promotional pricing of the earlier GPT-5.6 technology. The firm attributes the fee discount to enhancements in caching and inference infrastructure, the financial savings from which it says are being handed on to customers.

Under the brand new pricing, GPT-6 Sol prices $2 per million enter tokens and $10 per million output tokens, down from $4 and $20 respectively, whereas GPT-6 Luna is priced at $0.10 per million enter tokens and $0.50 per million output tokens, down from $0.20 and $1.20. OpenAI positions Astra as its uncompromising finest mannequin for probably the most demanding tasks, whereas Sol and Luna are designed to make superior AI sensible for higher-volume, on a regular basis workloads.

Benchmark Results and Technical Improvements

On AutomationBench, which exams brokers throughout 47 enterprise instruments in gross sales, advertising, operations, assist, finance, and HR, GPT-6 Sol at xhigh effort scored 33.2% at $0.27 per activity, outperforming Claude Opus 5 at max effort (26.9%) at roughly 9% of its price per activity. On Agents’ Last Exam, Sol at max effort reached 56.4%, exceeding Claude Opus 5’s highest rating at about 60% decrease price. For coding, GPT-6 Sol scored 68.8% on DeepSWE v1.1, inside 1.1 share factors of Claude Fable 5’s finest consequence at roughly 80% decrease price, whereas Luna’s 66.6% was corresponding to Opus 5 and Fable 5 at medium effort at 93–96% decrease price. On OSWorld 2.0, Sol matched Claude Opus 5 (60.5% versus 60.3%) at about one-fifth of the fee.

OpenAI additionally stories halved factual error charges for Sol relative to its predecessor on its inside analysis of de-identified conversations the place customers flagged errors, with Luna at greater effort matching GPT-5.6 Sol’s reliability at roughly one-hundredth of the fee. Alignment evaluations present each fashions enhancing over their predecessors, together with decrease charges of deception in intentionally difficult coding eventualities.

Beyond pricing, OpenAI has improved immediate caching to boost cache hit charges by default, providing a 90% low cost on cached input-token reads, alongside a monitoring dashboard, diagnostics instruments, and controls that permit builders alter reasoning effort and gear availability with out invalidating cached context. GitHub stories these adjustments have reduce the share of immediate tokens requiring contemporary processing by greater than 50% throughout billions of requests.

GPT-6 Sol and Luna can be found beginning in the present day in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu customers, with Luna additionally accessible to Free and Go customers through the desktop app; each are provided within the API as gpt-6-sol and gpt-6-luna, with a gradual rollout all through the day.

Independent Analysis Confirms Cost Gains, Mixed Results

Third-party analysis by Artificial Analysis largely corroborates OpenAI’s effectivity claims whereas portray a extra nuanced image of the fashions’ capabilities. Running its Intelligence Index, the agency discovered that GPT-6 Sol at max effort prices roughly $1.06 per activity, roughly half the $1.99 of its predecessor, whereas Luna drops from $0.18 to $0.07 per activity — inserting each releases firmly on the cost-efficiency Pareto frontier. Notably, the financial savings stem fully from the value reduce, as each fashions devour barely extra output tokens per activity than their predecessors.

Performance, nonetheless, is a mixture of progress and regression. In the Coding Agent Index, Sol improved by 2 factors to 57, with positive factors in Terminal-Bench 4.0 and SWE-Atlas-QnA, whereas Luna misplaced 2 factors, falling behind in SWE-Atlas-QnA and DeepSWE v1.1. Hallucination charges fell sharply on the AA-Omniscience benchmark — from 92% to 60% for Sol — although this was partly achieved by declining to reply extra questions, which lowered Sol’s accuracy by 5 factors. The most vital regressions appeared in knowledge-work evaluations: Sol dropped roughly 100 Elo factors in GDPval-AA v2.1 and Luna about 75, with handbook inspection attributing the declines to shorter deliverables that omit required rubric parts and lowered presentation high quality.

The submit GPT-6 Sol And Luna Debut With 50% Price Cuts, Stronger Caching, But Independent Tests Show Mixed Results appeared first on Metaverse Post.

Similar Posts