LLMs

OpenAI cuts GPT-5.6 Luna pricing by 80%, introduces Fast mode for Sol

OpenAI reduces GPT-5.6 Luna costs by 80% and Terra by 20%, while launching Fast mode for GPT-5.6 Sol to deliver 2.5× faster inference at doubled pricing.

Last verified:

Bottom Line Up Front

OpenAI reduced GPT-5.6 Luna pricing by 80% and GPT-5.6 Terra by 20%, effective July 30, 2026, while introducing Fast mode for GPT-5.6 Sol that achieves 2.5× faster inference speeds at double the cost. The price cuts reflect efficiency improvements across model architecture, inference infrastructure, and serving optimization, enabling customers to right-size model selection to workflow requirements.

GPT-5.6 Luna’s Dramatic Price Reduction

According to OpenAI’s announcement, GPT-5.6 Luna—positioned as the family’s fastest and most affordable option—now costs 80% less than its prior pricing tier. The company describes Luna as capable of using tools and completing multi-step workflows, making high-volume AI applications more practical at scale. OpenAI claims Luna delivers performance comparable to frontier-class models from one year prior at roughly 6 cents per task. The pricing reduction also applies to usage counted against ChatGPT Work and Codex subscriptions, which means existing paid-tier users benefit from the lower rates on those integrations.

GPT-5.6 Terra’s Balanced-Model Repricing

OpenAI positioned GPT-5.6 Terra as its “balanced model for everyday work,” and reduced its pricing by 20%. Unlike Luna’s steeper cut, Terra’s 20% reduction reflects a middle-ground option between cost-minimization and maximum capability. The company frames these tiered options as enabling customers to “match intelligence to the outcome”—selecting the right model based on task stakes, error cost, urgency, and scale rather than defaulting to maximum capability across all workloads.

Fast Mode and Sol’s Speed-Cost Tradeoff

The API now features Fast mode, which replaces OpenAI’s prior Priority Processing tier. According to OpenAI, Fast mode delivers up to 2.5× faster inference speeds than Standard processing for GPT-5.6 Sol, at double the standard price. The company notes that Fast mode maintains intelligence parity with Standard Sol—faster execution comes at the cost of premium pricing, not reduced model quality. Fast mode is backward compatible, meaning existing API requests tagged as “priority” will automatically route to Fast mode without code changes.

Efficiency Gains Across Model Layers

OpenAI attributes the pricing decreases to improvements across three operational areas: the models themselves, inference infrastructure, and serving optimization. The company frames these as years of accumulated gains in “how our models are built, served, and put to work.” However, OpenAI does not disclose specific technical details about the efficiency improvements—no quantified reductions in parameter count, inference latency, or memory footprint are provided in the announcement.

Why This Matters

The pricing restructuring forces teams to be intentional about model selection within workflows. A coding task that previously defaulted to premium models now has economic incentive to route uncertainty-resolution steps (where maximum reasoning is valuable) to Sol, and well-specified implementation tasks to Luna. This granularity benefits customers with variable-compute workloads—batch processing, customer support automation, and content generation pipelines can now optimize cost without sacrificing quality on high-stakes stages.

Fast mode’s 2.5× speedup at 2× cost unlocks use cases where latency, not throughput, is the constraint—real-time chat, live code completion, and interactive agents benefit from near-instant responses even at premium pricing. Organizations with latency-sensitive deployments can now trade cash for response time, whereas before they lacked this explicit option on Sol.

The announcement does not quantify how many customers will shift to Luna or whether the 80% price cut cannibalizes Sol revenue. Downstream impact depends on whether Luna’s performance truly meets real-world expectations in high-volume production—early adopter feedback over the next quarter will validate whether the positioning sticks.

Frequently Asked Questions

What are the new prices for GPT-5.6 Luna and Terra?

According to OpenAI, GPT-5.6 Luna is now 80% cheaper and GPT-5.6 Terra is 20% cheaper as of July 30, 2026. Specific dollar amounts were not disclosed in the announcement.

What is Fast mode and how does it differ from Standard processing?

Fast mode replaces OpenAI's Priority Processing offering and delivers up to 2.5× faster speeds than Standard processing for GPT-5.6 Sol at twice the price, with no change in model intelligence.

How does Luna compare to earlier frontier models?

OpenAI states that Luna delivers performance comparable to models that were frontier-class one year ago, at roughly 6 cents per task and nearly nine times the speed.

Does Fast mode work with existing API code?

Yes, according to OpenAI, Fast mode is backward compatible—requests tagged with priority will automatically use Fast mode.

#gpt-5.6 #pricing #inference #openai