The maker of ChatGPT has slashed the price of its GPT-5.6 Luna model by 80% and reduced the price of its mid-tier Terra model by 20%, with the flagship Sol model remaining at its current price.
These price adjustments reflect increasing scrutiny regarding AI spending as companies transition from flat-rate subscriptions to usage-based pricing, where costs are determined by the number of “tokens” processed by AI models.
Although token prices have generally decreased over the last year, the cost of executing tasks has risen, as workloads now require more tokens, making AI-related expenses less predictable.
Price battle
Under the new structure, OpenAI will charge 20 cents per million input tokens for Luna, down from $1, while output token pricing decreases to $1.20 from $6. The input price for Terra drops to $2 from $2.50, and output pricing is reduced from $15 to $12.
In comparison, Anthropic’s mid-tier Claude Sonnet 4.6 is priced at $3 per million input tokens and $15 per million output tokens, making OpenAI’s Terra more affordable on both counts.
OpenAI attributed the lower prices to efficiency improvements in GPT-5.6, which include advancements in coding and internal performance optimization. The company noted that Luna and Terra can now compete in tasks that previously required top-tier models, allowing businesses to achieve similar results at a reduced cost.
China challenge
This pricing strategy emerges as OpenAI and Anthropic encounter intensified competition from cost-effective Chinese open-source models like Z.ai’s GLM-5.2. Analysts indicate these models nearly match the performance of top US AI models while being more economical.
Analysts believe the price reductions could facilitate broader adoption of OpenAI’s models and bolster its position in the enterprise sector, but they may also impact the financial health of both OpenAI and Anthropic ahead of their expected initial public offerings.
Token rethink
The broader AI sector is reevaluating how businesses utilize AI following surging token costs, which sparked backlash against “tokenmaxxing”—maximizing AI usage in work settings.
Tokens are the units AI models use for processing and generating text, with each token equating to roughly three-quarters of a word. Under usage-based pricing, businesses incur costs based on the number of tokens their AI workloads consume.
Earlier this year, OpenAI CEO Sam Altman expressed enthusiasm for the innovations that “tokenmaxxing startups” could generate. Nvidia CEO Jensen Huang also noted that “if your $500K engineer isn’t burning $250K in tokens, something is wrong,” while Meta organized an internal competition to reward token utilization.
However, as AI-related expenses escalated, many companies determined that increased token consumption did not necessarily lead to enhanced productivity.
Smarter AI
Vincent Gusdorf, Head of AI Analytics at Moody’s Ratings, indicated that companies have become more judicious after recognizing the costs associated with scaling AI deployments.
Bain & Company reported that token costs for some large enterprises have nearly doubled every couple of months, prompting businesses to analyze their AI investment returns more closely.
Rather than relying on premium AI models for all tasks, numerous organizations are implementing “model routing,” where simpler requests are managed by lower-cost models and more complex tasks are assigned to advanced systems.
This transition has also heightened interest in open-source AI models from Chinese startups like Moonshot and Zhipu, which offer capabilities nearly comparable to leading US systems at lower operating expenses.
Mozilla Chief Technology Officer Raffi Krikorian noted that productivity measured by token usage resembles the software industry’s outdated practice of assessing programmers by the quantity of lines of code produced—an approach that ultimately fell out of favor for failing to accurately represent real productivity.
(With input from agencies)