OpenAI lost roughly $5 billion in fiscal 2024 — more than the combined losses of Uber, Tesla, Amazon, and Spotify that same year. Yet ChatGPT Plus still costs $20 a month.

How both of these facts hold at once is the puzzle. A loss of that size should have forced a price increase long ago. If prices aren't being raised rather than can't be raised, the reason lives somewhere in the cost structure.

Where the Costs Pile Up — and Where They Flow

The single biggest line item in AI service costs is GPU compute. Text generation, image analysis, and voice processing all run on GPUs — and Nvidia controls more than 90% of that market.

In the first half of 2026, Nvidia's data center segment posted roughly $44 billion in quarterly revenue, 92% of the company's total sales, at a 78% operating margin. While OpenAI, Anthropic, Google DeepMind, and Mistral are all running losses, the company selling them chips is pulling a 78% margin. Those numbers show, before anything else, which side of the AI ecosystem the profits are actually flowing to.

Some analyses put the monthly compute cost of a single ChatGPT Plus user's conversation volume at tens of dollars — meaning OpenAI collects $20 while spending several times that. Nvidia's market capitalization stood at roughly $5.4 trillion as of September 2026. Where the value the AI industry has created is actually piling up shows up in that number.

A 78% margin naturally raises the question of whether competitors will erode it. High margins usually attract competition, and competition usually brings prices down. AMD has launched its MI300 series, and both Meta and Microsoft have shifted some workloads over to AMD.

Yet the market-share split has barely moved in years — still roughly 90% Nvidia to 6-7% AMD.

Why CUDA Slows Down Price Competition

The reason GPU prices resist falling has less to do with hardware than with software. CUDA is Nvidia's development ecosystem, built up over more than a decade. The optimized code inside PyTorch and TensorFlow — the training frameworks used across AI research — thousands of published pretrained models, and the experiment code attached to academic papers are all built on CUDA.

Even as AMD's chips close the gap on hardware benchmarks, porting an existing training pipeline to run on AMD means rewriting code for ROCm, AMD's software layer. At data-center scale, that work takes months to years. Training a large model can burn through billions of dollars in a single run, which makes swapping out infrastructure mid-project impractical. When switching costs run this high, price competition moves slowly.

One distinction complicates this picture. AI development work splits into "training" and "inference." Training — building or updating a model — is heavily CUDA-dependent. Inference — generating responses from an already-finished model — already has viable alternatives in production, such as Google's TPUs and Amazon's Inferentia. That's how Google runs Gemini inference on TPUs.

But a service like ChatGPT or Claude, which keeps updating its models, has to do both training and inference. Because training is hard to pull away from Nvidia, the core of the cost structure stays locked to CUDA.

When GPU prices don't come down easily, AI service costs don't either. This is where the original question shifts. Once you understand why prices aren't falling, the next question follows naturally: is a subscription price hike just a matter of time? And if so, when?

What's Keeping Subscription Prices Low For Now

Current subscription prices for ChatGPT, Claude, and Gemini are set below cost. Investor money is what covers the gap. OpenAI alone raised tens of billions of dollars in funding in 2025; Anthropic has raised billions from Google and Amazon.

Investors are willing to absorb these losses on the reasoning that this is a land-grab phase — that once the market consolidates, pricing power will follow. The scenario: lock in users with low prices now, then capture profitability once competition thins out. The $20 a month users pay today covers only part of the actual cost; investment capital fills in the rest.

How AI subscriptions stay priced below costGPU compute cost (Nvidia's 78% margin)AI services: cost > subscription revenueVC funding covers the lossesPrices frozen during the market-sharerace

The moment this flow breaks is the condition for a price increase.

The Conditions That Would Trigger a Price Hike

The pace of VC funding could shift first. Investors won't cover losses forever. As pressure mounts to prove profitability, AI services will have to either raise prices or scale back what they offer. OpenAI has already raised its ChatGPT Pro price, and Anthropic has adjusted Claude Pro pricing — both can be read as early signals of this shift. Base subscription tiers are still frozen, but the pattern of raising premium tiers first has already begun.

Consolidation of the competitive field is another path. As long as ChatGPT, Claude, Gemini, and Perplexity are all competing head-to-head, none of them can easily move first on price. They all run on Nvidia GPUs, so their cost structures look similar — whichever one raises prices first risks losing subscribers to the others. But if even one of them falls noticeably behind on quality or scales back its service, that opens room for the rest to raise prices.

Corporate financial filings offer a way to gauge how close these two conditions are. The annual (10-K) and quarterly reports from Google (Alphabet), Meta, and Microsoft show, in hard numbers, how much AI infrastructure capital expenditure has grown year over year, and whether AI-related revenue is keeping pace with that spending. A capex spike paired with sluggish AI revenue growth signals rising investment pressure. Conversely, once a quarter shows an AI segment turning operating-profit positive, that company gains room to hold the line on price without a hike.

What's Worth Doing Now, Before Prices Change

Exactly when prices will rise is impossible to know. But the numbers above make it hard to bet on them staying flat.

It helps to design your workflows now, while API prices are still low. Building automation pipelines today gives you room to optimize toward fewer tokens for the same output once per-token costs rise. Building that architecture from scratch after a price hike often costs twice as much in both money and time.

It's also worth checking each service's GPU dependence in advance. Google Gemini relies heavily on its own TPUs, so it may absorb an Nvidia price increase with less shock than other services. Amazon Bedrock includes models that support its Inferentia chips. By contrast, small and mid-size AI API providers that depend entirely on Nvidia GPUs pass through GPU price increases the fastest. A price gap is likely to open up down the line between services built by companies developing their own chips — Google, Amazon, Meta — and those that aren't.

Which service you build into your core workflow now quietly sets the price risk you'll be carrying later.