
AI gross margins typically run lower than classic SaaS margins because every task a customer runs costs model inference, which scales with usage. ICONIQ's 2026 State of AI report puts the average AI product gross margin at 45 percent in 2025 and projects 53 percent in 2026, against the 60 to 80 percent plus a16z describes for traditional software. In our view, a lower margin can still be fundable with a credible path up.
Definition: AI gross margin is revenue minus the cost of delivering the product (inference and model API fees, hosting, retrieval and data infrastructure, and any human review), divided by revenue.
In SaaS, the thousandth customer is nearly free to serve. In AI, the thousandth customer sends you a bill.
That path up usually runs through three things: model routing, falling token prices and a pricing model that tracks usage. Below we show how to count inference in COGS, work through a per-seat margin you can adapt, compare seat, usage and outcome pricing on the same numbers, and close with our view on margin versus growth. For the general formulas behind CAC, LTV and payback, see startup business models and unit economics; for which parts of an AI product are defensible, see AI startup moats.
Why AI gross margins run below SaaS
We'd be careful about holding an AI product to an 80 percent SaaS bar. Traditional software has close to zero marginal cost: one more user barely moves the hosting bill. AI is different. Every query, agent step and document burns tokens you pay for.
The published benchmarks show the gap:
| Benchmark | Gross margin | Source and date |
|---|---|---|
| Traditional software (SaaS) | 60 to 80 percent plus | a16z, "The New Business of AI" (Feb 2020) |
| AI companies, early wave | 50 to 60 percent | a16z, same essay (Feb 2020) |
| AI "Shooting Stars" | about 60 percent | Bessemer, State of AI 2025 (Aug 2025) |
| AI "Supernovas" (fastest growers) | about 25 percent, often negative | Bessemer, State of AI 2025 (Aug 2025) |
| AI products, average | 45 percent (2025), 53 percent projected (2026), 59 percent projected (2027) | ICONIQ, State of AI 2026 (Jul 2026) |
Two things jump out. The fastest-growing AI companies have the thinnest margins: Bessemer's Supernovas reached about $125 million of ARR in their second year at gross margins near 25 percent. And margins are getting better. ICONIQ surveyed more than 300 executives at software companies building AI products. Two thirds saw better per-query unit economics, helped by inference cost management and model routing. As products scale, inference takes a larger share of cost and talent's share falls.
Back in 2020, a16z found AI companies often spent 25 percent or more of revenue on cloud resources. In 2026 the equivalent line is the model API bill, and it can be much larger.
How to calculate AI gross margin with inference costs
The formula is the same as for any software company. The difference is how strictly you define cost of goods sold. One way to set it up (the variables are illustrative):
AI gross margin = (Revenue - COGS) / Revenue
COGS per unit = inference cost + hosting and retrieval + third-party data or tools
+ human review or support attributable to delivery
Inference cost per task = (input tokens x input price) + (output tokens x output price)
- savings from caching and batch discounts
Four counting conventions we find useful:
- Model API fees and GPU time are COGS, not R&D, when they serve paying customers. Training runs and experiments are usually R&D; production inference is COGS.
- Credits do not reduce COGS. Free cloud or model credits hide a cost. Many founders find it clearer to model the margin at list price.
- Measure per task, then per customer. Averages hide power users, so it often helps to track the 90th percentile account.
- Include agent overhead. Agents make many model calls per user request, and each call has a cost.
Worked example: per-seat AI gross margin with inference in COGS
Here's an illustrative example. A B2B contract-review agent sells for $200 per seat per month. The average seat runs 300 tasks a month. Each task sends about 60,000 input tokens (the contract plus instructions and retrieved context) and generates 6,000 output tokens across several agent steps. Other delivery costs (hosting, retrieval, monitoring, a share of support) come to $20 per seat. Prices are Anthropic's published API list prices as of September 2026: Claude Opus 5.5 at $4 input and $20 output per million tokens, Claude Haiku 4.5 at $1 and $5.
Scenario A, everything on the frontier model.
| Line | Average seat (300 tasks) | Power user (1,500 tasks) |
|---|---|---|
| Inference per task | $0.24 input + $0.12 output = $0.36 | $0.36 |
| Inference per seat | $108.00 | $540.00 |
| Other COGS | $20.00 | $20.00 |
| Total COGS | $128.00 | $560.00 |
| Revenue | $200.00 | $200.00 |
| Gross margin | 36 percent | minus 180 percent |
A 36 percent margin sits well below Bessemer's Shooting Stars and ICONIQ's average. And every power user loses money. That's a problem that gets worse the more customers love the product.
Scenario B, routing plus prompt caching. Send 70 percent of tasks (clause extraction, classification, summaries) to Haiku 4.5 and keep 30 percent (judgment calls, redlines) on Opus 5.5. Cache the shared half of each prompt; Anthropic's pricing page lists cache reads at $0.20 per million for Opus 5.5 and at a tenth of the base input price for Haiku 4.5. (For simplicity this ignores the premium on cache writes.)
| Line | Frontier task (cached) | Small-model task (cached) | Blended task |
|---|---|---|---|
| Input cost | $0.12 + $0.006 | $0.03 + $0.003 | |
| Output cost | $0.12 | $0.03 | |
| Cost per task | $0.246 | $0.063 | 0.3 x $0.246 + 0.7 x $0.063 = $0.118 |
The average seat's inference falls from $108 to about $35, total COGS to about $55, and gross margin climbs from 36 percent to about 72 percent. The power user still costs about $197 against $200 of revenue, a margin near 2 percent.
Routing fixed the average. It didn't fix the price.
AI pricing: seat vs usage vs outcome, and what each does to margin
One way to protect margin is to tie revenue to the thing that drives your cost. Keep the Scenario B costs, change only how you charge, and you get these illustrative results:
| Pricing model | Average account (300 tasks) | Power user (1,500 tasks) |
|---|---|---|
| Seat: $200 per month | Revenue $200, margin 72 percent | Revenue $200, margin 2 percent |
| Usage: $0.70 per task | Revenue $210, margin 74 percent | Revenue $1,050, margin 81 percent |
| Outcome: $1.50 per accepted review, 60 percent acceptance | Revenue $270, margin 79 percent | Revenue $1,350, margin 85 percent |
Usage pricing tends to protect margin because revenue rises with cost. Outcome pricing can earn more per task, but you carry the cost of failed attempts: at a 40 percent acceptance rate the same outcome price would earn $180 on 300 tasks. Seat pricing is the easiest to buy.
Our lean is a hybrid: a platform fee plus included usage and overage. Some founders prefer pure usage pricing for transparency; others keep seats because procurement is simpler. Either way, we treat pricing as something you revisit, not something you set once. The market seems to agree. ICONIQ's 2026 report found companies use about 1.7 pricing models on average, with consumption-based pricing rising from 35 to 42 percent in six months and outcome-based from 18 to 23 percent.
For the mechanics of choosing a pricing metric and packaging tiers, see how to price your product. Where people rather than models deliver much of the work, margins behave differently again; see AI-enabled services.
What falling token prices mean for your margin plan
Our read: plan on cheaper capability, not a cheaper frontier, and don't bet on a fixed timetable. Here's the evidence as we see it. Others may weigh it differently.
Capability gets cheap fast. The frontier, less so. a16z's "Welcome to LLMflation" (November 2024) measured roughly a 10x annual decline in the price of constant capability, from $60 per million tokens for GPT-3 in 2021 to $0.06 for a model with a similar benchmark score in 2024. Frontier list prices tell a more mixed story. Anthropic's pricing page lists retired Opus 4 and 4.1 at $15 input and $75 output per million tokens against $4 and $20 for Opus 5.5, a 3.75x cut. It also lists newer top-tier models at $10 and $50. The newest model isn't necessarily the cheapest option; older models tend to fall in price over time.
Industry forecasts are slower than the historical curve. Executives on the record in mid-2026 put long-term token prices at about a tenth of today's, with big moves over three to five years. That works out to roughly a 37 to 54 percent decline per year. Lin Qiao, co-founder and CEO of the inference company Fireworks AI, gives a similar estimate: about a 10x cost reduction in three years once chip and memory supply constraints ease, driving about 100x more usage.
So far, cheaper tokens have meant bigger bills. This is the Jevons paradox at work: when a resource gets cheaper, total consumption often rises faster than the price falls. OpenRouter's CEO reported that after OpenAI cut the price of one model by 10x in total on OpenRouter over about two weeks, its usage there grew about 13x. He also described the AI market as growing 10x to 15x a year, with customers repeatedly underestimating the inference they will need. Budget for more volume, not a smaller bill.
Your customers now watch token spend too. Venture commentators in August 2026 described companies that began the year rewarding heavy token use, then got bills of around $20,000 per employee. They responded with caps (limits of around $200 to $500 for non-engineers were floated) and moves to open-weight models. The expectation for 2027 is that the heavy-spending phase gives way to demands for proven ROI. That cuts two ways. Customers will scrutinize what they pay you, and your own team's AI tool spend is part of your burn. One way to look at it: treat each employee's cost as salary plus the inference they consume, and compare that with their output.
Most work can move below the frontier. Michael Mignano, a general partner at USV, estimates that about 80 percent of non-coding enterprise tasks, such as summarization and document drafting, can run on models that are not at the frontier, while coding still benefits from top models. That's the case for routing, and it's why we think Scenario B is plausible rather than optimistic, though results vary by product.
"Margins can wait. Growth can't."
That's the growth-first case, and it deserves a fair hearing. Optimizing margin too early slows the product. Qiao treats margin optimization as a constraint that slows innovation during hypergrowth, so Fireworks optimizes a system once it knows it will scale it. Several AI coding companies started with negative gross margins and heavily subsidized usage and still became very valuable.
But not every negative-margin company recovers. Some start with poor margins and end with them, and many investors argue that companies with strong gross margins tend to make better investments. A startup can also run out of time before it grows out of a margin problem, if a foundation model company or a bundled competitor squeezes it first. That's one reason we don't treat the model itself as a moat.
Frontier or open-weight? Moving to open-weight models cuts COGS, but it doesn't reliably improve the product, and engineer time may be better spent on growth. One investor estimate puts open-weight hosts at about 30 percent gross margins against about 70 percent for closed labs, which is part of why cheaper models are cheaper.
The timing math. Jean-Denis Greze, co-founder and CEO of Town, prices for the margin he expects to have later: a level that should give 20 to 30 percent margins in about 18 months, on the assumption that the cost of a given capability halves every nine to twelve months, while most of Town's work still runs on frontier models. Run that rule forward and the cost of waiting becomes clear. If token costs halve every nine to twelve months, a task costs 25 to 35 percent of today's price in 18 months. An illustrative product priced for 25 percent gross margin at that point, with tokens as its only cost, runs at roughly minus 110 to minus 200 percent today.
Where we land
Growth-first margins can be fundable. We'd want two things: a written margin bridge, and investors who accept the gap with their eyes open. Reasonable investors disagree on how wide that gap can be. Pricing for tomorrow's costs isn't a pricing detail. It's a funding decision.
Should AI startups buy or rent GPUs?
Most seed-stage teams rent. API access or cloud GPUs keep capital free and let you switch models. The math tends to change only for steady, heavy workloads. Speechify's published figures make the comparison concrete: an H100 costs about $30,000 to buy, while renting one runs about $3.50 to $5 an hour. At those rates a year of continuous rental costs about $30,700 to $43,800 (Speechify's own estimate rounds this to $35,000 to $50,000), so owning pays back within about a year at full use. Speechify owns capacity for its base load, rents for peaks such as back-to-school season, and co-locates memory with GPUs for large training runs.
One caution from Qiao: hardware generations now arrive faster than old depreciation schedules assumed, which changes the build-versus-buy math. It may be worth buying only for load you're confident you'll use, and depreciating conservatively.
AI gross margin checklist and common mistakes
A simple checklist before a board meeting or a raise:
- [ ] Inference, hosting, retrieval and delivery labor all sit in COGS at list price, with no credits netted off.
- [ ] Cost per task is tracked by model, with a routing plan and a target frontier share.
- [ ] Margin is shown for the average and the 90th percentile account.
- [ ] Pricing includes a usage component or cap, so power users cannot sink margin.
- [ ] A 12 to 24 month margin bridge shows how routing, caching, price declines and pricing changes get you to target.
- [ ] Your team's own AI tool spend is budgeted and tied to output.
Mistakes we'd steer you away from:
- Reporting margin before inference. Investors are likely to recompute it, and the gap can cost you trust.
- Assuming price cuts arrive on schedule. Build the plan on routing and pricing you control. Treat falling token prices as upside.
- Unlimited plans on a variable-cost product. Heavy users can turn a seat into a loss.
- Defaulting to the newest model. Newer isn't necessarily cheaper, and many tasks don't need it.
- Ignoring buyer ROI scrutiny. If 2027 is the year customers demand proven ROI, it may help to price against outcomes you can measure.
Investors often look hard at this bridge at Series A, where a clear path to healthy margins can matter as much as growth; see when to raise a Series A. The return math behind high AI valuations is in high seed valuations and venture return math.
Fixing a margin bridge usually means a harder sales conversation: moving customers from unlimited seats to usage or outcome pricing is selling value, not features. 1752vc's Accelerate program is built for founders at that point, combining a $100K investment (at a valuation cap of up to $3.5M) with founder-led go-to-market and sales training, plus access to 850+ investors who are likely to ask these margin questions.
The bottom line
A thin margin today isn't a death sentence. A thin margin with no plan is harder to fund. Know your cost per task, price so heavy users pay their way, and write down how the gap closes.
Token prices move on the market's schedule.
Your pricing moves on yours.
Key takeaways
- AI gross margins typically run below SaaS because inference scales with usage; ICONIQ puts the 2025 average at 45 percent, projected to reach 53 percent in 2026.
- Token prices keep falling, with forecasts of about 10x over three to five years, but so far cheaper tokens have meant more usage rather than smaller bills.
- In the illustrative worked example, routing 70 percent of tasks to a smaller model plus caching lifts margin from 36 to about 72 percent, while seat pricing still leaves power users near break-even.
- Usage or outcome pricing tends to protect margin when usage varies; ICONIQ's data suggests most companies now blend pricing models.
- In our view, growth-first margins can be fundable with a credible, written margin bridge, but not every negative-margin company recovers.
Frequently asked questions
There is no single bar in 2026, and views vary. ICONIQ's 2026 report puts the average AI product at 45 percent in 2025, projected to reach 53 percent in 2026, and Bessemer's 2025 report shows steady growers near 60 percent while the fastest growers sit near 25 percent. Many investors care most about the trend and a credible plan toward 60 to 70 percent.
In most cases, yes. Inference and model API fees that serve paying customers are generally treated as cost of goods sold, because they rise with each unit delivered. Training runs and internal experiments usually belong in research and development. It often helps to record inference at list price even while you use free credits, since credits hide the cost rather than remove it, and investors tend to adjust for them.
A common approach is to tie revenue to the cost driver. Usage-based pricing (per task, document or call) or a hybrid of a platform fee plus included usage and overage keeps margins stable when some customers use far more than others. Outcome pricing can earn more per task, but you pay for failed attempts, so it tends to work best where success rates are high and measurable.
Most forecasts suggest so, though not on a fixed schedule. Industry forecasts in 2026 put long-term token prices at about a tenth of today's over three to five years, and a16z measured roughly 10x a year for constant capability. The newest frontier models have not fallen as fast, so we tend to think routing work to cheaper models is a safer plan than waiting for the frontier to get cheap.
Many teams use both, routed by task. By one USV estimate, most non-coding enterprise work, such as summarization and drafting, runs fine below the frontier, while coding and complex judgment still benefit from top models. One approach is to measure quality per task with evaluations, move each task type to the cheapest model that meets the bar, and revisit the split as prices change.
Sources
- 20VC: Rory O'Driscoll and Jason Lemkin, weekly roundup (August 2026)
- 20VC: Alex Atallah, OpenRouter (August 2026)
- 20VC: Nikesh Arora, Palo Alto Networks, and the weekly roundup (June 2026)
- 20VC: Lin Qiao, Fireworks AI (July 2026)
- 20VC: Cliff Weitzman, Speechify, and the weekly roundup (September 2026)
- 20VC: Michael Mignano, USV (July 2026)
- 20VC: Jean-Denis Greze, Town (September 2026)
- Anthropic: Claude API Pricing
- a16z: The New Business of AI (and How It's Different From Traditional Software)
- a16z: Welcome to LLMflation, LLM Inference Cost Is Going Down Fast
- ICONIQ: 2026 State of AI Report, The Builder's Economy
- Bessemer Venture Partners: The State of AI 2025
Disclaimer: This guide is for general education only and is not legal, tax or investment advice. Laws, market data and program terms change, so it may not reflect the latest developments or fit your situation. Treat it as a starting point, not a source of truth, and talk to a qualified lawyer, accountant or financial adviser before you make decisions.


