Cost of revenue: inference that delivers a paid product
When production inference is a direct input to a product a customer pays for, it belongs in cost of revenue, not operating expense. The cash cost is identical, but the line choice sizes AI gross margin and is what investors read for unit economics ASC 350-40-35. Misclassifying inference COGS as OpEx overstates gross margin and understates the cost of serving.
When inference is COGS
The test is directness: is the token spend a cost of delivering the specific product or service the customer pays for? If serving a customer request consumes tokens, that consumption is a cost of revenue. Internal or back-office use is operating expense instead.
No entry
Why the line matters
Cost per million tokens is the unit cost of the AI product. Booking it to COGS makes gross margin reflect the true cost of serving, which matters more for an AI business than for classic SaaS because the marginal cost of a request is real and variable.
Instruments and mechanics that land here
Primary sources
- [S4] Weaver: Navigating internally developed software costs: U.S. GAAP vs tax treatment (US GAAP)
- [S1] KPMG: Hot Topic: Accounting for internal-use software (ASC 350-40) (US GAAP)
Ledger current as of 2026-07-24. A position and a citation, not accounting advice. See how we cite.