More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Fast-growing AI companies fall into two camps: those marking up inference and those packaging it into outcomes. Markups follow cost-plus pricing, tacking a fixed margin—say 30%—onto raw API costs. You wrap a model in product features, a smooth workflow, a clean UI. But once inference commoditizes, customers spot cheaper raw calls and bypass you. Your premium shrinks until there’s no margin left.
Value-based pricing avoids that trap by charging for results, not compute. Sierra bills per resolved support ticket, nothing for failures. Devin sells “Agent Compute Units,” an abstraction like Databricks’ credits or Snowflake’s credits, so customers never see token prices. You set fees as a share of the surplus you create—say $5 per generated report or a percentage of cost savings—so your revenue stays detached from fluctuating inference rates.
Across both models, you can boost gross margin by cutting inference costs through tactics and model distillation. Caching and smart routing trim waste but are easy to copy. Distillation runs heavy “teacher” models on incoming requests, trains a smaller “student” under eight billion parameters, then deploys it on cheap hardware. That gives you a proprietary model competitors can’t immediately match, driving per-call costs from, for example, $1.00 down to $0.70.
When customers bring their own API key, cost-plus breaks down: they see your markup on their cloud bill and skip you. Value-based still works because they pay you for outcomes and manage raw inference separately. Optimization stays in play too—you charge a platform fee while routing, caching, or running distilled models under their key. At that point, you’re selling a software platform, not just token transactions.
Questions about this article
No questions yet.