Skip to main content

Command Palette

Search for a command to run...

The cost asymmetry most product teams ignore: input vs. output pricing

Updated
•3 min read•View as Markdown
A
I lead the strategy and delivery of products and systems that solve real problems, support commercial goals and scale across teams and markets. My work spans platform foundations, user-facing experiences, commercial tools and developer products, and I adapt my approach to the needs of each domain. My experience covers early-stage, scale-up and enterprise product environments, combining hands-on delivery with team leadership and award-winning work recognised across several innovative organisations.

When you pay for AI (LLMs), you may be paying for two different things.

You pay to feed the model information. And you pay for what it writes back.

Because it is not super clear, people treat these as the same thing or just ignore. Input and output tokens are priced differently, and the gap between them helps you understand your token cost.

Simple way to think about it.

Imagine you're hiring a consultant. You pay them to read a brief. Then you pay them to write a report. Reading is faster and cheaper than writing. Writing takes time, effort, and fills pages.

LLMs work the same way. Reading your input is the fast part. Writing the response is the expensive part. That's why output tokens cost more, sometimes two to three times more, than input tokens across most frontier models.

The model reads everything you send in one parallel pass. Fast. Then it writes the response one word at a time, each word depending on the previous one. Slow, sequential, and billed accordingly.

Product teams should model this split and not look at a blended price per token alone.

The interesting part

A new category of model is emerging that flips this entirely. Jev from TypeSafe is the most talked-about example right now. It doesn't write anything back to you. You give it a question with options, is this support ticket urgent or not, which category does this message belong to, score this response from 0 to 1, and it returns a typed decision as a choice, a number, or a probability. Not a sentence or paragraph. Just the answer.

Because there's nothing to write, output is free. TypeSafe's own phrase is "too cheap to meter." You pay only for input at $0.042 per million tokens i.e roughly 1/48th of what frontier models charge on the input side, let alone the output.

Featherless followed up with Simple Jev, an open-source version that goes further: input at $0.03 per million, output free, and public endpoints you can test without even creating an account.

The pattern is emerging. When a model doesn't generate text, output pricing disappears as a cost category entirely.

Why should product teams care?

Because a lot of what AI does inside a product doesn't need to generate text. Classifying a request before routing it to the right model. Deciding whether a response meets quality threshold before sending it to a user. Scoring inputs for relevance, urgency, sentiment or guardrailing outputs.

All of these are decisions, not documents. And until recently, product teams were paying frontier LLM prices to make those decisions, because those were the only models available.

That's changing fast. Cheap decision-layer models sitting in front of your main model stack means you can route intelligently, evaluate quality, and moderate outputs at a cost so low it barely shows up in your budget. The savings will likely come from what you route away from the expensive model, not from what the classifier itself costs. You need to watch out for model performance as over optimisation could lead to poor user experience.

A

The decision-versus-document distinction gives a concrete way to revisit a product cost model. One cost I would include beside the classifier request is the input passed downstream: if routing requires sending the same large context to both the decision model and the selected generator, the extra reading can offset part of the savings.

That suggests benchmarking a smaller routing payload containing only the fields needed for the decision, then checking disagreements against the full-context version. The useful product metric would be cost per successfully handled request, including unnecessary escalations and missed escalations, rather than savings calculated from classifier pricing alone.