You are building a production LLM pipeline where costs matter — a document processing API, a translation service, a data extraction system. You have compared provider pricing. You have a spreadsheet. It shows GPT at X cents per thousand tokens, Claude at Y, and a purpose-built model like Sarvam at what looks like a simple per-character rate.
if you are not a medium member, you can read the story at
That spreadsheet is missing at least three line items. The system prompt you send with every single call — repeated identically on every segment, multiplied by every document. The rate limit that prevents the concurrency your cost model assumed. The chunk size ceiling that turns a 5,000-character clause into six separate API calls. We ran these workloads in production across all three providers and logged everything. Here is what actually happened.
The Three Models — and Why They Are Not Directly Comparable
The first thing to understand is that per-token pricing and per-character pricing produce very different cost structures depending on your workload. You cannot put them side by side and declare a winner without running the actual numbers.
GPT and Claude charge per token and give you full usage data. Sarvam is a purpose-built translation model with no system prompt overhead — but it charges ₹20 per 10,000 characters (approximately $0.24 at current exchange rates), has a hard 999-character limit per call, and a 60-request-per-minute ceiling on the Starter plan. The character-based pricing looks simple, but the actual per-document cost surprises most teams when they run the numbers.
What One Document Actually Costs
Take a real workload: a banking sanction letter with 25 translatable segments, roughly 3,000 characters total, being translated to Hindi. First request, no cache.
For GPT-5.4 Mini, each segment call carries approximately 150 input tokens (system prompt plus the segment text) and returns around 100 output tokens. Across 25 segments run in parallel: 3,750 input tokens at $0.25/MTok and 2,500 output tokens at $2.00/MTok. Total: ~$0.006 per document. Latency: 300–500ms because all 25 calls fire concurrently.
For Claude Haiku 4.5, the token counts are similar but the rates are 4x higher on input and 2.5x higher on output. Total: ~$0.016 per document — 2.7x the GPT cost for the same work. Latency is comparable because the rate limit is the same and concurrency is identical.
For Sarvam, the pricing is ₹20 per 10,000 characters — pay-per-use, no subscription required on the Starter plan. For our 25-segment document with ~3,000 source characters: (3,000 / 10,000) × ₹20 = ₹6.00 ≈ $0.071 per document at current exchange rates. That is roughly 12x GPT and 4x Claude for the same workload. The 999-character limit also means some segments get split into multiple API calls, though since billing is per character, splitting does not increase cost — only latency. At 60 RPM, concurrent requests queue. Total: ~$0.071 per document, 500–1,500ms latency — Sarvam’s premium reflects Indian language quality, not cost efficiency.
Where Cache Changes Everything
The single-document numbers miss the most important effect. After the first few hundred documents, segment-level caching starts hitting. Banking documents share a large vocabulary of common phrases. We observed 40–60% cache hit rates after the initial warm-up period.
A cache hit costs a DynamoDB read — roughly $0.00000025. Effectively free compared to any LLM call.
Here is what that does to monthly costs at 1,000 documents per day:
Cache hit rate is a significant cost lever across all three providers. At 50% cache hits, GPT drops from $180 to $90 per month, Claude from $480 to $240, and Sarvam from $2,130 to $1,065. Sarvam remains the most expensive at every cache level — the per-character rate simply runs ahead of per-token costs at typical document volumes. The caching strategy still matters, but it does not close the gap between providers.
The Hidden Costs That Per-Token Pricing Ignores
This is where most cost estimates go wrong. The pricing page shows you input and output rates. It does not show you any of the following.
System prompt overhead. GPT and Claude require a system prompt on every single segment call. Our prompt is approximately 150 tokens. For a 100-segment legal document, that is 15,000 extra input tokens carrying the same instructions, identically, on every call.
- GPT overhead: 15,000 × ($0.25 / 1M) = $0.00375 per document
- Claude overhead: 15,000 × ($1.00 / 1M) = $0.015 per document
- Sarvam overhead: $0.00 — no system prompt exists
As a percentage of total cost, prompt overhead represents 30–50% of the per-document expense for LLM-based models. This is the strongest argument for keeping prompts short. Every token you add to the system prompt gets paid 100 times for a 100-segment document.
Chunk size multiplication. Sarvam’s 999-character limit means long segments are split at sentence boundaries, translated independently, and reassembled. A 5,000-character legal clause is one API call on GPT or Claude. On Sarvam, it is six calls — and each of those counts against the 60 RPM ceiling.
For a legal document with 100 segments where 10 are long clauses averaging 3,000 characters, GPT makes 100 API calls. Sarvam makes approximately 130. Since Sarvam bills by character, splitting does not multiply the cost — but at 60 RPM on the Starter plan, those extra calls push processing time from 1–2 seconds to 15–30 seconds. For latency-sensitive workloads, upgrading to the Pro plan (₹10,000, 200 RPM) or Business plan (₹50,000, 1,000 RPM) removes the bottleneck but adds upfront cost.
Latency as infrastructure cost. AWS Lambda bills by execution time. A GPT-translated document with 500ms of LLM calls costs roughly $0.000004 in Lambda compute. The same document routed through Sarvam at 1,500ms due to rate limiting costs $0.000012. This is negligible per request — but it compounds with rate-limit-induced latency, and at scale the Lambda bill becomes measurable.
Sanitization failures. GPT and Claude occasionally return formatting artifacts — code fences, smart quotes, preamble text — that require a sanitization layer and occasionally fall back to untranslated source text. Sarvam, being purpose-built, produces clean output with no sanitization needed. Fallback segments are not a financial cost, but they are a quality cost that can trigger human review, which absolutely has a financial cost.
When to Use Each Model
The decision is not about which model is cheapest. It is about matching the model’s properties to the workload.
The routing logic in practice:
Sarvam is the right choice when Indian language translation quality is the top priority and cost is secondary, or when operational simplicity matters — no system prompt to maintain, no sanitization layer needed, transparent per-character billing. At ~$0.071 per document it is the most expensive option, but for compliance-grade content, regulatory notices, or customer-facing material in Indian languages where quality directly affects trust, the premium may be justified.
GPT-5.4 Mini is the default for everything else — standard banking documents, notification templates, correspondence. Best cost-to-quality ratio for real-time translation with no chunk size concerns.
Claude Haiku 4.5 earns its 2.7x price premium when documents contain long clauses that would require multiple splits on Sarvam, or when translation quality directly affects downstream human review costs. Legal agreements and regulatory notices are the right use cases — not routine document processing.
The routing decision should be made at the client configuration level, not left to the caller at request time. Our system enforces this through per-client model authorization — a client approved for budget processing cannot accidentally switch to the premium model, and a compliance team that needs Claude cannot be downgraded without a configuration change.
The Number That Matters More Than Model Choice
After running this system in production, one metric has proven more predictive of cost than any other: cache hit rate.
Every request logs this data:
Aggregate this across 30,000 documents per month and the pattern is clear. The infrastructure — Lambda, DynamoDB, API Gateway, Secrets Manager — costs about $7.50 per month regardless of model. The LLM API cost dominates by an order of magnitude and responds almost linearly to cache hit rate.
A team optimizing model selection while ignoring cache hit rate is tuning the wrong variable — even for Sarvam at $0.071/document, a 50% cache hit rate halves your monthly bill. Build the segment-level cache first regardless of provider. Then optimise model selection.
The pricing page is not wrong — it just tells you the easy half of the story.