How We Unified Three LLM Providers Behind One Interface
Ganta Hemanth
The naive multi-provider approach looks like this.
You add a second LLM. You add an if provider == "openai" check in your service layer. It works. Then you add a third. The if becomes elif. You add cost tracking — now the cost logic branches too. You add caching — the cache key logic needs provider awareness. Six months later, every new feature requires touching five different if/elif chains, and adding a fourth provider means finding all of them.
If you are not a member of medium, check the article here
This is the trap. And it’s avoidable.
When we added OpenAI and Anthropic to a translation API that started with Sarvam AI, we had one constraint: the translation service must not know which provider it’s using. The interface was the solution. Everything else — the factory, the registry, the per-client authorization — followed from that single decision.
The Interface: Make the Contract Explicit
Every provider must implement one contract:
Base translation client interface contract
Two decisions worth noting.
TranslationResult includes token counts even though Sarvam doesn't report them — it returns zeros. This keeps cost tracking uniform across all providers. The service layer always gets the same shape back regardless of which provider answered.
doc_type is part of the interface because it affects the prompt. Banking documents use different terminology than legal documents, and the LLM needs to know. It's not a detail that belongs in the service layer — it belongs in the contract.
Translation client architecture diagram
Three Providers, One Shape
Here’s what each provider actually looks like under the hood:
Comparison table of OpenAI, Anthropic, and Sarvam translation providers
These differences are real. OpenAI and Anthropic both accept system prompts but structure their APIs differently. Sarvam is a purpose-built translator — no system prompt, no token counts, a tight character limit per request. The interface absorbs all of it. The service layer sees none of it.
The Anthropic client, for example:
Anthropic translation client implementation example
The Sarvam client returns the same TranslationResult shape with zeros for token fields. From the service layer's perspective, these are identical.
The Model Registry: Config as Data
Rather than scattering model configuration across client classes, everything lives in a central registry:
Central model registry configuration
frozen=True matters here. Model configurations should never mutate at runtime. Frozen dataclasses prevent accidental mutations and are hashable, which helps with caching.
The cost fields earn their place too. Every API response includes an estimated_cost_usd computed in real time:
Estimated cost calculation
Clients see their per-request cost without needing their own pricing tables. When rates change, there’s one place to update.
The Factory: One Client Per Model
Translation client factory implementation
Clients are singletons keyed by model code. Each client maintains a connection pool and authentication state — creating a new one per request wastes resources and risks hitting rate limits during initialization. Class-level _clients persists across Lambda warm invocations, so the OpenAI client is created once per container, not once per request.
Python’s match/case (3.10+) is the right tool for factory dispatch. It's more explicit than if/elif chains and trivially extensible — adding a fourth provider is one new case clause.
Per-Client Authorization
Not every API client should have access to every model. Some clients are restricted to cheaper models. Some providers are off-limits for compliance reasons. We built this in from the start as a simple dictionary:
Per-client model authorization mapping
The client ID comes from the JWT token validated by API Gateway. In the endpoint:
Authorization validation in API endpoint
This returns a 403 Forbidden with enough detail for the client to understand what went wrong and what they can use instead.
The lesson here: per-client authorization is cheap to build at the start and expensive to retrofit. A dictionary mapping is a five-minute addition. Doing it six months later, after clients have been calling the API without constraints, means auditing all existing usage first.
Translation service architecture overview
The Service Layer: No Provider Knowledge
With all this in place, the translation service looks like this:
Translation service abstraction example
The only place the service interacts with provider-specific details is model_config.max_chars_per_chunk for segment splitting. Everything else is abstracted. The service doesn't import any provider SDK. It doesn't branch on provider name. It calls translate() and gets back a TranslationResult.
That’s the goal. The interface is doing its job.
Adding a Fourth Provider
This is the real test of any abstraction. To add Google Translate as TX-04:
  1. Create google_client.py implementing BaseTranslationClient
  2. Add a TX-04 entry to MODEL_REGISTRY
  3. Add a "google" case to the factory's match statement
  4. Store the Google API key in Secrets Manager
  5. Update CLIENT_MODEL_MAPPING to authorize clients for TX-04
No changes to the translation service. No changes to the endpoint. No changes to the caching layer, the observability layer, or the middleware stack. The abstraction does what abstractions are supposed to do: contain change.
The Principle
Multi-provider LLM architecture isn’t inherently complex. It gets complex when provider-specific logic leaks into business logic — when your service layer starts caring about which model it’s calling, or your cost tracking branches on provider name, or your tests need to mock three different API clients.
The interface is the architecture. Everything else — the registry, the factory, the authorization mapping — is mechanical once the interface is right. Get the contract clear first. The rest follows.