The naive multi-provider approach looks like this.
You add a second LLM. You add an if provider == "openai" check in your service layer. It works. Then you add a third. The if becomes elif. You add cost tracking — now the cost logic branches too. You add caching — the cache key logic needs provider awareness. Six months later, every new feature requires touching five different if/elif chains, and adding a fourth provider means finding all of them.
If you are not a member of medium, check the article here This is the trap. And it’s avoidable.
When we added OpenAI and Anthropic to a translation API that started with Sarvam AI, we had one constraint: the translation service must not know which provider it’s using. The interface was the solution. Everything else — the factory, the registry, the per-client authorization — followed from that single decision.
The Interface: Make the Contract Explicit
Every provider must implement one contract:
Two decisions worth noting.
TranslationResult includes token counts even though Sarvam doesn't report them — it returns zeros. This keeps cost tracking uniform across all providers. The service layer always gets the same shape back regardless of which provider answered.
doc_type is part of the interface because it affects the prompt. Banking documents use different terminology than legal documents, and the LLM needs to know. It's not a detail that belongs in the service layer — it belongs in the contract.
Three Providers, One Shape
Here’s what each provider actually looks like under the hood:
These differences are real. OpenAI and Anthropic both accept system prompts but structure their APIs differently. Sarvam is a purpose-built translator — no system prompt, no token counts, a tight character limit per request. The interface absorbs all of it. The service layer sees none of it.
The Anthropic client, for example:
The Sarvam client returns the same TranslationResult shape with zeros for token fields. From the service layer's perspective, these are identical.
The Model Registry: Config as Data
Rather than scattering model configuration across client classes, everything lives in a central registry:
frozen=True matters here. Model configurations should never mutate at runtime. Frozen dataclasses prevent accidental mutations and are hashable, which helps with caching.
The cost fields earn their place too. Every API response includes an estimated_cost_usd computed in real time:
Clients see their per-request cost without needing their own pricing tables. When rates change, there’s one place to update.
The Factory: One Client Per Model
Clients are singletons keyed by model code. Each client maintains a connection pool and authentication state — creating a new one per request wastes resources and risks hitting rate limits during initialization. Class-level _clients persists across Lambda warm invocations, so the OpenAI client is created once per container, not once per request.
Python’s match/case (3.10+) is the right tool for factory dispatch. It's more explicit than if/elif chains and trivially extensible — adding a fourth provider is one new case clause.
Per-Client Authorization
Not every API client should have access to every model. Some clients are restricted to cheaper models. Some providers are off-limits for compliance reasons. We built this in from the start as a simple dictionary:
The client ID comes from the JWT token validated by API Gateway. In the endpoint:
This returns a 403 Forbidden with enough detail for the client to understand what went wrong and what they can use instead.
The lesson here: per-client authorization is cheap to build at the start and expensive to retrofit. A dictionary mapping is a five-minute addition. Doing it six months later, after clients have been calling the API without constraints, means auditing all existing usage first.
The Service Layer: No Provider Knowledge
With all this in place, the translation service looks like this:
The only place the service interacts with provider-specific details is model_config.max_chars_per_chunk for segment splitting. Everything else is abstracted. The service doesn't import any provider SDK. It doesn't branch on provider name. It calls translate() and gets back a TranslationResult.
That’s the goal. The interface is doing its job.
Adding a Fourth Provider
This is the real test of any abstraction. To add Google Translate as TX-04:
- Create google_client.py implementing BaseTranslationClient
- Add a TX-04 entry to MODEL_REGISTRY
- Add a "google" case to the factory's match statement
- Store the Google API key in Secrets Manager
- Update CLIENT_MODEL_MAPPING to authorize clients for TX-04
No changes to the translation service. No changes to the endpoint. No changes to the caching layer, the observability layer, or the middleware stack. The abstraction does what abstractions are supposed to do: contain change.
The Principle
Multi-provider LLM architecture isn’t inherently complex. It gets complex when provider-specific logic leaks into business logic — when your service layer starts caring about which model it’s calling, or your cost tracking branches on provider name, or your tests need to mock three different API clients.
The interface is the architecture. Everything else — the registry, the factory, the authorization mapping — is mechanical once the interface is right. Get the contract clear first. The rest follows.