Api-Management
Filed under
One Door to the Models, Part 8: Orchestration on Top of the Gateway
An orchestration framework that retries and reroutes is competing with the gateway. Part 8 keeps the framework thin and puts tools under the same identity.
One Door to the Models, Part 6: Semantic Caching and Its Failure Modes
A semantic cache is a correctness surface, not just a cost lever. Part 6 tunes the threshold, isolates tenants, and handles the day the cache is gone.
One Door to the Models, Part 5: Identity, Quota, and Chargeback
You get five custom metric dimensions, and time series multiply. Part 5 builds tenant identity, quota, and a chargeback model that survives cardinality.
One Door to the Models, Part 3: The Provider Abstraction and Streaming
Set stream to true and token counting becomes estimation. Open a WebSocket and load balancing stops existing. Part 3 maps guarantees to transports.
One Door to the Models, Part 2: Terraform, Bicep, or ARM
Bicep has no state file and cannot touch Entra ID. Part 2 picks the IaC layer for the gateway, and finds the model-version default that upgrades production.
One Door to the Models, Part 1: The Case for a Central LLM Gateway
A company runs five GenAI apps and cannot say what any of them cost. Part 1: the scenario, build versus buy on Azure, and the gateway this series builds.