Six Real-Time Patterns in Azure: What the Diagram Leaves Out

I see a diagram of real-time communication patterns pop up in my feed every few weeks. Polling, long polling, Server-Sent Events, WebSockets, webhooks, gRPC streaming. Six neat boxes, and one piece of sensible advice underneath: choose the simplest model that fits the use case.

The diagram is correct. However, it is not the part that costs you a sprint.

Picking a pattern takes about five minutes. Making that pattern survive API Management, Azure Front Door, a scale-out event, and a rolling deployment takes considerably longer. This post covers that second part. It also comes with a working sample that runs all six patterns from one image, against one shared event source. Azure needs two apps to host it, and that’s one of the more interesting findings below.

Four questions, not six boxes

A grid of six options invites you to shop. Instead, answer four questions in order. The pattern then picks itself.

Which direction does data flow? The client pulls, the server pushes, or both sides talk at once. This single question removes at least three options.

What is the latency budget? Seconds are cheap. Milliseconds are not. Most business dashboards tolerate a delay that their designers never measured.

Who owns the connection when it drops? Someone must handle reconnect, resume, and replay. If your answer is “the browser does that automatically”, read the reconnect semantics again first.

Who pays per open socket? Idle connections still consume replicas, units, and money.

When two patterns both fit, pick the one that keeps state out of the connection. A dropped request costs a retry. A dropped session costs a reconnect, a replay, and a support ticket.

The six patterns, and what bites in Azure

Polling

The client asks again on its own schedule. Azure itself uses this pattern constantly. The asynchronous request-reply pattern returns 202 Accepted with a Location header, and Durable Functions exposes a status query endpoint that works exactly this way.

What bites: the missing Retry-After header. Without it, every client invents its own interval. Ten thousand clients then converge on the same second after a deployment, and your scale rule reacts to a spike you created yourself.

Long polling

The server holds the request open until data arrives or the clock runs out. Azure Service Bus applies the same idea inside its SDK, where a receive call waits for a configurable maximum wait time instead of returning empty.

What bites: timeouts you do not control. App Service and Azure Functions enforce a fixed 230-second ceiling on HTTP requests. That number comes from the Azure Load Balancer underneath, which idles connections out at 240 seconds by default. You cannot raise it because it is a platform constraint, not an application setting. Front Door then applies its own origin response timeout on top. Therefore, derive your hold time from the shortest timeout in the path, not from the longest. The sample holds for 25 seconds, which clears every layer.

Server-Sent Events

The server pushes over one long-lived HTTP response. SSE is unfashionable and quietly excellent. It runs over plain HTTP, it survives proxies that understand chunked responses, and browsers reconnect on their own.

You already depend on it. Token streaming from Azure OpenAI and Microsoft Foundry arrives as Server-Sent Events. Every chat interface you have built this year uses this pattern, whether or not the architecture diagram says so.

What bites: buffering, which gets its own section below. Also reconnect gaps. The browser resends the Last-Event-ID header automatically, but the server must honor it. Otherwise, every reconnect silently drops the events that arrived while the socket was down.

WebSockets

Both sides talk over one persistent connection. This is the right answer for chat, collaborative editing, and live trading. It is the wrong answer for a dashboard that changes twice an hour.

Azure gives you two managed options. That is Azure Web PubSub, which handles raw WebSocket clients and works well outside .NET. And Azure SignalR Service fits when you already use hubs and want fallbacks.

What bites: the deployment. A rolling revision in Azure Container Apps drops every open socket at once. All those clients then reconnect together, which looks exactly like an attack to your scale rules. Managed services exist mainly to move that problem off your replicas.

Webhooks

One system calls another when an event happens. Azure Event Grid delivers this way, with retries, dead-lettering, and support for the CloudEvents schema.

What bites: the unglamorous eighty percent, and it starts before your first event arrives.

First, you must pass a validation handshake, and the shape depends on your schema. The native Event Grid schema sends a POST carrying a SubscriptionValidationEvent. You read validationCode from the data object and echo it back as {"validationResponse": "<code>"} with a 200, within 30 seconds. As a result, an endpoint that authenticates, queues, and logs before responding can miss that window on a cold start. The CloudEvents schema works differently. It sends an HTTP OPTIONS preflight carrying a WebHook-Request-Origin header, which you echo back as WebHook-Allowed-Origin. Afterward, every real delivery carries a matching Origin header that you can check.

Second, verify signatures over the raw body, because serialization isn’t byte-stable.

Third, deduplicate. At-least-once delivery makes duplicates a certainty rather than an edge case.

Fourth, answer 400 rather than 500 when a payload will not parse. Event Grid retries 5xx responses because a server error suggests a temporary problem. A body your parser cannot read is not temporary. So an unhandled exception turns one bad payload into a retry cycle that runs until the event dead-letters, and no attempt in that cycle can ever succeed. I introduced this exact bug in the sample while writing this post, then watched a shell-quoting mistake surface it.

Ordering bites too. Event Grid delivers at least once and never guarantees sequence. A failed event enters an exponential backoff queue, while later events sail straight past it. Retries therefore scramble order actively rather than occasionally. So build consumers around sequence numbers or timestamps in the payload. Alternatively, when order is structural rather than incidental, pull from Event Hubs or a Service Bus session instead.

gRPC streaming

Service-to-service, strongly typed, over HTTP/2. ASP.NET Core supports server, client, and bidirectional streaming out of the box.

What bites: the protocol has to match on both sides of the ingress, and the two halves are configured in different places.

Start with ingress. The transport property governs the protocol between the ingress proxy and your container, not only at the edge. Container Apps needs transport: http2 for gRPC. However, the WebSocket handshake relies on the HTTP/1.1 Upgrade mechanism, which HTTP/2 replaces with stream multiplexing. So the two protocols sit badly behind one ingress.

Now the container. Kestrel needs an explicit Protocols: Http2 setting to speak cleartext HTTP/2. The Http1AndHttp2 value looks like it covers both, yet it quietly falls back to HTTP/1.1 without TLS, because negotiation depends on ALPN. Container Apps terminates TLS at the ingress, so your container never sees a handshake to negotiate over.

Kestrel says so during startup, in plain language:

HTTP/2 is not enabled for [::]:8080. The endpoint is configured to use
HTTP/1.1 and HTTP/2, but TLS is not enabled. HTTP/2 requires TLS application
protocol negotiation. Connections to this endpoint will use HTTP/1.1.

That line sits in your container logs from the first boot. Nobody reads startup logs when the application starts successfully, so it waits there until a gRPC call fails hours later.

Get one half right, and the other wrong, and the client receives this instead:

upstream connect error or disconnect/reset before headers.
retried and the latest reset reason: remote refused stream reset

That message names the upstream connection, which sends you to inspect ingress, networking, and scale rules. Meanwhile, the container is the one answering in the wrong protocol. So read the container startup logs first when a gRPC call fails at the edge. The answer usually arrives before the question.

The sample settles both halves at once. It listens on port 8080 for HTTP, SSE, and WebSockets over HTTP/1.1, and on port 8081 for gRPC over h2c. It then deploys the same image twice, once with transport: auto pointed at 8080, and once with transport: http2 pointed at 8081. One image, two ports, two ingress configurations, no compromise.

What the edge does to your stream

Here is the failure I keep seeing, and it never appears on the pattern diagram.

You build an SSE endpoint. It works perfectly on your laptop. You then publish it through API Management, and every event stops arriving. Nothing appears for two minutes, and then the whole stream lands at once, or the connection simply times out.

Your code is fine. The gateway buffers the response by default. It collects chunks from the backend, typically in 8 KB increments, and forwards them only when the buffer fills, or the stream ends. That is reasonable behavior for a normal API and fatal for a stream.

The fix is one attribute in the backend policy:

<backend>
<forward-request buffer-response="false" timeout="240" />
</backend>

Set buffer-response to false, and the gateway forwards each chunk as it arrives. Note the hyphen. An underscore looks close enough to survive a code review and fails policy validation.

Meanwhile, three more layers deserve a look before you ship:

  • Kestrel buffers writes unless you call DisableBuffering and flush after each event.
  • Front Door applies an origin response timeout. Standard and Premium default to 60 seconds, while the Classic SKU defaults to 30. You can raise it, but only to 240 seconds. That ceiling matters more than the default. No SSE stream published through Front Door survives past four minutes, so build the reconnect path and honor Last-Event-ID from the start. Otherwise, a silent disconnect looks exactly like a bug in your code.
  • App Service and Functions cap HTTP requests at 230 seconds. The Azure Functions hosting documentation states it plainly, and no setting overrides it. Microsoft’s own recommendation is the asynchronous pattern: return 202 Accepted and let the client poll. In other words, the platform pushes long streams back toward pattern number one.

Notice what those numbers have in common. App Service stops at 230 seconds. Front Door caps at 240. API Management accepts a higher timeout on forward-request, yet values above 240 are discouraged, because the network drops idle connections at that same boundary. All three inherit one limit from the load balancer underneath.

As a result, four minutes is the practical ceiling for a single connection in Azure. No pattern on the diagram changes that. Plan the reconnect instead.

In short, the pattern lives in your code, but the behavior lives in your platform configuration. Test through the full path, not against localhost.

Cost and scale change the answer

Persistent connections change how you scale. A stateless request occupies a replica for milliseconds. A WebSocket occupies one for hours.

Container Apps scales on concurrent requests, and a long-lived connection counts as one. Therefore, set that threshold low. Otherwise, your replicas fill with idle sockets long before CPU tells you anything is wrong.

Managed services price per unit and per connection, so check the current pricing pages before you compare. Still, the service fee is rarely the real cost. Backplane configuration, connection affinity, graceful drain during scale-in, and reconnect storms after each deployment cost far more engineering time than the invoice suggests.

Where this is the wrong answer

Every pattern has a place where it becomes a liability. Here are the ones I argue about most.

WebSockets for a dashboard that changes hourly. You pay for a connection per user to deliver one update, and you now maintain reconnect logic forever. Nobody has ever thanked an architect for putting a WebSocket on a management report.

SSE for mobile clients on unreliable networks. Reconnect works, yet you still need Last-Event-ID and server-side replay to avoid gaps. When the app spends half its day in the background, use push notifications instead.

Webhooks when ordering matters. Event Grid retries, so your receiver sees events twice and out of sequence. If order affects correctness, pull from a log or a session instead.

gRPC streaming to a browser. You need grpc-web plus a translating proxy. As a result, you added infrastructure to solve a problem that SSE already solved.

Long polling in new code. It exists for compatibility with systems you cannot change. For greenfield work, SSE is simply better.

A private socket server to avoid managed service pricing. The service fee is not what hurts. Sticky sessions, drain behavior, and reconnect storms are what hurt.

Real-time at all. If a user reads the number once per hour, batch it. Real-time is a requirement, not a compliment.

Try it yourself

The companion repository runs all six patterns from one image, against one price tick produced every second.

Because the data is identical everywhere, the six-pane test page shows exactly what differs. You watch the request counter climb under polling, the lag column drop under SSE, and the WebSocket pane accept a filter sent back up the same connection. You can also put API Management in front and reproduce the buffering failure in about ten minutes.

Deploy it with one script:

Bash, macOS or Linux:
./infra/deploy.sh
PowerShell on Windows:
./infra/deploy.ps1

Choose the boring one

The original advice still holds. Choose the simplest pattern that meets your latency, reliability, and scalability requirements.

I would add one line to it. Choose the simplest pattern that survives your gateway, your scale rules, and your next deployment. That constraint eliminates more options than the latency budget ever will.

Production Readiness: Closing the Execution Gap

Every post in this series so far has covered a layer: compute, messaging and orchestration, data patterns, governance and identity, observability and FinOps. This one isn’t a layer. It’s a question that cuts across all of them: when is the platform actually ready for production?

That question is harder than it looks, because “we built it” and “we can run it” are different claims. I’ve watched more than one integration platform pass every technical checkpoint and still fall short of production-ready, not because the design missed anything, but because nobody had turned that design into something enforceable, operable, and provable. So this post is about the gap between those two states, and how you close it. It’s the through-line under every layer, and it’s the thing that turns a strong architecture into a platform you can responsibly put load on.

The shift: from “is it built?” to “can we operate it?”

Here’s the single most useful reframe I know for this stage. Stop asking “is the platform technically built?” and start asking “can we operate it safely, recoverably, auditably, and predictably?”

Those are not the same question. The first is about whether the components exist and connect. The second is about whether, when something goes wrong at 2 am, someone can see what happened, understand it, recover from it, and prove afterward that they handled it correctly. A platform can pass the first test comfortably and fail the second completely. And the second test is the one that actually determines whether you should go live. So the moment you catch yourself saying “it works,” push on: does it work in a demo, or does it work under a failure you didn’t plan for?

The core move: make the implicit explicit

Most integration platforms at this stage share the same shape. The design is good, and someone has largely written it down. But a lot of what matters lives in documentation, in code that’s still evolving, or in the heads of the people who built it. That works fine while a small team builds the foundation. It stops working the moment the first real production use case lands, because implicit choices become whatever the first team happens to decide.

The fix is a production baseline: an explicit, enforceable statement of what production use demands. It names the mandatory components and patterns, pins down the environment profiles, fixes the security controls and monitoring standards, and settles the recovery agreements and release criteria. Its job is to pull those decisions out of documents and habits and into something the platform enforces, so the first production use case inherits the decisions rather than reinventing them.

Without that baseline, every early integration is free to make its own choices, and you accumulate inconsistency, technical debt, and a future re-platforming you didn’t budget for.

Design versus demonstrable: the recurring gap

The same gap shows up in every layer, and once you see it, you can’t unsee it. On security, the design names Zero Trust, least privilege, and pipeline-driven change, yet permanent broad access rights sit in the environment, quietly contradicting it. Observability design specifies OpenTelemetry and required fields, but the alert rules and dashboards never make it into the infrastructure. And the CI/CD design describes a full release chain with quality gates, while the pipelines really only cover the dev environment.

In each case the design is right and the demonstrable working is missing. That’s the pattern to hunt for when you assess readiness: not “did someone design this?” but “does the platform actually enforce this design somewhere it can prove?” Anything that lives only as intent is a gap, however good the intent.

What operational readiness actually covers

Technical function is necessary but not sufficient. Operational readiness asks a distinct set of questions, and that set decides go-live. In practice, it comes down to whether you can demonstrably answer these:

Can you see what’s happening: monitoring, chain-level tracing, message-level insight? When an incident hits, does someone triage it, and do they have the information they need to do so? For recovery, can the platform handle errors, replay from a known point, and fall back on a real Business Continuity and Disaster Recovery plan rather than an RTO written in a document? Afterward, can you prove what happened through an audit trail, an access log, and change history? On access control, do you separate permanent from elevated rights, and dev from production? For cost, can you attribute it, budget against it, and tier retention? And finally, can you hand it over — have you defined operational ownership, or does every incident route back to the people who built it?

If any of those answers is “only in the design,” please fix it before go-live, not after the first incident teaches you the hard way.

The boundary that decides scale: central versus decentral

The other thing production-readiness has to settle is a boundary, not just a checklist. Most modern integration platforms want value-stream teams to deliver independently within central guardrails. That’s the right ambition. But it only works when you draw the line between central platform ownership and team autonomy deliberately, because that line runs through every layer: API governance, messaging configuration, RBAC, pipeline use, monitoring, error handling, lifecycle, support.

Get the line wrong and you recreate the exact problem the platform set out to solve. The platform team becomes the bottleneck again — the single point through which every change, approval, and incident has to pass. So team autonomy isn’t just a tooling question. It needs explicit ownership agreements, release paths, access models, quality controls, and operational responsibilities. The tooling enables autonomy; the agreements make it safe.

Shared components are where this bites hardest. A shared API gateway, a shared message broker, a shared logging workspace these touch technology, security, governance, operations, cost, and autonomy all at once. They make the platform economical, and they carry the biggest scaling risk. For each one, decide deliberately: why does it stay shared rather than isolated, who owns it, who may change it, how do you monitor it, and how do you attribute its cost? Leave those implicit and the shared component quietly becomes everyone’s dependency and no one’s responsibility.

Readiness isn’t one moment — it’s three

The last reframe worth making: “ready” isn’t a single bar. It’s three different bars at three different moments, and conflating them is how platforms either over-build early or under-prepare for scale.

First go-live. The bar here is the minimum production baseline and demonstrable operational readiness. Not every capability has to be complete, but the ones that are preconditions for running safely in production do. This is where the baseline, the recovery plan, and the release criteria have to be real.

First team onboarding. The bar shifts to whether the federated model actually works in practice. Can one real team deliver independently within the central guardrails, without quality, security, or consistency buckling under the first real use? This is a practice test, and it’s better, even, to let some things get concrete here rather than designing them fully in the abstract.

Scaling to many teams. Now the bar is repeatability. Anything that worked at one team through direct conversation now has to become standardised, documented, and reproducible: lifecycle policy, cost allocation, onboarding, support model, versioning, exception handling. Direct alignment doesn’t scale; product-steering does.

Naming which moment a given concern belongs to is half the battle. It stops you from demanding scale-grade rigour before first go-live, and from discovering at team five that nobody built the repeatable version.

Where this thinking gets over-applied

Consistent with the series, the honesty section. “Production baseline” thinking is right, but it can tip into paralysis.

Not everything has to be complete before first go-live. The three-moments split exists precisely so you don’t. Demanding full lifecycle policy, mature FinOps, and a complete federation model before a single use case runs is how a platform never ships. Match the rigour to the moment.

A baseline that only flags is a baseline that gets ignored. The whole point of the production baseline is that the platform enforces it. A pile of documented-but-unenforced standards manufactures the appearance of readiness without the substance, which is more dangerous than an honest gap, because it invites false confidence.

You can make a decision deliberately without making it central. Drawing the central-versus-decentral line carefully doesn’t mean pulling everything central. Sometimes the deliberate call is “this is team-owned,” and recording that reasoning is the point, not the direction.

The shape of it

For an integration architect, production-readiness isn’t a technical checkpoint; it’s the shift from a platform that’s built to one you can operate responsibly. Make the implicit explicit in a production baseline. Hunt the gap between what’s designed and what’s demonstrable. Answer the operational-readiness questions before go-live, not after. Draw the central-versus-decentral line deliberately, especially for shared components. And treat “ready” as three moments, not one. Do that, and the layers from this series stop being a good architecture on paper and become a platform you can actually run.

This is the lens that ties the series together. The Azure PaaS map has the layer-by-layer foundation; this post is the question you hold every layer up against before you put production load on it.

AI Foundry Spoke Model Deployment: Why It Still Happens

A comment on the first post in this series asked why AI Foundry Spoke model deployment happens at all in the Citadel pattern, a question worth answering properly in its own post rather than buried in a reply thread.

Great article as usual! Just wondering, in the given Governance Hub & Agent Spoke architecture, what is the purpose of deploying the models both to the Spoke and the Hub? Shouldn’t they only be deployed to the Hub and provided from there?

That’s a sharp question, and it points at a real tension in the Citadel pattern that the original post didn’t call out explicitly enough. Here’s the direct answer, followed by the reasoning behind it.

The short answer

No, they shouldn’t both serve your application’s inference needs. Only the Hub deployment should. The model deployment sitting in the Spoke exists today because of how Azure AI Foundry’s Agent Service currently works, not because the architecture intends a second, governance-free inference path.

That distinction matters, so it’s worth walking through why the Spoke deployment is there at all.

Why AI Foundry Spoke Model Deployment Happens at All

In the ideal version of the Citadel pattern, every model call flows through the Hub’s APIM gateway. That’s the entire point of centralizing governance in one place. It’s what the rest of the series demonstrated: every agent call routed through apim-wpvlimv4ngkns, with token tracking, content safety, cost attribution, and the kill switch all enforced at that single choke point.

The Spoke still ends up with its own local model deployment because the AI Foundry Agent Service needs one for two reasons. This is a direct consequence of how the AI Landing Zone Bicep templates provision the Spoke, not a choice made anywhere in this series.

It powers the Agent Service’s own internal capabilities. Thread management, agent orchestration, and built-in tools like Code Interpreter or File Search (when enabled) call the model directly, through Foundry’s own runtime, rather than through any external endpoint you control. That runtime doesn’t route through APIM. It talks to whatever model deployment sits alongside it in the same project.

It satisfies the Foundry project’s provisioning requirements. A Foundry project currently expects an associated model deployment to exist as part of setting up the project, even if your actual application traffic never calls that deployment directly.

Neither of these is a governance decision. They’re artifacts of how the Agent Service is architected right now.

Two paths, not one

The practical result is that a Citadel deployment ends up with two separate paths to a model, and they serve different purposes.

The Spoke’s local deployment exists for Foundry’s own internal agent runtime. It’s not meant to see your production traffic, and if it does, none of the governance you built in the Hub applies to those calls.

The Hub’s deployment, reached through APIM, is what your application code should use. That’s what we wired up explicitly in Part 2 of this series, with the standard OpenAI SDK pointed at the gateway rather than directly at the Foundry endpoint. The Hub itself is built on the AI Hub Gateway Solution Accelerator, and its AI gateway capabilities are exactly what give APIM the token metering, content safety, and audit trail features this series has leaned on throughout.

The second path exists precisely because of a limitation the series already documented. The Agent Service SDK, in its current preview state, doesn’t route its own LLM calls through APIM. It bypasses the gateway entirely, which means using it directly would mean giving up token metering, policy enforcement, and audit trails on every call the agent makes. That’s why Part 2 used the standard OpenAI SDK pointed at APIM instead of the native Agent Service SDK, and it’s the same underlying issue this reader’s question is really about.

What this means in practice

If you’re building on this pattern today, treat the Spoke’s model deployment as infrastructure the platform needs to exist, not as a second inference endpoint your application is allowed to call. Point your application code at the Hub, through APIM, every time. Leave the Spoke deployment alone to do the job Foundry needs it for internally, and don’t build anything that calls it directly for your own traffic.

If you’re reviewing someone else’s Citadel-pattern deployment, this is worth checking explicitly. A model deployment sitting in a Spoke isn’t wrong by itself, but it’s worth confirming nothing in the application is quietly calling it and skipping the gateway.

Where this is heading

I’d expect this to tighten up as the Agent Service SDK matures out of preview and gains native APIM routing support. When that happens, the two-path situation described here becomes a one-path situation, and the Spoke deployment stops being something you need to actively route around.

Until then, the answer to the original question stands: deploy to the Hub, govern everything through APIM, and treat the Spoke’s model deployment as plumbing the platform needs rather than a second front door.

Thanks to the reader who asked the original question. It’s exactly the kind of detail that’s easy to leave implicit in an architecture diagram and much more useful said out loud.

AI Gateway Commercial vs Open Source: How to Choose the Right Control Plane

The AI gateway commercial vs. open-source decision is one that most organizations reach not by planning but by accident. One team has already integrated directly with Azure OpenAI. Another is using LiteLLM to wrap a few models. A third wants to use the enterprise API management platform you already have. Suddenly, you need to make a choice, and the conversation gets complicated fast.

This companion post to my APIM for AI Workloads series takes a step back from the Azure API Management specifics and addresses the question that comes before all of it: which gateway should you be using in the first place? The series covers APIM in depth because it’s the right answer for the Microsoft ecosystem. But it’s not the only answer, and for some organizations it’s not the right one.

Here is how to think through the decision properly.

Why the AI Gateway Commercial vs Open Source Choice Matters More Than You Think

Most API gateway decisions are relatively low-stakes. If you pick the wrong one, you migrate. But the AI gateway decision carries more weight for two reasons.

First, the gateway sits in the critical path of every AI interaction in your organization. Its policy language, authentication model, and observability hooks become embedded in the way your teams build AI-powered applications. Switching later is not impossible, but it is disruptive.

Second, the governance patterns you establish now, how you handle token limits, cross-charging, PII, and compliance logging, are much harder to retrofit than to design in from the start. The Team Rockstars IT AI Gateway whitepaper, published this month, makes this point well: organizations that set up audit logging via an AI gateway from day one build a direct compliance advantage under the EU AI Act. Those who add it later risk complex and costly rework.

So the choice deserves deliberate thought, not a default.

The Commercial Options for AI Gateway

Commercial AI gateways offer a faster path to production and offload operational complexity to the vendor. The main options in the market today are:

Azure API Management is the right choice if you are already in the Microsoft ecosystem. Its AI-specific policy extensions for token limits, token metrics, semantic caching, and load balancing across PTU and PAYG backends are mature and tightly integrated with Azure Monitor and Application Insights. The series covers this in depth from Part 1 onwards.

Kong Konnect is a strong option for organizations that already use Kong for API management and want to extend it into AI. Its plugin ecosystem covers rate limiting, authentication, and observability, with AI-specific plugins growing quickly.

Portkey is purpose-built as an AI gateway with a lightweight footprint and fast time-to-value. It supports a broad range of model providers, has built-in semantic caching and observability, and is a practical option for teams that want AI governance without the overhead of a full enterprise API management platform.

Apigee (Google Cloud) is the natural choice for GCP-centric organizations. Like APIM in the Microsoft world, its AI gateway capabilities are deepening with each release as Google embeds Gemini and Vertex AI integrations.

The common advantages across all commercial options are faster deployment, built-in compliance features, vendor support contracts, and operational burden offloaded to the vendor. The common risks are licensing costs, proprietary policy languages that create switching friction, and dependency on the vendor’s roadmap.

The Open Source Options for AI Gateway

Open-source gateways offer maximum control and no licensing costs, but they require your organization to own what the vendor would otherwise handle.

LiteLLM is the most widely adopted open source AI gateway today. It provides a unified API across more than 100 model providers, with built-in rate limiting, spend tracking, and a proxy server that is straightforward to self-host. The community is active, and the feature velocity is high. The supply chain risk is real, though: a 2025 attack targeting LiteLLM and Trivy demonstrated that even widely used security-adjacent tools can become attack vectors. If you run LiteLLM in production, you own the patching cadence.

Agent Gateway from Anthropic is purpose-built for MCP and agentic traffic. If your primary use case is governing tool calls from AI agents rather than managing completion API traffic, it is worth evaluating alongside the broader options.

One API provides a unified, OpenAI-compatible interface across multiple providers and is widely used by organizations seeking provider-agnostic routing without vendor lock-in.

HelixML focuses on self-hosted deployments with strong data-sovereignty properties, making it relevant for organizations where data-residency requirements rule out SaaS-based gateway options.

AI Gateway Commercial vs Open Source: Five Decision Factors

Five factors consistently determine which direction is right for a given organization:

Time to value. In my experience, commercial gateways can be production-ready in days to weeks. Open source deployments typically take weeks to months to reach production quality, depending on how much custom policy logic you need to build. If you have an urgent compliance or cost control problem to solve, commercial is the pragmatic choice.

Compliance and data residency. For Dutch and European organizations operating under AVG, NIS2, and the EU AI Act, commercial gateways offer contractual guarantees: data processing agreements, certified regions, and SLAs with defined incident response times. Open source can meet the same requirements, but you are responsible for demonstrating compliance yourself rather than relying on a vendor certification.

Internal platform capability. Open source is not free. The licensing cost is zero, but according to the CNCF’s platform engineering maturity model. Organizations without a dedicated platform engineering team that can credibly own the gateway long-term should not choose open source. The operational gap will become visible at the worst possible moment.

Flexibility and lock-in risk. Open source wins on long-term flexibility. Proprietary policy languages in commercial gateways create switching friction that grows over time as you invest in custom policies. If multi-cloud strategy and provider-agnosticism are strategic priorities, design your gateway layer with that in mind from the start, even if you begin with a commercial option, applying the strangler fig pattern to abstract away proprietary dependencies over time.

Supply chain risk. This factor is underweighted in most evaluations. The 2025 supply chain attack targeting LiteLLM and Trivy demonstrated that open source security tooling itself can become an attack vector. Commercial vendors have contractual obligations around vulnerability disclosure and patching. With open source, that obligation falls to your team.

A Decision Framework for AI Gateway Commercial vs Open Source

The flowchart above works through the most decisive questions in order. A few practical observations from applying it:

Regulated industries almost always land in commercial. Healthcare, financial services, and insurance organizations operating under Dutch or European regulation have compliance requirements that are significantly easier to satisfy with contractual vendor guarantees than with self-operated open source tooling. At my company, the AVG and healthcare-specific data processing requirements made APIM the clear choice.

The hybrid pattern is underused. Many organizations run a commercial gateway in production for governed workloads, while developer teams use LiteLLM or a lightweight open source option in lower environments for experimentation. This gives you the compliance and operational properties you need in production while keeping the innovation surface open. It is more work to maintain two gateway patterns, but the tradeoff is often worth it.

Design for replaceability regardless of what you choose. The Team Rockstars whitepaper frames this well: choose your first gateway deliberately, but design for replacement. Use open standards, abstract your policy logic where possible, and avoid deep coupling to proprietary features without open-source equivalents. The gateway landscape is evolving fast enough that what is the right choice today may not be in two years.

Where This Fits in the APIM for AI Workloads Series

The rest of the series goes deep on Azure API Management specifically: the token metric policy, load balancing and circuit breaking, semantic caching, and MCP gateway for agentic workloads. If you have landed on APIM as your gateway of choice or if you are in a Microsoft-centric organization where it is the natural fit, the series covers the production patterns you need.

  • Part 1: Why your AI APIs need a gateway.
  • Part 2: Authentication and authorization.
  • Part 3: Token limit policy.
  • Part 4: Token metric policy and cross-charging.
  • Part 5: Load balancing and circuit breaking.
  • Part 6: Semantic caching.
  • Part 7: APIM as MCP gateway for agentic AI workloads.

Azure API Management for AI: Securing Your AI APIs with Authentication and Authorization

Part 2 of 7 in the “APIM for AI Workloads” series

In Part 1 of this series, I made the case for why Azure API Management for AI workloads is the right control plane for governing AI traffic across an organization. This post gets practical: how do you actually secure access to your AI backends with APIM without creating a credential-management nightmare?

Security is where many AI projects cut corners, and understandably so. When you’re moving fast to prove value with a new model, authentication feels like overhead. But AI endpoints are expensive, and an unsecured Azure OpenAI endpoint is a real risk: anyone with the URL and key can start consuming tokens at your cost. At scale, that’s a significant financial and compliance exposure.

APIM addresses this with a three-layer security model. Let’s walk through each layer.

Azure API Management for AI Security: A Three-Layer Model

The authentication and authorization pattern in APIM is deliberately layered. Each layer answers a different question and operates independently, so a failure at any layer stops the request before it reaches the AI backend.

The three layers are:

  • Subscription keys to identify and track API consumers.
  • JWT validation to enforce fine-grained access control based on claims.
  • Managed Identity to authenticate APIM to Azure OpenAI without storing credentials.

Each layer has a distinct role. Confusing them is a common mistake, so it’s worth being explicit about what each one does and does not do.

Layer 1: Subscription Keys

Subscription keys are APIM’s mechanism for identifying API consumers. When you create an API product in APIM and require a subscription, callers must include their key in the Ocp-Apim-Subscription-Key header. APIM validates the key, maps it to a subscriber, and lets the request proceed.

This is important for AI workloads specifically because subscription keys enable per-consumer token tracking. When you combine subscription key validation with the Token Metric policy we’ll cover in Part 4, you get usage data broken down by subscriber, which is the foundation of any internal cross-charging model.

Subscription keys answer the question: Who is calling? They don’t answer what the caller is allowed to do. For that, you need JWT validation.

Layer 2: JWT Validation and Claims-Based Authorization

The validate-jwt policy is where you enforce what a caller is permitted to do. It validates the JWT token in the Authorization header against your identity provider, and can inspect any claim in the token to make authorization decisions.

For Azure OpenAI specifically, this is where you control which teams or applications can access which model deployments. A team working on an internal chatbot should not be able to call a GPT-4o deployment reserved for a production workload. JWT claims let you enforce that boundary at the gateway layer, with no changes required in the calling application.

A typical policy checks the token signature against your Azure AD tenant’s OpenID Connect configuration, then validates that a required scope or role claim is present:

The failed-validation-httpcode=”401″ attribute ensures unauthenticated callers get a clean rejection before they ever reach the backend. You can also use failed-validation-error-message to return a specific error message, which helps consumers debug auth failures without exposing internal details.

For multi-provider setups where you’re routing to non-Azure backends like Mistral or Cohere, the same JWT policy applies. The claims model is provider-agnostic, which is one of the advantages of centralizing auth in APIM rather than handling it per-backend.

Layer 3: Managed Identity for Backend Authentication

Managed Identity is the most important security improvement you can make when setting up Azure API Management for AI. It replaces the pattern of storing an Azure OpenAI API key in APIM’s named values with a system-assigned or user-assigned Managed Identity that APIM uses to authenticate directly to Azure OpenAI via Azure AD.

The practical difference is significant. With API key authentication, you have a long-lived secret that needs to be stored, rotated, and kept out of source control. With Managed Identity, there is no secret. APIM requests a short-lived token from Azure AD at runtime, and Azure AD issues it based on the APIM instance’s identity. Nothing is stored. Nothing can leak.

The configuration is a single policy element in the inbound section: <authentication-managed-identity resource=”https://cognitiveservices.azure.com”/&gt;. APIM handles the rest, automatically fetching and refreshing the token.

On the Azure OpenAI side, you grant the APIM instance’s Managed Identity the Cognitive Services User role on the Azure OpenAI resource. That’s the minimum required permission. You can scope it further to specific deployments if needed.

For organizations in regulated industries, such as healthcare, financial services, and government, Managed Identity is not optional. It satisfies Zero Trust authentication requirements and produces a full audit trail in Azure Monitor, tied to the APIM instance identity rather than a shared key.

Azure API Management for AI: Putting the Three Layers Together

In a production setup, all three layers run sequentially within the inbound policy pipeline. A request arrives with a subscription key and a JWT. APIM validates the key first (fast, no external call), then validates the JWT against Azure AD, then forwards the request to Azure OpenAI using its Managed Identity token. The AI backend never sees the caller’s JWT, and APIM never stores an API key.

The result is a clean separation of concerns:

  • The calling application manages its own JWT (issued by Azure AD based on its own identity or the user’s identity).
  • APIM enforces the authorization policy without the backend needing to know anything about it.
  • The AI backend trusts only APIM’s Managed Identity, not arbitrary callers.

This is the architecture you want before you go to production with any AI workload that touches sensitive data or incurs meaningful cost.

What’s Next in This Series

Part 3 covers the Token Limit policy: how to enforce tokens-per-minute limits per consumer, configure throttling behavior, and handle the differences between the azure-openai-token-limit and llm-token-limit policy variants.