I see a diagram of real-time communication patterns pop up in my feed every few weeks. Polling, long polling, Server-Sent Events, WebSockets, webhooks, gRPC streaming. Six neat boxes, and one piece of sensible advice underneath: choose the simplest model that fits the use case.
The diagram is correct. However, it is not the part that costs you a sprint.
Picking a pattern takes about five minutes. Making that pattern survive API Management, Azure Front Door, a scale-out event, and a rolling deployment takes considerably longer. This post covers that second part. It also comes with a working sample that runs all six patterns from one image, against one shared event source. Azure needs two apps to host it, and that’s one of the more interesting findings below.
Four questions, not six boxes
A grid of six options invites you to shop. Instead, answer four questions in order. The pattern then picks itself.

Which direction does data flow? The client pulls, the server pushes, or both sides talk at once. This single question removes at least three options.
What is the latency budget? Seconds are cheap. Milliseconds are not. Most business dashboards tolerate a delay that their designers never measured.
Who owns the connection when it drops? Someone must handle reconnect, resume, and replay. If your answer is “the browser does that automatically”, read the reconnect semantics again first.
Who pays per open socket? Idle connections still consume replicas, units, and money.
When two patterns both fit, pick the one that keeps state out of the connection. A dropped request costs a retry. A dropped session costs a reconnect, a replay, and a support ticket.
The six patterns, and what bites in Azure
Polling
The client asks again on its own schedule. Azure itself uses this pattern constantly. The asynchronous request-reply pattern returns 202 Accepted with a Location header, and Durable Functions exposes a status query endpoint that works exactly this way.
What bites: the missing Retry-After header. Without it, every client invents its own interval. Ten thousand clients then converge on the same second after a deployment, and your scale rule reacts to a spike you created yourself.
Long polling
The server holds the request open until data arrives or the clock runs out. Azure Service Bus applies the same idea inside its SDK, where a receive call waits for a configurable maximum wait time instead of returning empty.
What bites: timeouts you do not control. App Service and Azure Functions enforce a fixed 230-second ceiling on HTTP requests. That number comes from the Azure Load Balancer underneath, which idles connections out at 240 seconds by default. You cannot raise it because it is a platform constraint, not an application setting. Front Door then applies its own origin response timeout on top. Therefore, derive your hold time from the shortest timeout in the path, not from the longest. The sample holds for 25 seconds, which clears every layer.
Server-Sent Events
The server pushes over one long-lived HTTP response. SSE is unfashionable and quietly excellent. It runs over plain HTTP, it survives proxies that understand chunked responses, and browsers reconnect on their own.
You already depend on it. Token streaming from Azure OpenAI and Microsoft Foundry arrives as Server-Sent Events. Every chat interface you have built this year uses this pattern, whether or not the architecture diagram says so.
What bites: buffering, which gets its own section below. Also reconnect gaps. The browser resends the Last-Event-ID header automatically, but the server must honor it. Otherwise, every reconnect silently drops the events that arrived while the socket was down.
WebSockets
Both sides talk over one persistent connection. This is the right answer for chat, collaborative editing, and live trading. It is the wrong answer for a dashboard that changes twice an hour.
Azure gives you two managed options. That is Azure Web PubSub, which handles raw WebSocket clients and works well outside .NET. And Azure SignalR Service fits when you already use hubs and want fallbacks.
What bites: the deployment. A rolling revision in Azure Container Apps drops every open socket at once. All those clients then reconnect together, which looks exactly like an attack to your scale rules. Managed services exist mainly to move that problem off your replicas.
Webhooks
One system calls another when an event happens. Azure Event Grid delivers this way, with retries, dead-lettering, and support for the CloudEvents schema.
What bites: the unglamorous eighty percent, and it starts before your first event arrives.
First, you must pass a validation handshake, and the shape depends on your schema. The native Event Grid schema sends a POST carrying a SubscriptionValidationEvent. You read validationCode from the data object and echo it back as {"validationResponse": "<code>"} with a 200, within 30 seconds. As a result, an endpoint that authenticates, queues, and logs before responding can miss that window on a cold start. The CloudEvents schema works differently. It sends an HTTP OPTIONS preflight carrying a WebHook-Request-Origin header, which you echo back as WebHook-Allowed-Origin. Afterward, every real delivery carries a matching Origin header that you can check.
Second, verify signatures over the raw body, because serialization isn’t byte-stable.
Third, deduplicate. At-least-once delivery makes duplicates a certainty rather than an edge case.
Fourth, answer 400 rather than 500 when a payload will not parse. Event Grid retries 5xx responses because a server error suggests a temporary problem. A body your parser cannot read is not temporary. So an unhandled exception turns one bad payload into a retry cycle that runs until the event dead-letters, and no attempt in that cycle can ever succeed. I introduced this exact bug in the sample while writing this post, then watched a shell-quoting mistake surface it.
Ordering bites too. Event Grid delivers at least once and never guarantees sequence. A failed event enters an exponential backoff queue, while later events sail straight past it. Retries therefore scramble order actively rather than occasionally. So build consumers around sequence numbers or timestamps in the payload. Alternatively, when order is structural rather than incidental, pull from Event Hubs or a Service Bus session instead.
gRPC streaming
Service-to-service, strongly typed, over HTTP/2. ASP.NET Core supports server, client, and bidirectional streaming out of the box.
What bites: the protocol has to match on both sides of the ingress, and the two halves are configured in different places.
Start with ingress. The transport property governs the protocol between the ingress proxy and your container, not only at the edge. Container Apps needs transport: http2 for gRPC. However, the WebSocket handshake relies on the HTTP/1.1 Upgrade mechanism, which HTTP/2 replaces with stream multiplexing. So the two protocols sit badly behind one ingress.
Now the container. Kestrel needs an explicit Protocols: Http2 setting to speak cleartext HTTP/2. The Http1AndHttp2 value looks like it covers both, yet it quietly falls back to HTTP/1.1 without TLS, because negotiation depends on ALPN. Container Apps terminates TLS at the ingress, so your container never sees a handshake to negotiate over.
Kestrel says so during startup, in plain language:
HTTP/2 is not enabled for [::]:8080. The endpoint is configured to useHTTP/1.1 and HTTP/2, but TLS is not enabled. HTTP/2 requires TLS applicationprotocol negotiation. Connections to this endpoint will use HTTP/1.1.
That line sits in your container logs from the first boot. Nobody reads startup logs when the application starts successfully, so it waits there until a gRPC call fails hours later.
Get one half right, and the other wrong, and the client receives this instead:
upstream connect error or disconnect/reset before headers.retried and the latest reset reason: remote refused stream reset
That message names the upstream connection, which sends you to inspect ingress, networking, and scale rules. Meanwhile, the container is the one answering in the wrong protocol. So read the container startup logs first when a gRPC call fails at the edge. The answer usually arrives before the question.
The sample settles both halves at once. It listens on port 8080 for HTTP, SSE, and WebSockets over HTTP/1.1, and on port 8081 for gRPC over h2c. It then deploys the same image twice, once with transport: auto pointed at 8080, and once with transport: http2 pointed at 8081. One image, two ports, two ingress configurations, no compromise.
What the edge does to your stream
Here is the failure I keep seeing, and it never appears on the pattern diagram.
You build an SSE endpoint. It works perfectly on your laptop. You then publish it through API Management, and every event stops arriving. Nothing appears for two minutes, and then the whole stream lands at once, or the connection simply times out.

Your code is fine. The gateway buffers the response by default. It collects chunks from the backend, typically in 8 KB increments, and forwards them only when the buffer fills, or the stream ends. That is reasonable behavior for a normal API and fatal for a stream.
The fix is one attribute in the backend policy:
<backend> <forward-request buffer-response="false" timeout="240" /></backend>
Set buffer-response to false, and the gateway forwards each chunk as it arrives. Note the hyphen. An underscore looks close enough to survive a code review and fails policy validation.
Meanwhile, three more layers deserve a look before you ship:
- Kestrel buffers writes unless you call
DisableBufferingand flush after each event. - Front Door applies an origin response timeout. Standard and Premium default to 60 seconds, while the Classic SKU defaults to 30. You can raise it, but only to 240 seconds. That ceiling matters more than the default. No SSE stream published through Front Door survives past four minutes, so build the reconnect path and honor
Last-Event-IDfrom the start. Otherwise, a silent disconnect looks exactly like a bug in your code. - App Service and Functions cap HTTP requests at 230 seconds. The Azure Functions hosting documentation states it plainly, and no setting overrides it. Microsoft’s own recommendation is the asynchronous pattern: return
202 Acceptedand let the client poll. In other words, the platform pushes long streams back toward pattern number one.
Notice what those numbers have in common. App Service stops at 230 seconds. Front Door caps at 240. API Management accepts a higher timeout on forward-request, yet values above 240 are discouraged, because the network drops idle connections at that same boundary. All three inherit one limit from the load balancer underneath.
As a result, four minutes is the practical ceiling for a single connection in Azure. No pattern on the diagram changes that. Plan the reconnect instead.
In short, the pattern lives in your code, but the behavior lives in your platform configuration. Test through the full path, not against localhost.
Cost and scale change the answer
Persistent connections change how you scale. A stateless request occupies a replica for milliseconds. A WebSocket occupies one for hours.
Container Apps scales on concurrent requests, and a long-lived connection counts as one. Therefore, set that threshold low. Otherwise, your replicas fill with idle sockets long before CPU tells you anything is wrong.
Managed services price per unit and per connection, so check the current pricing pages before you compare. Still, the service fee is rarely the real cost. Backplane configuration, connection affinity, graceful drain during scale-in, and reconnect storms after each deployment cost far more engineering time than the invoice suggests.
Where this is the wrong answer
Every pattern has a place where it becomes a liability. Here are the ones I argue about most.
WebSockets for a dashboard that changes hourly. You pay for a connection per user to deliver one update, and you now maintain reconnect logic forever. Nobody has ever thanked an architect for putting a WebSocket on a management report.
SSE for mobile clients on unreliable networks. Reconnect works, yet you still need Last-Event-ID and server-side replay to avoid gaps. When the app spends half its day in the background, use push notifications instead.
Webhooks when ordering matters. Event Grid retries, so your receiver sees events twice and out of sequence. If order affects correctness, pull from a log or a session instead.
gRPC streaming to a browser. You need grpc-web plus a translating proxy. As a result, you added infrastructure to solve a problem that SSE already solved.
Long polling in new code. It exists for compatibility with systems you cannot change. For greenfield work, SSE is simply better.
A private socket server to avoid managed service pricing. The service fee is not what hurts. Sticky sessions, drain behavior, and reconnect storms are what hurt.
Real-time at all. If a user reads the number once per hour, batch it. Real-time is a requirement, not a compliment.
Try it yourself
The companion repository runs all six patterns from one image, against one price tick produced every second.

Because the data is identical everywhere, the six-pane test page shows exactly what differs. You watch the request counter climb under polling, the lag column drop under SSE, and the WebSocket pane accept a filter sent back up the same connection. You can also put API Management in front and reproduce the buffering failure in about ten minutes.
Deploy it with one script:
Bash, macOS or Linux:./infra/deploy.shPowerShell on Windows:./infra/deploy.ps1
Choose the boring one
The original advice still holds. Choose the simplest pattern that meets your latency, reliability, and scalability requirements.
I would add one line to it. Choose the simplest pattern that survives your gateway, your scale rules, and your next deployment. That constraint eliminates more options than the latency budget ever will.