Six Real-Time Patterns in Azure: What the Diagram Leaves Out

I see a diagram of real-time communication patterns pop up in my feed every few weeks. Polling, long polling, Server-Sent Events, WebSockets, webhooks, gRPC streaming. Six neat boxes, and one piece of sensible advice underneath: choose the simplest model that fits the use case.

The diagram is correct. However, it is not the part that costs you a sprint.

Picking a pattern takes about five minutes. Making that pattern survive API Management, Azure Front Door, a scale-out event, and a rolling deployment takes considerably longer. This post covers that second part. It also comes with a working sample that runs all six patterns from one image, against one shared event source. Azure needs two apps to host it, and that’s one of the more interesting findings below.

Four questions, not six boxes

A grid of six options invites you to shop. Instead, answer four questions in order. The pattern then picks itself.

Which direction does data flow? The client pulls, the server pushes, or both sides talk at once. This single question removes at least three options.

What is the latency budget? Seconds are cheap. Milliseconds are not. Most business dashboards tolerate a delay that their designers never measured.

Who owns the connection when it drops? Someone must handle reconnect, resume, and replay. If your answer is “the browser does that automatically”, read the reconnect semantics again first.

Who pays per open socket? Idle connections still consume replicas, units, and money.

When two patterns both fit, pick the one that keeps state out of the connection. A dropped request costs a retry. A dropped session costs a reconnect, a replay, and a support ticket.

The six patterns, and what bites in Azure

Polling

The client asks again on its own schedule. Azure itself uses this pattern constantly. The asynchronous request-reply pattern returns 202 Accepted with a Location header, and Durable Functions exposes a status query endpoint that works exactly this way.

What bites: the missing Retry-After header. Without it, every client invents its own interval. Ten thousand clients then converge on the same second after a deployment, and your scale rule reacts to a spike you created yourself.

Long polling

The server holds the request open until data arrives or the clock runs out. Azure Service Bus applies the same idea inside its SDK, where a receive call waits for a configurable maximum wait time instead of returning empty.

What bites: timeouts you do not control. App Service and Azure Functions enforce a fixed 230-second ceiling on HTTP requests. That number comes from the Azure Load Balancer underneath, which idles connections out at 240 seconds by default. You cannot raise it because it is a platform constraint, not an application setting. Front Door then applies its own origin response timeout on top. Therefore, derive your hold time from the shortest timeout in the path, not from the longest. The sample holds for 25 seconds, which clears every layer.

Server-Sent Events

The server pushes over one long-lived HTTP response. SSE is unfashionable and quietly excellent. It runs over plain HTTP, it survives proxies that understand chunked responses, and browsers reconnect on their own.

You already depend on it. Token streaming from Azure OpenAI and Microsoft Foundry arrives as Server-Sent Events. Every chat interface you have built this year uses this pattern, whether or not the architecture diagram says so.

What bites: buffering, which gets its own section below. Also reconnect gaps. The browser resends the Last-Event-ID header automatically, but the server must honor it. Otherwise, every reconnect silently drops the events that arrived while the socket was down.

WebSockets

Both sides talk over one persistent connection. This is the right answer for chat, collaborative editing, and live trading. It is the wrong answer for a dashboard that changes twice an hour.

Azure gives you two managed options. That is Azure Web PubSub, which handles raw WebSocket clients and works well outside .NET. And Azure SignalR Service fits when you already use hubs and want fallbacks.

What bites: the deployment. A rolling revision in Azure Container Apps drops every open socket at once. All those clients then reconnect together, which looks exactly like an attack to your scale rules. Managed services exist mainly to move that problem off your replicas.

Webhooks

One system calls another when an event happens. Azure Event Grid delivers this way, with retries, dead-lettering, and support for the CloudEvents schema.

What bites: the unglamorous eighty percent, and it starts before your first event arrives.

First, you must pass a validation handshake, and the shape depends on your schema. The native Event Grid schema sends a POST carrying a SubscriptionValidationEvent. You read validationCode from the data object and echo it back as {"validationResponse": "<code>"} with a 200, within 30 seconds. As a result, an endpoint that authenticates, queues, and logs before responding can miss that window on a cold start. The CloudEvents schema works differently. It sends an HTTP OPTIONS preflight carrying a WebHook-Request-Origin header, which you echo back as WebHook-Allowed-Origin. Afterward, every real delivery carries a matching Origin header that you can check.

Second, verify signatures over the raw body, because serialization isn’t byte-stable.

Third, deduplicate. At-least-once delivery makes duplicates a certainty rather than an edge case.

Fourth, answer 400 rather than 500 when a payload will not parse. Event Grid retries 5xx responses because a server error suggests a temporary problem. A body your parser cannot read is not temporary. So an unhandled exception turns one bad payload into a retry cycle that runs until the event dead-letters, and no attempt in that cycle can ever succeed. I introduced this exact bug in the sample while writing this post, then watched a shell-quoting mistake surface it.

Ordering bites too. Event Grid delivers at least once and never guarantees sequence. A failed event enters an exponential backoff queue, while later events sail straight past it. Retries therefore scramble order actively rather than occasionally. So build consumers around sequence numbers or timestamps in the payload. Alternatively, when order is structural rather than incidental, pull from Event Hubs or a Service Bus session instead.

gRPC streaming

Service-to-service, strongly typed, over HTTP/2. ASP.NET Core supports server, client, and bidirectional streaming out of the box.

What bites: the protocol has to match on both sides of the ingress, and the two halves are configured in different places.

Start with ingress. The transport property governs the protocol between the ingress proxy and your container, not only at the edge. Container Apps needs transport: http2 for gRPC. However, the WebSocket handshake relies on the HTTP/1.1 Upgrade mechanism, which HTTP/2 replaces with stream multiplexing. So the two protocols sit badly behind one ingress.

Now the container. Kestrel needs an explicit Protocols: Http2 setting to speak cleartext HTTP/2. The Http1AndHttp2 value looks like it covers both, yet it quietly falls back to HTTP/1.1 without TLS, because negotiation depends on ALPN. Container Apps terminates TLS at the ingress, so your container never sees a handshake to negotiate over.

Kestrel says so during startup, in plain language:

HTTP/2 is not enabled for [::]:8080. The endpoint is configured to use
HTTP/1.1 and HTTP/2, but TLS is not enabled. HTTP/2 requires TLS application
protocol negotiation. Connections to this endpoint will use HTTP/1.1.

That line sits in your container logs from the first boot. Nobody reads startup logs when the application starts successfully, so it waits there until a gRPC call fails hours later.

Get one half right, and the other wrong, and the client receives this instead:

upstream connect error or disconnect/reset before headers.
retried and the latest reset reason: remote refused stream reset

That message names the upstream connection, which sends you to inspect ingress, networking, and scale rules. Meanwhile, the container is the one answering in the wrong protocol. So read the container startup logs first when a gRPC call fails at the edge. The answer usually arrives before the question.

The sample settles both halves at once. It listens on port 8080 for HTTP, SSE, and WebSockets over HTTP/1.1, and on port 8081 for gRPC over h2c. It then deploys the same image twice, once with transport: auto pointed at 8080, and once with transport: http2 pointed at 8081. One image, two ports, two ingress configurations, no compromise.

What the edge does to your stream

Here is the failure I keep seeing, and it never appears on the pattern diagram.

You build an SSE endpoint. It works perfectly on your laptop. You then publish it through API Management, and every event stops arriving. Nothing appears for two minutes, and then the whole stream lands at once, or the connection simply times out.

Your code is fine. The gateway buffers the response by default. It collects chunks from the backend, typically in 8 KB increments, and forwards them only when the buffer fills, or the stream ends. That is reasonable behavior for a normal API and fatal for a stream.

The fix is one attribute in the backend policy:

<backend>
<forward-request buffer-response="false" timeout="240" />
</backend>

Set buffer-response to false, and the gateway forwards each chunk as it arrives. Note the hyphen. An underscore looks close enough to survive a code review and fails policy validation.

Meanwhile, three more layers deserve a look before you ship:

  • Kestrel buffers writes unless you call DisableBuffering and flush after each event.
  • Front Door applies an origin response timeout. Standard and Premium default to 60 seconds, while the Classic SKU defaults to 30. You can raise it, but only to 240 seconds. That ceiling matters more than the default. No SSE stream published through Front Door survives past four minutes, so build the reconnect path and honor Last-Event-ID from the start. Otherwise, a silent disconnect looks exactly like a bug in your code.
  • App Service and Functions cap HTTP requests at 230 seconds. The Azure Functions hosting documentation states it plainly, and no setting overrides it. Microsoft’s own recommendation is the asynchronous pattern: return 202 Accepted and let the client poll. In other words, the platform pushes long streams back toward pattern number one.

Notice what those numbers have in common. App Service stops at 230 seconds. Front Door caps at 240. API Management accepts a higher timeout on forward-request, yet values above 240 are discouraged, because the network drops idle connections at that same boundary. All three inherit one limit from the load balancer underneath.

As a result, four minutes is the practical ceiling for a single connection in Azure. No pattern on the diagram changes that. Plan the reconnect instead.

In short, the pattern lives in your code, but the behavior lives in your platform configuration. Test through the full path, not against localhost.

Cost and scale change the answer

Persistent connections change how you scale. A stateless request occupies a replica for milliseconds. A WebSocket occupies one for hours.

Container Apps scales on concurrent requests, and a long-lived connection counts as one. Therefore, set that threshold low. Otherwise, your replicas fill with idle sockets long before CPU tells you anything is wrong.

Managed services price per unit and per connection, so check the current pricing pages before you compare. Still, the service fee is rarely the real cost. Backplane configuration, connection affinity, graceful drain during scale-in, and reconnect storms after each deployment cost far more engineering time than the invoice suggests.

Where this is the wrong answer

Every pattern has a place where it becomes a liability. Here are the ones I argue about most.

WebSockets for a dashboard that changes hourly. You pay for a connection per user to deliver one update, and you now maintain reconnect logic forever. Nobody has ever thanked an architect for putting a WebSocket on a management report.

SSE for mobile clients on unreliable networks. Reconnect works, yet you still need Last-Event-ID and server-side replay to avoid gaps. When the app spends half its day in the background, use push notifications instead.

Webhooks when ordering matters. Event Grid retries, so your receiver sees events twice and out of sequence. If order affects correctness, pull from a log or a session instead.

gRPC streaming to a browser. You need grpc-web plus a translating proxy. As a result, you added infrastructure to solve a problem that SSE already solved.

Long polling in new code. It exists for compatibility with systems you cannot change. For greenfield work, SSE is simply better.

A private socket server to avoid managed service pricing. The service fee is not what hurts. Sticky sessions, drain behavior, and reconnect storms are what hurt.

Real-time at all. If a user reads the number once per hour, batch it. Real-time is a requirement, not a compliment.

Try it yourself

The companion repository runs all six patterns from one image, against one price tick produced every second.

Because the data is identical everywhere, the six-pane test page shows exactly what differs. You watch the request counter climb under polling, the lag column drop under SSE, and the WebSocket pane accept a filter sent back up the same connection. You can also put API Management in front and reproduce the buffering failure in about ten minutes.

Deploy it with one script:

Bash, macOS or Linux:
./infra/deploy.sh
PowerShell on Windows:
./infra/deploy.ps1

Choose the boring one

The original advice still holds. Choose the simplest pattern that meets your latency, reliability, and scalability requirements.

I would add one line to it. Choose the simplest pattern that survives your gateway, your scale rules, and your next deployment. That constraint eliminates more options than the latency budget ever will.

Cosmos DB Agent Kit: Auditing My Own Sample Repo

A bonus post, series-adjacent, running the Cosmos DB Agent Kit against cosmos-agent-memory-lab to see what six posts of hand-checked rules missed.

Every rule in this series hierarchical partition keys, TTL modes, vector index types, the composite ID trick came from reading Microsoft Learn docs and debugging real errors one at a time. The Cosmos DB Agent Kit packages that same category of knowledge differently: 100+ best-practice rules across 12 categories that AI coding agents apply while writing or reviewing Cosmos DB code. This post checks the rules against the actual sample repo behind posts 2 through 4, rule by rule, instead of taking the kit’s word for it.

What the Cosmos DB Agent Kit Actually Ships

It’s not a linter or a static-analysis tool. It’s a set of Markdown rule files, one per practice, each with an incorrect example, a corrected example, and a short explanation of why, following the Agent Skills format, so tools like Claude Code, GitHub Copilot, and Gemini CLI can load them and apply them while generating or reviewing code. npx skills add AzureCosmosDB/cosmosdb-agent-kit installs it. The categories track the same ground this series covered: data modeling and partition keys rank CRITICAL, queries and SDK usage rank HIGH, vector search and full-text search get their own dedicated categories.

Nine Things the Repo Already Gets Right

Cross-checking models.py, schema.py, and checkpointer.py against the kit’s rule files turned up a solid list of matches. The hierarchical partition key orders tenantId before threadId, broad to narrow, exactly as partition-hierarchical recommends. The composite turn ID uses: as a separator — the kit’s model-id-constraints rule calls out #, ?, /, and \ as characters that break Cosmos DB’s REST auth signing, and: sits on its explicit safe list. TTL follows the kit’s own “correct” pattern precisely: container default -1, item-level ttl overriding it per turn. The vector embedding policy and DiskANN index both match vector-embedding-policy and vector-index-type field for field, and the indexing policy excludes the embedding path from the regular range index to avoid double-indexing cost.

The full-text policy uses en-US, case-sensitive, as the kit’s fts-define-policy rule insists, and the content field sits in fullTextIndexes without also cluttering excludedPaths incorrectly. Turns live in their own container, separate from checkpoints, precisely the pattern-langgraph-chat-history-separate rule, which exists because checkpoint blobs make poor chat history. And checkpointer.py reaches for a point read wherever it already has both the id and the partition key, rather than running a query that costs roughly 2.5x more.

Three Things It Would Flag

Not everything cleared. seed.py never normalizes its mock embeddings to unit length vector-normalize-embeddings flags exactly this, and it’s a genuine miss, not a style preference: unnormalized vectors produce inconsistent cosine-similarity scores, and the fix is one line (v / ||v||₂) that never made it into the SHA256-based fake_embedding() function. Second, the checkpointer is a hand-rolled BaseCheckpointSaver implementation against the synchronous azure-cosmos SDK, not the official async CosmosDBSaver from langchain-azure-cosmosdb that sdk-langchain-cosmosdb-saver recommends. That one’s a deliberate tradeoff rather than an oversight — building it by hand is what made post 4’s checkpointing section possible to explain from the inside. Still, a production app should almost certainly reach for the maintained package instead. Third, turn upserts carry no ETag check so that sdk-etag-concurrency would flag the read-modify-write path as vulnerable to lost updates under concurrent writes. The demo’s single-writer pattern never triggers it, but two agents racing to update the same turn would silently drop one of them.

Where a Rulebook Like This Actually Helps

A tool that catches “your mock embeddings aren’t normalized” in seconds, instead of after a confusing test failure, earns its place in a workflow. What it doesn’t replace is the reason each rule exists — Cosmos DB Agent Kit tells you: is a safe ID separator, but this series spent a paragraph on why # breaks HMAC signing specifically in Gateway mode. Both matter: the rulebook for speed, the explanation for judgment calls like the checkpointer tradeoff above, which no automated check can make for you. Six posts of hand-checked Cosmos DB rules and a 100-rule kit landed on the same answers almost everywhere — that convergence is worth more than either source alone.


Sources

Rate Limiting Algorithms on Azure: 6 Compared

Your API works perfectly, until someone hammers it with 10,000 requests in a second. Rate limiting stands between a stable system and an outage, and the usual list of rate limiting algorithms has six entries: fixed window counter, sliding window log, sliding window counter, token bucket, leaky bucket and concurrency limiter. However, few guides show what each one actually lets through, or where it should run. So I built all six in C# on Azure, load-tested them, and wrote down what I learned. The code is in github.com/steefjan1/rate-limiting-azure.

Four of the six are already in .NET

In fact, on .NET you mostly do not implement these. System.Threading.RateLimiting ships the fixed window, sliding window, token bucket and concurrency limiter; ASP.NET Core wires them in with AddRateLimiter and RequireRateLimiting. The whole fixed window:

options.AddPolicy("fixed-window", ctx =>
RateLimitPartition.GetFixedWindowLimiter(PartitionKey(ctx), _ =>
new FixedWindowRateLimiterOptions { PermitLimit = 10, Window = TimeSpan.FromSeconds(1) }));

The missing two, sliding window log and leaky bucket, are small RateLimiter subclasses in the repo. The log keeps a queue of timestamps per client. The leaky bucket, by contrast, keeps one number, the next free drain slot, and either holds the request until then or rejects it when the queue is full, making it the only limiter that adds latency on purpose.

One nuance: .NET’s sliding window is segmented, not the weighted “current plus previous window” version most explanations describe; the repo has the weighted one in Redis.

What the load test shows

Every endpoint gets the same budget: 10 requests per second per client. What matters most is the highest number of requests the backend saw inside any one-second span.

Scenario 1: a burst across the window boundary

First, a burst across a window boundary: one request opens the window, nine more arrive at 900 ms, ten more at 1100 ms.

AlgorithmMax accepted in any 1 sRejected of 20
fixed window190
sliding window log109
sliding window counter1010
token bucket181
leaky bucket110 (the last ones waited 999 ms)

The fixed window result is the textbook flaw, measured: a “10 per second” limit let 19 through in 200 ms. The log and the segmented counter, by contrast, hold at 10.

The token bucket result surprised me. A full bucket plus a couple of refilled tokens should give 12 or 13, not 18. The cause is in .NET: when the bucket is full, TokenBucketRateLimiter.ReplenishInternal returns early without moving its last-replenishment timestamp, and the PartitionedRateLimiter that hosts it refills on elapsed time. So after a quiet spell, the first refill credits all the time the bucket spent full, roughly doubling your burst allowance. The Redis token bucket in the repo, however, updates the timestamp on every call and does not do this.

Scenario 2: sustained overload

Next, sustained overload: 30 requests per second for three seconds. Fixed window and sliding log each accept 30 of 90; the segmented counter accepts 26. The token bucket accepts 39 by design, since capacity sets the burst and refill rate sets the average. Meanwhile, the leaky bucket accepts 40 with a p95 latency of 1002 ms, because it queued ten and drained them at the fixed rate.

Scenario 3: concurrency at a slow backend

Finally, 20 simultaneous calls at a backend that takes 250 ms. Without a limiter, all 20 hit it at once; with the concurrency limiter at 4, four get through and sixteen get a 429 within 2 ms. In other words, a service handling 10,000 cheap requests a minute can still fall over when 500 slow ones run at once.

The decisions that come before choosing a rate limiting algorithm

Which layer

“Combine two or three at different layers” is advice everyone gives and nobody draws. Here is mine.

On Azure, the gateway is API Management, where two of the six already live as policy XML: rate-limit-by-key counts requests per key, and limit-concurrency caps in-flight calls to your backend. Through the deployed gateway, the load test gave the same numbers for 2 to 5 ms of added latency, and APIM rejected nothing, since its policy counts only successful responses. Note, though, that rate-limit-by-key is a sliding window on the classic tiers and a token bucket on the v2 tiers: same XML, different burst behaviour.

Which unit

Ten cheap reads and ten model calls put different pressure on a system, so one threshold everywhere is the wrong shape. Instead, limit against the bottleneck: in-flight calls for a slow dependency, tokens for an AI backend, connections for a database. For example, APIM’s llm-token-limit is the token bucket with tokens as currency, and the RateLimiter API takes a permitCount, so an expensive endpoint can cost five permits while a cheap one costs one.

Where the counter lives

Every in-process limiter counts per instance. Because the sample deploys the API to Container Apps with two replicas on purpose, the same load test shows what that means: the fixed window and the sliding log each accepted 60 of 90 instead of 30, and the concurrency limiter let 8 calls reach the slow backend instead of 4. Every configured number doubled, silently, as one Log Analytics query over the API’s rejection log makes visible:

Each in-process policy rejected about 30 of 90, split across two replicas, while each Redis-backed policy rejected about 60, from one counter. APIM has the same shape: counters are per gateway node, never aggregated. Since the only way to get one counter per client across replicas is shared state, the repo runs each algorithm as one atomic Lua script in Azure Managed Redis: INCR plus PEXPIRE for the fixed window, a sorted set for the log, two weighted keys for the counter, a hash for the token bucket. Against the same two replicas, the Redis versions accepted 30, 30, 29 and 39. It costs one round trip per request, including rejects, but added no measurable latency in the same region.

Shared state, however, brings two failure modes of its own. First, clocks: the scripts take “now” from the calling replica, so drifting clocks disagree about the window; if that matters, read TIME inside the script and let Redis be the clock. Second, network: the limiter fails open when Redis is unreachable, because failing closed would be a worse outage than the one it prevents. That is a choice, and it needs an alert. Also, Azure Cache for Redis closes to new creations on 1 October 2026, so the sample uses Azure Managed Redis, clustered by default with hash tags to keep a client’s keys in one slot.

Who is being limited

Per IP punishes everyone behind a corporate NAT, and per subscription is what APIM gives. Per-user needs a validated token, so the limiter sits after authentication, and unauthenticated floods reach your identity provider. So real systems combine more than one key: IP at the edge, subscription at the gateway, user in the service. Fairness between tenants is policy, not algorithm. The sample uses an X-Client-Id header that APIM sets from the subscription ID, so the gateway and service partition on the same identity.

What the client gets

The client needs a 429 with Retry-After, or it retries at once and your limiter becomes a load generator. Every limiter computes Retry-After from its own state: the window’s remaining time, the log’s oldest entry, the leaky bucket’s next free slot.

The other half is on the caller: honour the header, add jitter, and tune retries with the limits they will hit. Microsoft.Extensions.Http.Resilience does the first two out of the box.

What you watch

Rejections and saturation are different signals. A rising 429 count per policy says a client is over budget, while a concurrency limiter pinned at its permit limit, or a leaky bucket queue that never drains, says the system is at capacity. So the sample tags every rejection with X-RateLimit-Policy and ships APIM gateway logs to Log Analytics. The repo’s docs/kql.md has the queries, starting with the one that says whether the gateway or the service produced a 429 (BackendResponseCode empty versus 429).

Where each one is the wrong answer

A fixed window is wrong when the thing you protect cannot survive 2x for a moment. A sliding log, however, is wrong at high limits: 10,000 per minute means 10,000 timestamps per client. A token bucket is wrong when the downstream needs a smooth rate rather than an average rate. A leaky bucket is wrong at an HTTP edge where clients time out before they drain; it belongs in front of a fragile dependency you own. A concurrency limiter is wrong for fairness between clients, because it says nothing about rate. And every in-process limiter is wrong the moment you have two replicas and still expect the number you configured.

What I would do

Token bucket at the gateway, keyed per client, loose enough that honest bursts pass. Concurrency limiter at the service, in front of the slow thing, sized to what it can take. For more than one replica, use a per-client counter in Redis, and return Retry-After so your own clients respect it. Two rate limiting algorithms are usually enough for a public API on a side project. Then run the load test, because two of the five let through nearly twice what their configuration says.

Foundry Model Ledger: Which Model, at What Price, Until When, and Which of Mine

Every model conversation with a team ends with the same four questions. Is the model available in our region? What can it do? What does it cost per million tokens? And when does Microsoft retire it? Then the platform team asks a fifth: which of our own deployments are affected? Microsoft’s Foundry Model Explorer answers the first question well. However, it does not answer the other four. It cannot, because the answers live in your subscription and in a price list that was never meant to be read by software. So I built the Foundry Model Ledger to read them together. The interesting part is what the data looks like once you do.

What the Foundry Model Explorer already does

Microsoft’s explorer is a region availability reference. It shows one row per model and version, with lifecycle status, deployment SKUs, retirement date, and the list of regions that carry it, exportable to CSV. If your question is “where can I deploy gpt-5.6-terra”, it is the fastest answer there is. However, it is a periodic snapshot rather than a live read. It shows no prices in any currency. And it knows nothing about what you have deployed. Those three gaps are the Foundry Model Ledger.

Three sources, one Foundry Model Ledger

The model catalog lives in Azure Resource Manager. A single call, GET /subscriptions/{id}/providers/Microsoft.CognitiveServices/locations/{region}/models, returns every model and version a region carries. Each entry has its capabilities (chat completion, tool calling, embeddings, context window), its lifecycle status, its deployment SKUs, and its retirement date under deprecation.inference. This is the same data the Foundry portal shows. Moreover, it needs nothing more than Reader on the subscription.

Prices live in the Azure Retail Prices API. It is public, needs no authentication, and returns every meter Azure bills. Filter on serviceName eq 'Foundry Models' and a region, and you get the list prices for tokens, images, and hours, in any currency the API knows.

Your own deployments live under each Cognitive Services or Foundry account. At first I reached for Azure Resource Graph, which is the natural place to list resources across a subscription. It returned zero rows against a subscription with twelve accounts. Resource Graph does not index the accounts/deployments child type. So the Ledger lists the accounts through ARM and then calls each account’s /deployments endpoint, six at a time.

How the Foundry Model Ledger is built

The Foundry Model Ledger is a .NET 8 isolated Azure Function on Flex Consumption with a single-page UI. It deploys with azd up, in the same shape as the [INTERNAL LINK: RAG in 8 Steps on Azure] sample. A user-assigned managed identity with Reader on the subscription reads the first and third source. The second source needs no identity at all.

Pick a region and a currency, and the table shows model, version, lifecycle, capabilities, and retirement date with a days-left bar. Next to those sit the input and output prices per million tokens. Click a row for every SKU with its capacity range, every capability ARM reports, and every price meter the tool matched. The same panel has a button that checks every other region for the same model and version.

Finally, a second tab joins your deployments to the catalog of their own region and sorts them by soonest retirement.

What the data says

Sweden Central, on the day I took these screenshots, carried more than 300 catalog entries from 11 publishers. 216 were generally available, 66 in preview, and 16 marked as deprecating. 70 versions retire within 90 days. Furthermore, 30 are already past their retirement date but still listed. That last group matters. The catalog endpoint tells you what the region knows about, not what you can still deploy. For example, gpt-4o-mini has been closed to new deployments for a long time and still appears. The SKU list and the retirement date are the better signals.

The same model and version often appears twice, once for account kind OpenAI and once for AIServices, each with its own SKU list. The Foundry Model Ledger merges those into one row. If you script against the endpoint yourself, expect the duplicates.

The Retail Prices API returned 1,727 meters for Foundry Models in that one region. None of them carries a model identifier.

Prices are written for invoices

A meter name is a billing label, not a key. The same model, gpt-5.6-sol, is spread across meters such as 5.6 sol ShortCo Inp Std Gl 1M Tokens, 5.6 sol LongCo Cd Wr PP DZ 1M Tokens, and 56sol ShCo Cd Wr Fl Gl 1M Tokens. Older meters read gpt 4.1 nano cached Inp glbl Tokens and bill per 1K tokens; newer ones bill per 1M. Similarly, grok-4.6 appears as 4.6 Inp DZ Tokens under the product Azure Grok Models, with no “grok” in the meter name at all.

The abbreviations are their own dialect. Input is Inp, inpt, or in. Output is Outp, opt, outpt, or out. Cached input is Cd. Global, data zone, and regional are Gl, DZ, and regnl. Batch, priority, and flex tiers have their own tokens.

How the matcher works

So the Foundry Model Ledger has a matcher. First, it narrows meters to the model’s publisher. Then it tokenizes both sides the same way and requires every token of the model name (minus the family prefix) to appear in the meter, in order. It also rejects meters that carry a sibling variant such as mini or pro the model does not have. When a meter carries a date token, o3 0416 or chat-latest 08062026, it pins the version. Next, it classifies direction, deployment type, tier, and context length, and normalizes the unit to a price per million tokens. Finally, it picks a headline: global standard, short context, uncached. Every price in the table carries a confidence label, exact, name, or loose. In addition, the detail panel shows every matched meter, so the headline number is never the only evidence.

What the first live run got wrong

On the first live run it priced 204 of the entries in Sweden Central. The misses were instructive. FLUX image models came out at “40,000 per million” because their meters bill per 1K images, and I had treated every 1K unit as tokens. Qwen sits under product Qwen models, which the publisher map did not know. Moreover, most of its meters are fine-tuning meters that must never become a headline price. And the Anthropic models, plus Cohere rerank and parse, are in the catalog with no Foundry Models meter in the region at all. The model exists; the public price does not.

Two tabs that came from using it

The first version stopped at the tabs above. Two more followed within a day, because the first questions people asked were not the ones I had built for.

Where else is this model available

The region check in the detail panel gives a list of names. That is correct and hard to read. So the Foundry Model Ledger now has an availability map. Every Azure region is plotted from the subscription’s own location metadata, on a world outline from Natural Earth. Type a model name, pick a version, and the regions that carry it light up in the brand blue. Regions without Azure AI services show as a dashed ring, and regions that host the service but not the model stay grey. Zoom presets for Europe, North America, Asia Pacific, the Middle East and Africa, and South America keep the labels readable where regions cluster. In Europe alone there are twenty. The data behind the map is one lightweight catalog call per region, cached, so a check across every region takes a few seconds the first time and is instant after that.

What else does the same job for less

The second question is the one that decides budgets. If we use this model for that workload, what else in the region can do the same job, and what would it cost instead? The Alternatives tab answers it with the data the Ledger already has. Pick the model you use or consider. Its capability flags from the catalog, such as chat completion, tool calling, JSON schema output, and image input, become chips, and every chip is required by default. Click one off if you do not need it and the list widens. Every other model in the region that reports all remaining capabilities is listed with its list price, what your monthly workload would cost on it, and the delta against your model, cheapest first. Enter the workload as millions of input and output tokens per month. Filters restrict the list to the same publisher or to generally available models, and retired versions are excluded. Other versions of the same model are listed last, so they do not pose as alternatives.

The comparison is on capabilities, not on quality. A nano model will always look like a 99 percent saving next to a frontier model, and the table does not know whether it would pass your evaluation set. What it does give you is the shortlist and the price gap in one view, which is the part that used to take an afternoon with the pricing page open in three tabs.

Where the Foundry Model Ledger is the wrong answer

Do not budget on it. These are list prices, matched heuristically, in whatever currency you pick. The matcher will be wrong somewhere the day Microsoft renames a meter. It also cannot price provisioned throughput, which is billed per hour per unit rather than per model. Use it to see the shape of a decision, then confirm on the pricing page.

Do not put it on the internet as is. The Function has no authentication, and it exposes your deployment list to anyone with the URL. Keep it internal, or put Easy Auth with Entra ID in front of it. If you already run model endpoints behind API Management, the same gateway can front the Ledger, with the policies from the APIM for AI Workloads series.

And do not read the catalog as an availability promise. Listed does not mean deployable, and a retirement date is a floor, not a schedule. For a plain “where is it available” question, Microsoft’s explorer remains the quicker tool.

What I would build next

The tab I did not plan is the one people ask about: my deployments, joined to the catalog, sorted by retirement. That is a governance question, not a browsing question. It sits next to the routing questions from Don’t Build Around Today’s Model. Build for the AI Control Plane. Consequently, the natural next steps are all on that side. All subscriptions in a tenant instead of one. An alert when a deployment sits on a version retiring within 90 days. A history, so you can see what appeared and what closed in a region since last month, which is a changelog Microsoft does not publish. And a cost delta for moving each deployment to its successor, which the Alternatives tab now does for one model at a time and should do for the whole deployment list.

The Foundry Model Ledger repo is at github.com/steefjan1/foundry-model-ledger. azd up deploys it; scripts/run-local.ps1 runs it against your az login; node tools/mock-server.mjs runs the UI on a sample snapshot without .NET or Azure. If the matcher misprices a model in your region, the detail panel shows why. An issue with that screenshot is the fastest way to get it fixed.

Cosmos DB Agent Memory Cost: Caching, RU Drivers, and a Pitfalls Roundup

Post 6 of 6 on Cosmos DB agent memory cost: the hard numbers post 1 promised back at the start of this series.

Post 1 opened with a 2017 CloudBrew reviewer calling a Cosmos DB proof of concept “an hour-long marketing pitch,” and my answer then was that cost means nothing without the revenue it enables. Five posts later, that argument still needs the numbers behind it.

Where This Post Picks Up

This post closes the series with them: what actually drives Cosmos DB agent memory cost at the RU level, how semantic caching cuts LLM spend specifically, what to monitor once an agent runs in production, and a pitfalls roundup that pulls every thread from posts 2 through 5 into one list.

Semantic Caching: Reusing What You Already Paid to Compute

An LLM call is almost always the most expensive, highest-latency step in an agent’s request path, far more than any Cosmos DB read or write. A semantic cache cuts that cost by skipping the LLM entirely when a close-enough answer already exists. Instead of matching prompts by exact string, it vectorizes the incoming prompt and runs a similarity search against the prompt-completion pairs already sitting in the cache. The mechanics are the same as the VectorDistance() query from post 3; only the container changes, from memory to cache.

Two details make this different from a normal cache, and both matter for cost control. First, the similarity threshold is a real trade-off, not a default to leave alone: set it too high, and near-identical questions still miss and hit the LLM anyway; set it too low, and the cache starts returning answers that don’t actually match what the user meant. Second, a semantic cache needs the same context window as an LLM.

Cache only the raw prompt, and two different users who each ask “what’s the second largest?” in unrelated conversations get whichever answer the cache stored first, correct for one thread, wrong for the other. Vectorize a slice of the conversation history alongside the latest prompt, the way post 2’s turn-based schema already structures it, and the cache lookup carries the same context the LLM would have used. TTL handles cleanup the same way it does for turns in post 2, with one addition worth considering: a hit-count field that increments on each cache hit lets a pruning pass keep frequently reused entries around longer than questions the cache only ever answered once.

What Actually Drives Cosmos DB Agent Memory Cost

Four decisions drive most of the RU bill for an agent workload, and they’re not evenly weighted. Partition skew usually costs the most: at Cosmos DB Conf 2026, an engineer described a production account running at 100% RU utilization, throttling and retrying under load, where the obvious fix looked like provisioning more throughput. The real cause turned out to be a single logical partition absorbing over 80% of traffic, one automated integration account driving most writes under a partition key that looked reasonable on paper. Fixing the data model, without adding a single RU of throughput, dropped utilization to 20–35% and made the throttling disappear entirely. More throughput would have masked that problem, not fixed it.

Item size and shape matter next, and this series already covered the mechanism in post 2: one document per turn keeps writes small and cheap. At the same time, one-document-per-thread turns every new message into a full-item rewrite that gets steadily more expensive as the thread grows. Vector index choice is the third lever. DiskANN’s sharding and approximate search solve a scale problem post 3 already flagged, and paying for that complexity below roughly ten thousand vectors buys nothing quantizedFlat wasn’t already providing. TTL is the fourth: expired short-term memory that never actually expires, because someone set a container-level default once and never came back to it, quietly inflates storage and index size on data nobody queries anymore.

The Fifth Lever: Consolidation

There’s a fifth lever underneath all four, and it’s the one this whole series has been arguing for since post 1: consolidation. Running a cache, a relational store, and a dedicated vector database as three separate systems means paying for three separate throughput allocations, three separate operational surfaces, and cross-system network cost on every request that touches more than one of them. One Cosmos DB account carrying memory, search, and cache together shares throughput across all three instead of over-provisioning each in isolation; the 2017 CloudBrew critique missed the same argument when it judged the account’s line-item cost without asking what running three systems instead of one would have cost by comparison.

Monitoring: What to Watch Once It’s Running

Three signals catch most problems before they become an incident. Change feed lag matters most for the multi-agent handoffs post 4 covered — a growing lag between a write and the Function that reacts to it means a specialist agent is falling behind the conversation, not just running a little slower. Break RU consumption out per container instead of watching one account-wide total, and it shows which specific workload is driving spend — turns, checkpoints, or the semantic cache — instead of leaving that as a guess. And for catching expensive patterns before they ship at all, the Azure Cosmos DB VS Code extension’s Query Insights and Index Advisor flag cross-partition queries, missing filters, and indexing gaps directly in the editor, well before a query shape becomes production traffic.

Pitfalls Roundup: Every Thread from This Series

  • Unsharded vector index in a multitenant app (post 3) — without a vectorIndexShardKey, semantic search scans every tenant’s vectors, not just the current one.
  • Thread-per-item growth (post 2) — an item that grows by one append per turn gets more expensive to write with every message, and eventually hits a hard size limit.
  • Missing or forgotten TTL (post 2) — short-term memory nobody set an expiration for keeps sitting in the container indefinitely, quietly inflating storage.
  • Cross-tenant memory leakage (posts 3 and 4) — a global vector index or an unscoped checkpoint container lets one tenant’s context bleed into another’s.
  • Treating change feed as globally ordered (post 4) — ordering holds within a partition key, never across the whole container.
  • Confusing Foundry Agent Service Classic and New containers (post 5) — the newest trap in the list, and already the most common source of “why is my thread storage empty” reports.

What I’d Ask the Product Team

Multi-region writes for a globally distributed agent multiply throughput cost by the number of regions. The guidance so far is “add regions only where traffic justifies it,” which is reasonable. Still, it leaves the actual crossover point (how much traffic, at what latency requirement) for each team to work out through trial and error rather than a documented formula. A cost calculator that takes a workload shape and a target latency and outputs a recommended region count would save a lot of that guesswork.

Where This Is the Wrong Answer

Not every agent workload belongs on one account. A workload with one enormous, narrowly specialized vector search needs tens of billions of vectors; nothing else can still get better unit economics from a dedicated vector database that specializes in exactly that shape, rather than a general-purpose store carrying memory, search, and cache together. The unified argument holds for the vast majority of agent workloads this series has covered, not for every workload unconditionally.

Closing the Series

That 2017 reviewer wasn’t wrong that the account cost more than a bare-minimum alternative; the miss was judging that cost without the workload it made possible, the same mistake the Figma AWS costs piece argued against in a completely different context. Six posts and one real production case study later, Cosmos DB agent memory cost comes down to the same handful of decisions this series has covered since post 2: partition key, item shape, index choice, and TTL, with semantic caching and consolidation compounding the savings on top. That’s the whole series in one sentence, and it’s the argument I’d have made at CloudBrew in 2017 if I’d had the RU numbers to back it up yet.


Sources

Cosmos DB Foundry Agent Service: Bring-Your-Own Thread Storage

Post 5 of 6 on Cosmos DB Foundry Agent Service integration, owning the thread store instead of leaving it opaque behind a managed API.

Posts 1 through 4 assumed you manage the Cosmos DB account directly. Foundry Agent Service changes that assumption by default: spin up an agent the basic way, and Microsoft manages the thread store for you, out of reach of a direct query. Standard agent setup flips that around. Cosmos DB Foundry Agent Service integration lets threads, system messages, and agent metadata land in a Cosmos DB account you own, sitting right where the schema, search, and checkpointing patterns from the rest of this series already apply.

Why Bring-Your-Own Thread Storage Matters

Data residency, security review, and auditability all get harder when a vendor holds conversation history in a store you can’t query. Standard setup solves that by provisioning three customer-owned resources instead of one managed black box: Azure Storage for uploaded files, Azure AI Search for the agent’s vector stores, and Azure Cosmos DB for everything: conversational messages, threads, and agent metadata. Cosmos DB carries the load that matters most for this series: it’s where the actual conversation lives.

Inside enterprise_memory: Cosmos DB Foundry Agent Service Containers

Standard setup names the resulting database enterprise_memory, and container names inside it depend entirely on which Foundry Agent Service runtime the agent runs on. Foundry Agent Service (Classic) writes to three containers: thread-message-store for end-user conversation messages, system-thread-message-store for internal system messages, and agent-entity-store for agent metadata like instructions and tools. Foundry Agent Service (New) writes to two different containers instead of agent-definitions-v1 and run-state-v1, and neither runtime reads the other’s containers. Check thread-message-store for an agent running on the New runtime, and it comes back empty, not because BYO thread storage failed, but because the data landed somewhere else entirely.

Provisioning Standard Agent Resources

Provisioning means more than a Cosmos DB account on its own. Standard setup also expects an Azure Storage account, an Azure AI Search resource, and an Azure Key Vault for secrets, alongside a deployed agent-compatible model. Once those exist, Microsoft’s Bicep template accepts the resource IDs of existing accounts and wires up the rest: account and project connections, role assignments, and the capability hosts that tell Foundry where agent state actually lives

Two things catch people off guard here, so it’s worth flagging both before you deploy. First, throughput: your Cosmos DB account needs at least 3,000 RU/s total 1,000 RU/s for each of the three baseline containers and that number scales up with every additional project sharing the account, since each project gets its own container set. Undershoot it, and the deployment doesn’t fail quietly; it throws CapabilityHostProvisioningFailed during the capability host step. Second, roles: the project’s managed identity needs Cosmos DB Operator at the account level to provision containers, plus Cosmos DB Built-in Data Contributor at the database level for enterprise_memory. The database-level scope covers every container inside it, so a single role assignment handles the whole set instead of one per container.

Querying Thread History Directly

This is what bring-your-own thread storage actually buys over the default: a direct line into conversation history that Foundry’s own API doesn’t expose. Once a thread exists in thread-message-store (Classic) or run-state-v1 (New), you can run the same vector, full-text, and hybrid queries from post 3 against it — same RANK RRF(...) syntax, same partition-scoped WHERE clause. The only difference is the container: Foundry manages it instead of your own code.

Treat this as a connection into Foundry Agent Service, though, not a replacement for it. Foundry still owns thread creation, run orchestration, and tool invocation; Cosmos DB Foundry Agent Service integration only changes where the resulting data sits and who can query it directly.

Pitfalls

Checking the wrong container set. The single most common source of “BYO thread storage isn’t working” reports is querying Classic’s containers for an agent running on the New runtime, or vice versa. Confirm which runtime a project uses before assuming a missing thread means a broken connection.

Underprovisioning throughput for multiple projects. The 3,000 RU/s floor covers one project’s container set. Add a second project to the same Cosmos DB account, and the container count and the RU/s requirement under it doubles. CapabilityHostProvisioningFailed almost always traces back to this, not to a misconfigured connection.

Assuming you can edit a capability host after creation. You can’t. Pointing a project capability host at the wrong Cosmos DB resource ID means deleting and recreating the project, not patching the connection — worth getting right on the first deployment rather than treating it as a setting to adjust later.

Next: Cost, Caching, and Production Pitfalls

That settles Cosmos DB Foundry Agent Service integration for readers who need full ownership of thread storage rather than a managed default. The last post in this series pulls back to the practitioner-notes view: RU cost drivers, semantic caching, and a closing pitfalls roundup across everything posts 2 through 5 have covered.


Sources

Five RAG Architectures in Real Azure Code

Over the past few months I kept running into the similar looking infographics, in one form or another: five or six boxes, each a named RAG architecture, arrows showing how a query flows through it. Hybrid RAG. GraphRAG. Agentic RAG. Corrective RAG. Multimodal RAG. They’re useful as vocabulary. They are not implementation guides. None of them show you the part that actually takes the time.

So I built all five RAG architectures, on Azure, against one shared corpus and one shared set of test questions, and measured what came out. This post is the result: what each diagram leaves out, what the equivalent Azure code actually looks like, where the real deployment pitfalls were, and a comparison table built from real runs, not from argument.

The corpus is a fictional Dutch health insurer, Zorgverzekeraar Meridiaan, the same one I’ve used in a couple of other posts in this series. Nine documents: dental and physiotherapy policies, a provider network, an authorization process, a member complaint and the quarterly report that restates it, a stale FAQ sitting next to the current policy, a reimbursement table, and a scanned claim form. Fifteen questions, tagged by which pattern they were designed to stress. All five patterns answer all fifteen questions, so the comparison is apples to apples.

The repo is at github.com/steefjan1/five-rag-patterns if you want to run it yourself.

What the diagram shows vs. what the Azure code does

Hybrid RAG

The diagram draws dense and sparse retrieval as two separate paths that merge into a box labeled Reciprocal Rank Fusion. That box is mostly a non-event on Azure. Azure AI Search’s hybrid query type takes a vector query and a text query together and fuses them server-side. There is no RRF code to write.

What actually takes engineering effort is the index schema: chunk granularity (I chunk by document section, not by a fixed token window, so a retrieval unit is a coherent answer, not an arbitrary slice), which fields are filterable versus searchable versus vector, and whether semantic ranking is worth its cost on top of the fusion you already get for free.

Measured: recall 1.00 across all fifteen questions, the best of any pattern on pure coverage. Precision sits at 0.38, diluted by a fixed top-5 retrieval regardless of how many documents a question actually needs. Cheapest sane baseline in the set: $0.0026 and 2.00 seconds per query.

GraphRAG

The diagram draws one static graph: entities, edges, a subgraph retrieval step, a box for community summaries. What it doesn’t draw is that the graph has a maintenance cost. Community detection (Louvain, via networkx, which runs fine at this corpus’s scale without a dedicated graph database) and community summarization are real compute and real Azure OpenAI spend, paid once at setup and again every time the graph changes enough to shift community boundaries. Nothing about that shows up in the box-and-arrow version.

Entity linking here is a cheap substring match against entity names, not an embedding call, which is part of why this pattern is the cheapest per query in the whole set. Retrieval is a graph walk: two hops, both directions, so a question like “which hospital did this referral come from, and which GP group refers into that hospital” resolves correctly even though no single document states the answer. It’s two separate edges, walked in sequence.

Measured: recall 1.00 on its own three relational questions, the two-hop case included, and 0.21 on the other twelve. No other pattern swings that hard between its own territory and everything else. $0.0013 per query, the cheapest pattern here, in the narrowest lane.

Agentic RAG

The diagram shows a planner routing to tools and a reasoner that loops “until confident.” There is no upper bound drawn anywhere on that loop. Left alone, that is a cost leak, not a reliability feature, so the actual implementation caps it at five iterations and reports hitting the cap as its own outcome rather than quietly forcing an answer and calling it clean.

The other thing worth knowing if you’re building this on Azure: the AI Foundry Agent Service SDK bypasses API Management for its own LLM calls. I found this the hard way on an earlier project in this series. If your governance model depends on APIM, that means routing tool-calling agents through the standard OpenAI SDK pointed at the gateway, not through the framework’s own agent runtime, or every rate limit and kill switch you built stops applying the moment the agent framework makes the call instead of your code.

Two of this pattern’s four tools aren’t retrieval at all. The dental waiting-period and annual-maximum arithmetic is transcribed from the policy documents as plain code, not left for a language model to compute from prose. Insurance eligibility math is exactly the kind of thing an LLM gets subtly wrong under pressure, and exactly the kind of thing code gets right every time.

Measured: precision 1.00, recall 1.00 on its own two questions, and it’s the only pattern that doesn’t collapse elsewhere: recall 0.85 on the other thirteen, because it always has a general search tool as a fallback when nothing more specific fits. That’s the real finding here. It’s not that Agentic RAG is “better,” it’s that it hedges.

Corrective RAG

The diagram shows retrieve, grade, then three branches: answer, rewrite the query and loop back, or fall back to a web search. The rewrite loop has an arrow pointing backward and no stated exit condition. A closed corpus also has no web to fall back to, so “incorrect” here means declining to answer rather than guessing.

The corpus has a document built specifically to test the grading step: an archived FAQ with a plausible, wrong number sitting right next to the current policy with the right one. A pattern with no grading step retrieves both and may cite either. This one grades the retrieval, asks the model to identify which passage is authoritative using published dates and explicit supersession language, and only feeds the authoritative passages to the final answer. What got fetched and what got used are tracked separately on purpose, so a working grader shows zero distractor citations even though the distractor was retrieved.

Measured: precision 1.00, recall 1.00 on the three distractor questions, confirmed live, not just in the design. The more interesting number is that its other twelve questions score better (precision 0.79) than its own target slice (0.67). The grading discipline isn’t just catching the one distractor it was built to catch, it generalizes. That comes at a real cost: 3.97 seconds average latency, roughly double every other pattern, and the highest cost per query in the set, because a full run can mean three model calls instead of one.

Multimodal RAG

The diagram’s box says “shared multimodal embedding model (e.g. CLIP or ColPali),” which means self-hosting an embedding model. That’s a heavier operational commitment than anything else in this comparison needs, and it’s avoidable. This uses caption-then-embed instead: Document Intelligence extracts the actual structure of the reimbursement table (tables are exactly where a vision model hallucinates a plausible-looking row that isn’t in the source, so that step doesn’t get skipped), the vision-capable chat deployment captions the scanned claim form directly, and both captions get embedded with the same text-embedding-3-large deployment every other pattern uses. Same index Hybrid RAG built, two more documents in it, no new index and no new field.

One implementation note that cost real iteration: a first version of the captioning prompt asked for verbatim transcription, which correctly produced the form’s Dutch date format and Dutch status text. A validation step checking the caption against a hand-written ground truth flagged that as a mismatch, because the ground truth expected ISO dates and English. That’s not a captioning bug, it’s a prompt that needed to ask for normalization, not transcription. Worth deciding on purpose, since a shared index with mixed date formats and mixed languages retrieves worse than a normalized one.

Measured: recall 1.00 across the board and groundedness 1.00 on its own two questions, with no degradation on the other thirteen. That composability is the finding: this pattern is Hybrid RAG’s exact retrieve-and-answer loop plus two documents, and the numbers confirm that composition was free.

The comparison table

Same corpus, same fifteen questions, one pass, all five RAG architectures.

PatternAvg latencyCost/queryOverall precisionOverall recallOwn-target recall
Hybrid2.00s$0.00260.381.001.00 (n=5)
GraphRAG2.32s$0.00130.200.371.00 (n=3)
Agentic2.09s$0.00450.450.871.00 (n=2)
Corrective3.97s$0.00510.770.971.00 (n=3)
Multimodal2.18s$0.00270.381.001.00 (n=2)

“Own-target” means the small subset of the fifteen questions each pattern was actually designed to answer (GraphRAG’s two-hop provider questions, Corrective RAG’s stale-document case, and so on). Every pattern hits recall 1.00 in its own lane. What separates them is what happens outside it: GraphRAG falls to 0.21 recall on the other twelve questions, Agentic RAG only falls to 0.85, and Hybrid, Corrective, and Multimodal don’t fall at all, because their retrieval isn’t scoped to a narrow entity set in the first place.

Two honest caveats on this table. Precision across every pattern is capped low by a fixed top-5 retrieval regardless of how many documents a question actually needs, so precision here measures retrieval breadth more than answer quality, read recall and the own-target column as the more meaningful columns. And the cost figures come from a placeholder price table, not a live Azure billing export, useful for comparing patterns against each other, not for a procurement conversation.

Deployment pitfalls

Every one of these was a real failure against a live Azure subscription, not a hypothetical.

Pinned model versions rot. A deployment written against gpt-4o-mini version 2024-07-18 failed eight months later with ServiceModelDeprecated. The fix wasn’t a newer pin, it was to stop pinning: leave the deployment’s model version empty and let Azure resolve the current default, and check az cognitiveservices model list -l <region> -o table before assuming a model name is still offered at all.

The account kind changed. Azure OpenAI is now provisioned through Foundry as kind: 'AIServices', not the older kind: 'OpenAI'. Same deployment mechanism underneath, different account kind and a newer API version. A template written against the old kind fails Cognitive Services preflight validation, not at compile time.

A malformed policy XML fails at ARM validation, not at Bicep build time. An APIM policy embedded as a Bicep string had a raw double-quoted path literal sitting inside an already double-quoted XML attribute. bicep build compiled it clean, because Bicep has no way to know a string is meant to be well-formed XML. The actual break only showed up against the live ARM validation API, after Azure AI Search, Cosmos DB, and the Foundry account had already finished provisioning. A small script that compiles the template and separately parses every embedded policy string as XML catches this before the next azd up, not during one.

RBAC role assignments alone don’t turn on Azure AD authentication. Azure AI Search kept returning a flat 403 on every data-plane call despite two correctly scoped role assignments, because the service still only accepted API-key authentication. Nothing had told it to accept AAD tokens at all. The fix is a separate property, disableLocalAuth: true, on the search service itself. If a resource with roles that look correct still refuses an authenticated caller, check the resource’s own auth settings before re-checking the role assignment.

Where none of this is the answer

None of these five patterns is the right first move for a small, stable knowledge base. Plain vector search, no fusion, no graph, no grading, no agent loop, is the correct answer until you can name the specific failure mode you’re buying insurance against. Every pattern here is a bet against one kind of failure, and every bet has a cost attached whether or not you ever collect on it.

Don’t build all five for one real system either. Pick based on the failure mode your domain actually has. GraphRAG only pays for itself if your questions are genuinely relational, multi-hop, the kind no single document answers. If they’re not, you’re paying setup cost and getting a narrower Hybrid RAG. Agentic RAG’s flexibility costs a planning call before any retrieval happens at all, worth it if your questions genuinely vary in shape, wasted overhead if they don’t. Corrective RAG’s discipline costs roughly double the latency of everything else in this comparison. That’s a fine trade when a wrong answer is expensive and a two-second wait isn’t. It’s a bad trade for a chat widget where speed is the product.

What this actually proves

The infographic’s taxonomy is real. These five RAG architectures are genuinely different, with genuinely different failure modes, and that part of the diagram holds up. What doesn’t hold up is the implication that the hard part is choosing between them. The hard part, in every case, was the piece the diagram didn’t draw: RRF turned out to be free because Azure AI Search already does it, but community detection is not free and has to be redone as the graph changes. An agent loop needs a hard cap or it’s an open-ended bill. A query rewrite loop needs the same cap for the same reason. A self-hosted multimodal embedding model turned out to be avoidable entirely, caption-then-embed onto infrastructure you already have gets you most of the way there.

The single most useful number in this whole exercise might be the smallest one: Corrective RAG’s grading step scored better on questions it wasn’t built for than on the one it was. That’s a pattern worth paying attention to. The things that make a RAG system more disciplined in one specific place often make it more disciplined everywhere, not just in the place you were testing for.

If you want to see the actual failure modes up close rather than the aggregate table, the earlier posts in this series go deeper on two of them: what naive RAG diagrams leave out covers the hybrid retrieval and groundedness gaps in more detail, and choosing between RAG, GraphRAG, and Agentic RAG when auditability is the constraint makes the conceptual case this post backs with numbers.

The full repo, including the corpus, the eval harness, and every pattern’s implementation, is at github.com/steefjan1/five-rag-patterns.

Multi-Agent State and Checkpointing with Cosmos DB

Post 4 of 6 on Cosmos DB multi-agent state coordinating what several agents know about the same conversation, without a separate message bus.

Post 3 settled retrieval for a single agent working alone. This post is about what changes once a second agent enters the picture. Coordinating what several agents know about the same conversation turns out to be a different problem from storing and retrieving one agent’s memory, and Cosmos DB multi-agent state ends up resting on two mechanisms this series already covered: hierarchical partitioning from post 2, and change feed from post 1’s retail monitoring callback.

Shared but Separable: What Changes with Multiple Agents

A single agent needs one memory scope. Moreover, a multi-agent system needs two at once: shared memory that every agent can read and write for coordination, and private memory that lets each agent keep its own persona, prompts, and reasoning history separate from the others. Lose the separation, and agents start bleeding into each other’s context. Lose the sharing, and they can’t coordinate at all.

A triage agent, a product agent, and a specialist agent a common pattern in production multi-agent apps each hold their own scoped state. Still, all three write to the same underlying container, so any of them can pick up where another left off.

LangGraph Checkpointing on Cosmos DB Multi-Agent State

LangGraph’s checkpoint interface persists a graph’s state after every step, and Cosmos DB has more than one implementation of it: langgraph-checkpoint-cosmosdb on PyPI, and the checkpoint saver that ships inside langchain-azure-cosmosdb. Both plug into the same standard LangGraph pattern: compile the graph with a checkpointer, then pass a thread_id on every invocation:

from langgraph. graph import StateGraph
from langgraph_checkpoint_cosmosdb import CosmosDBSaver
checkpointer = CosmosDBSaver(
endpoint=cosmos_endpoint,
key=cosmos_key,
database_name="agentmemory",
container_name="checkpoints",
)
graph = StateGraph(AgentState)
# add_node / add_edge calls wire up triage -> specialist routing here
app = graph.compile(checkpointer=checkpointer)
config = {"configurable": {"thread_id": "contoso:thread-1234"}}
app.invoke({"messages": [...]}, config=config)

Encode tenantId:threadId into the thread_id string, and the checkpointer’s hierarchical partitioning lines up with the [tenantId, threadId] partition key from post 2 — the same pattern manages per-user, per-session state at scale, this time for graph checkpoints instead of turn-based memory items. Microsoft’s own multi-agent-langgraph sample builds a personal-shopper scenario on exactly this foundation: a triage agent routes requests, and a product agent answers them using retrieval-augmented generation against the same Cosmos DB account.

Change Feed as the Handoff Mechanism

Post 1 covered change feed as the primitive behind a 2023 retail monitoring solution, a new record in Cosmos DB firing a Function that could raise an incident. The same primitive coordinates agent handoffs: one agent writes a turn, a Function listening on the container’s change feed picks it up, and it hands the conversation to whichever agent should act next. No polling loop checks for new work; the write itself is the signal.

A minimal handoff trigger, using the turn-based schema from post 2:

python

import azure.functions as func
def main(documents: func.DocumentList) -> None:
for doc in documents:
if doc.get("targetAgent") == "specialist":
notify_specialist_agent(doc["threadId"], doc["turnIndex"])

The triage agent sets targetAgent on the turn it writes; the Function reacts to that write and wakes the specialist agent for that thread.

A Second Worked Example: Spring AI for Java Shops

Python and LangGraph aren’t the only path here. Spring AI 2.0 shipped with a Cosmos DB-backed vector store and memory integration for Java, and Microsoft’s multi-agent-spring-ai sample mirrors the LangGraph pattern in Java: multiple agents, one Cosmos DB account, the same shared-but-separable memory shape. Worth a look if the rest of the stack runs on the JVM rather than Python.

Pitfalls

Shared containers without tenant or session isolation. A checkpoint container that mixes every tenant’s graph state leaks context across customers the moment a partition or vector index goes unsharded; the same isolation failure post 3 flagged for vector search is now showing up in agent state instead of retrieved memories. Apply the same [tenantId, threadId] discipline to checkpoints that post 2 applied to turns.

Treating change feed as globally ordered. Change feed guarantees order within a single partition key, not across the whole container. Moreover, a handoff design that assumes “the Function always sees writes in the exact order they happened across every agent” breaks the moment two agents write to different partitions at close to the same time. Design handoffs so each step only depends on ordering within its own thread’s partition, not on a global sequence that Cosmos DB never promised.

Next: Cosmos DB Inside Microsoft Foundry Agent Service

That covers Cosmos DB multi-agent state when you manage the account directly. Post 5 covers the other path: Microsoft Foundry Agent Service’s bring-your-own thread storage, where Cosmos DB still does the work, but Foundry owns the orchestration layer on top of it.


Sources

Finding the Right Memory: Vector, Full-Text, and Hybrid Search in Cosmos DB

Post 3 of 6 on Cosmos DB agent memory search, because storing memory well doesn’t guarantee you’ll retrieve the right piece.

Post 2 ended with the schema settled and retrieval still open. This post closes that gap: the practical mechanics of Cosmos DB agent memory search, one container, four query patterns. Run the same question four different ways against that container, and it comes back with four different answers, because “find the right memory” isn’t one query pattern; it’s at least three, and knowing which one to reach for is most of the job.

Vector Indexing for Cosmos DB Agent Memory Search

Cosmos DB supports two vector index types, and the right one depends almost entirely on how many vectors you’re searching, not on anything specific to agents.

quantizedFlat compresses each vector and scans the compressed space exactly. It suits smaller workloads (tens of thousands of vectors) and trades a small amount of accuracy for lower RU cost and faster scans. For a single tenant’s short-term memory, this is often enough on its own.

DiskANN, on the other hand, indexes vectors for approximate nearest-neighbor search and scales to hundreds of thousands or billions of embeddings, with dynamic updates and strong recall even at that size. Post 1 already leaned on DiskANN as part of the case for Cosmos DB as a unified store; this is the mechanism behind that claim.

Sharding the Vector Index for Multitenant Isolation

DiskANN doesn’t have to search across every vector in the container. A vectorIndexShardKey partitions the index itself by a property you choose: session, user, or tenant, so a query only searches candidates within that shard instead of the whole container.

That maps directly onto the partition key work from post 2: set the vectorIndexShardKey to tenantId, or to the same [tenantId, threadId] pair you already use as the partition key, and semantic search for one tenant never touches another tenant’s vectors. A global, unsharded index still works and makes searching everything at once simpler, but it’sonly appropriate for a single-tenant app or a genuinely shared knowledge base where cross-tenant recall is the point rather than a leak.

Full-Text Search: When Precision Beats Semantics

Vector search finds what’s semantically similar. Sometimes semantically similar isn’t what you want — a customer asking about “the refund policy” needs the actual refund policy language, not five conceptually related passages about returns in general.

Full-text search on Cosmos DB handles that case through BM25, a statistical ranking function that scores by term frequency and document length. Cosmos DB applies linguistic processing automatically: tokenization, stemming, case normalization, so “running” still matches “run” or “ran.” It’s the right tool whenever exact terms or phrases carry meaning that a vector embedding would blur.

Hybrid Search: Combining Both with RRF

Most agent memory queries don’t need to choose between semantic and lexical relevance; they need a blend of both. That’s what Reciprocal Rank Fusion (RRF) does: it takes the vector-similarity ranking and the BM25 ranking for the same result set and merges them into one combined rank, instead of forcing a pick between the two.

In practice, this shows up as a single ORDER BY RANK RRF(...) clause, which the next section demonstrates directly.

Four Ways to Ask the Same Question

Take the turn-based schema from post 2 — tenantId, threadId, turnIndex, messages, embedding, content and run the same underlying question against it four ways. (content is a flat, denormalized copy of the turn’s text, added specifically because Cosmos DB doesn’t support wildcard array paths like /messages/*/content in a full-text policy or index the full-text and hybrid queries below point at c.content rather than c.messages for exactly that reason.)

Most recent, by recency:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
ORDER BY c.turnIndex DESC

Semantic, by vector similarity:

SELECT TOP 5 c.messages, VectorDistance(c.embedding, @queryVector) AS score
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
ORDER BY VectorDistance(c.embedding, @queryVector)

Hybrid, blending both with RRF:

SELECT TOP 5 c.messages, VectorDistance(c.embedding, @queryVector) AS score
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
ORDER BY VectorDistance(c.embedding, @queryVector)

Keyword, by exact phrase:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
AND FULLTEXTCONTAINS(c.content, @phrase)
ORDER BY c.turnIndex DESC

Run all four against a thread where a customer asked about refunds three times, in different words, across twenty turns, and the differences stop being theoretical fast: recency surfaces whichever turn happened most recently, even if it’s off-topic; semantic search pulls in every conceptually related turn, including the ones that used different words entirely; hybrid balances the two; keyword search returns only the turns that used the customer’s actual phrase, and ranks them by recency underneath that filter.

Running These Queries in Data Explorer

The four queries above use parameterized SQL, the same form search.py, from the companion repo behind this series, sends through the Python SDK, which binds @tenantId, @queryVector, and @phrase properly before the query runs. Paste them as-is into the Azure Portal’s Data Explorer query pane instead, and two things break, neither of which is a schema or code bug:

Data Explorer’s query box doesn’t bind named parameters. A query that leaves @tenantId unresolved either matches nothing and returns “No results” silently, or for VectorDistance() inside ORDER BY and FullTextScore() fails to compile outright, because both functions require their arguments to resolve to literal values at query-compile time rather than at execution time.

Swap every @parameter for a literal value and all four run cleanly. Against the seeded sample data (tenantId = "contoso", threadId = "thread-1234", searching for "refund"):

Recency, with literals:

SELECT TOP 5 c.messages, c.turnIndex
FROM c WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
ORDER BY c.turnIndex DESC

Semantic, with literals:

SELECT TOP 5 c.messages, VectorDistance(c.embedding, [0.8196, 0.6392, -0.2471, 0.1608, -0.8667, 0.4902, -0.2549, 0.2235]) AS score
FROM c
WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
ORDER BY VectorDistance(c.embedding, [0.8196, 0.6392, -0.2471, 0.1608, -0.8667, 0.4902, -0.2549, 0.2235])

Hybrid, with literals:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
ORDER BY RANK RRF(
VectorDistance(c.embedding, [0.8196, 0.6392, -0.2471, 0.1608, -0.8667, 0.4902, -0.2549, 0.2235]),
FullTextScore(c.content, "refund")
)

Keyword, with literals:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
AND FULLTEXTCONTAINS(c.content, "refund")
ORDER BY c.turnIndex DESC

Pitfalls

Reaching for DiskANN on a small dataset. DiskANN’s approximate search and sharding options solve a scale problem. Below roughly ten thousand vectors, quantizedFlat gets equivalent recall for less operational complexity and lower RU cost. Default to DiskANN because it sounds like the “serious” choice, and you’ve added index-shard decisions to a workload that never needed them.

A global vector index in a multitenant app. Skip the vectorIndexShardKey, and a semantic query searches every candidate in the entire container, tenant boundaries or not. Nothing stops the query from surfacing another tenant’s conceptually similar memory in the result set unless a WHERE clause happens to filter it back out after the fact, and relying on a filter to catch what the index itself should have scoped is the kind of gap that shows up in an audit, not in testing.

Forgetting WHERE filters still apply. Vector and hybrid queries look like they replace normal filtering, but ORDER BY VectorDistance(...) or ORDER BY RANK RRF(...) still runs inside a WHERE-scoped query, same as any other. Leave the WHERE c.tenantId = @tenantId AND c.threadId = @threadId clause off a semantic query, and it searches everything the container holds, not just the thread the agent is currently in.

Next: Coordinating Multiple Agents

That settles Cosmos DB agent memory search for a single agent working alone. Coordinating what several agents know about the same conversation is a different problem, and it’s where change feed, a mechanism post 1 already covered as a callback to the 2023 retail monitoring work, comes back to tie multi-agent state together. That’s post 4.


Sources

The Microsoft AI Stack in 2026 and the Certification Trail That Runs Through It

I spent some time observing what’s inside the Microsoft AI stack. Foundry, Agent Framework, Logic Apps, AI Search, Purview, Entra. After a while, I had a picture of how the pieces fit and decided to draw one myself to share.

Then recently Microsoft published AI-500, an expert certification for multi-agent systems. That made me curious whether Microsoft’s view of the platform matches my own. So I plotted the certification trail against my diagram. This post shows the result.

The Microsoft AI stack as I see it

Six layers. Five of them stack vertically. The sixth runs down the side, through all the others.

Models sit at the bottom: GPT-5, Claude, Mistral, Grok, Microsoft’s own MAI and Phi, Llama, DeepSeek, and the open catalog. This is the least differentiated layer. In Foundry, swapping one model for another is a configuration change. It gets the most attention and deserves the least.

Infrastructure comes next. Foundry is the hub, alongside Azure OpenAI, Azure ML, AKS, Container Apps, App Service, and Foundry Local for the edge. This is hosting, serving, and compute. Solid and well understood.

Data and context: the real moat

I split the data layer in two, because it hides the most important part of the platform. Layer 3a holds the sources of truth: Microsoft Graph, SharePoint, Exchange, Fabric and OneLake, Dataverse, and the vector stores in Cosmos DB, Azure SQL, and PostgreSQL.

Layer 3b is context. Work IQ, Fabric IQ, Foundry IQ, and Azure AI Search turn enterprise data into grounding that respects who is asking. Permission-trimmed retrieval must honor Entra ACLs at query time, not filter after ranking. AI Search supports this through document-level access control. Get it wrong and answers leak, or recall collapses.

This is where the hard engineering hours go. Swapping a model is a config change. Graph-grounded context is months of work.

Agents, Copilots, and the layer vendors leave out

The agentic platform sits above the data. Microsoft Agent Framework merged Semantic Kernel and AutoGen. Next to it sit Foundry Agent Service, Copilot Studio, Logic Apps, Azure Functions, Service Bus, and Event Grid. MCP and A2A handle tools and agent-to-agent communication.

The Copilot layer is on top: Microsoft 365, GitHub, Security, Dynamics 365, Power Platform, and Teams. This is distribution. These are the surfaces people already live in. As a result, enterprise AI adoption is easier when data, permissions, infrastructure, and applications already exist in one ecosystem.

Then comes the sixth layer of the Microsoft AI stack, the one vendor diagrams leave out: governance and evidence. Entra ID and Entra Agent ID handle identity. Purview covers labels, DLP, and audit. Content Safety and API Management provide guardrails and the AI gateway. Defender, Sentinel, Azure Monitor, Log Analytics, and Azure Policy deliver traces, evaluations, and control evidence.

I work for a regulated organization. Before anything goes live, Legal and Internal Audit ask four questions. Who approved the agent? What data did it access? Which controls applied? What happened when it made a wrong decision? Distribution makes deployment easier. However, without traceable evidence it does not make the system production ready. That is why this layer runs vertically through my drawing.

The Microsoft AI certification trail

I knew AI-900. I had not followed what replaced it. This week I learned that Microsoft now offers a full set of AI certifications, from fundamentals to an expert exam. The expert exam, AI-500, is in beta. I read the study guide. Its scope says a lot: orchestration patterns, agent-to-agent protocols, observability, guardrails, and cost control. Architecture work, end to end.

What made it click was laying the whole trail side by side. Each step has its own verb.

  1. AI-901, Azure AI Fundamentals. Understand the concepts and services.
  2. AI-103, AI Apps and Agents Developer. Build applications and agentic solutions on Foundry.
  3. AI-200, Azure AI Cloud Developer. Engineer cloud-native AI properly.
  4. AI-300, ML Operations Engineer. Operate models in production.
  5. GH-300 and GH-600, Copilot and Agentic AI Developer. Ship with agents working beside you.
  6. AI-500, Multi-Agent AI Solutions Expert. Orchestrate systems of agents at scale.

Plotting the trail on the Microsoft AI stack

Six exams, six layers of the Microsoft AI stack. I mapped each exam to its layers, in its study guide exercises, and its center of gravity.

The fundamentals exam touches every layer at concept depth. AI-103 lives in models, infrastructure, context, and the agentic platform. AI-200 moves down into infrastructure and data. AI-300 sits on infrastructure and the evidence plane, because operating models in production is mostly monitoring and evaluation. Meanwhile, the GitHub exams live at the top, where developers meet agents in the editor.

AI-500 is the interesting one. It spans context, the agentic platform, and governance. It does not test models at all. In fact, three of its five headline topics belong to the plane most stack diagrams leave out.

That was the moment the two pictures agreed. My diagram says the hard part of the Microsoft AI stack is context and evidence, not model choice. Microsoft’s expert exam tests context and evidence, not model choice. I did not expect a certification roadmap to confirm an architecture opinion, but here we are.

Where this is the wrong answer

Do not read the trail as a ladder you must climb in order. If you already run agents in production, AI-500 reflects your work, and AI-901 will teach you nothing. Platform engineers should look at AI-200 and AI-300 first. Developers should start with the GitHub exams.

Also, do not read my layer mapping as Microsoft’s. It is my reading of the study guides, and beta study guides move.

Finally, do not confuse the certificate with the evidence. Passing AI-500 shows you know what a control looks like. It does not produce the audit trail for your agent. That still takes Purview configured, Entra Agent ID issued, traces flowing to Log Analytics, and someone signing off. The exam is a map of the work. The work is still the work!

Credits

The Azure icons come from the official Azure architecture icon set. Product marks belong to their owners. Microsoft’s announcements are on the Skills Hub blog: Multi-Agent AI Solutions Expert, AI Apps and Agents Developer Associate, and GitHub Agentic AI Developer. The stack diagram builds on a five-layer picture that circulated on LinkedIn; the split data layer and the governance plane are my additions.