Six Real-Time Patterns in Azure: What the Diagram Leaves Out

I see a diagram of real-time communication patterns pop up in my feed every few weeks. Polling, long polling, Server-Sent Events, WebSockets, webhooks, gRPC streaming. Six neat boxes, and one piece of sensible advice underneath: choose the simplest model that fits the use case.

The diagram is correct. However, it is not the part that costs you a sprint.

Picking a pattern takes about five minutes. Making that pattern survive API Management, Azure Front Door, a scale-out event, and a rolling deployment takes considerably longer. This post covers that second part. It also comes with a working sample that runs all six patterns from one image, against one shared event source. Azure needs two apps to host it, and that’s one of the more interesting findings below.

Four questions, not six boxes

A grid of six options invites you to shop. Instead, answer four questions in order. The pattern then picks itself.

Which direction does data flow? The client pulls, the server pushes, or both sides talk at once. This single question removes at least three options.

What is the latency budget? Seconds are cheap. Milliseconds are not. Most business dashboards tolerate a delay that their designers never measured.

Who owns the connection when it drops? Someone must handle reconnect, resume, and replay. If your answer is “the browser does that automatically”, read the reconnect semantics again first.

Who pays per open socket? Idle connections still consume replicas, units, and money.

When two patterns both fit, pick the one that keeps state out of the connection. A dropped request costs a retry. A dropped session costs a reconnect, a replay, and a support ticket.

The six patterns, and what bites in Azure

Polling

The client asks again on its own schedule. Azure itself uses this pattern constantly. The asynchronous request-reply pattern returns 202 Accepted with a Location header, and Durable Functions exposes a status query endpoint that works exactly this way.

What bites: the missing Retry-After header. Without it, every client invents its own interval. Ten thousand clients then converge on the same second after a deployment, and your scale rule reacts to a spike you created yourself.

Long polling

The server holds the request open until data arrives or the clock runs out. Azure Service Bus applies the same idea inside its SDK, where a receive call waits for a configurable maximum wait time instead of returning empty.

What bites: timeouts you do not control. App Service and Azure Functions enforce a fixed 230-second ceiling on HTTP requests. That number comes from the Azure Load Balancer underneath, which idles connections out at 240 seconds by default. You cannot raise it because it is a platform constraint, not an application setting. Front Door then applies its own origin response timeout on top. Therefore, derive your hold time from the shortest timeout in the path, not from the longest. The sample holds for 25 seconds, which clears every layer.

Server-Sent Events

The server pushes over one long-lived HTTP response. SSE is unfashionable and quietly excellent. It runs over plain HTTP, it survives proxies that understand chunked responses, and browsers reconnect on their own.

You already depend on it. Token streaming from Azure OpenAI and Microsoft Foundry arrives as Server-Sent Events. Every chat interface you have built this year uses this pattern, whether or not the architecture diagram says so.

What bites: buffering, which gets its own section below. Also reconnect gaps. The browser resends the Last-Event-ID header automatically, but the server must honor it. Otherwise, every reconnect silently drops the events that arrived while the socket was down.

WebSockets

Both sides talk over one persistent connection. This is the right answer for chat, collaborative editing, and live trading. It is the wrong answer for a dashboard that changes twice an hour.

Azure gives you two managed options. That is Azure Web PubSub, which handles raw WebSocket clients and works well outside .NET. And Azure SignalR Service fits when you already use hubs and want fallbacks.

What bites: the deployment. A rolling revision in Azure Container Apps drops every open socket at once. All those clients then reconnect together, which looks exactly like an attack to your scale rules. Managed services exist mainly to move that problem off your replicas.

Webhooks

One system calls another when an event happens. Azure Event Grid delivers this way, with retries, dead-lettering, and support for the CloudEvents schema.

What bites: the unglamorous eighty percent, and it starts before your first event arrives.

First, you must pass a validation handshake, and the shape depends on your schema. The native Event Grid schema sends a POST carrying a SubscriptionValidationEvent. You read validationCode from the data object and echo it back as {"validationResponse": "<code>"} with a 200, within 30 seconds. As a result, an endpoint that authenticates, queues, and logs before responding can miss that window on a cold start. The CloudEvents schema works differently. It sends an HTTP OPTIONS preflight carrying a WebHook-Request-Origin header, which you echo back as WebHook-Allowed-Origin. Afterward, every real delivery carries a matching Origin header that you can check.

Second, verify signatures over the raw body, because serialization isn’t byte-stable.

Third, deduplicate. At-least-once delivery makes duplicates a certainty rather than an edge case.

Fourth, answer 400 rather than 500 when a payload will not parse. Event Grid retries 5xx responses because a server error suggests a temporary problem. A body your parser cannot read is not temporary. So an unhandled exception turns one bad payload into a retry cycle that runs until the event dead-letters, and no attempt in that cycle can ever succeed. I introduced this exact bug in the sample while writing this post, then watched a shell-quoting mistake surface it.

Ordering bites too. Event Grid delivers at least once and never guarantees sequence. A failed event enters an exponential backoff queue, while later events sail straight past it. Retries therefore scramble order actively rather than occasionally. So build consumers around sequence numbers or timestamps in the payload. Alternatively, when order is structural rather than incidental, pull from Event Hubs or a Service Bus session instead.

gRPC streaming

Service-to-service, strongly typed, over HTTP/2. ASP.NET Core supports server, client, and bidirectional streaming out of the box.

What bites: the protocol has to match on both sides of the ingress, and the two halves are configured in different places.

Start with ingress. The transport property governs the protocol between the ingress proxy and your container, not only at the edge. Container Apps needs transport: http2 for gRPC. However, the WebSocket handshake relies on the HTTP/1.1 Upgrade mechanism, which HTTP/2 replaces with stream multiplexing. So the two protocols sit badly behind one ingress.

Now the container. Kestrel needs an explicit Protocols: Http2 setting to speak cleartext HTTP/2. The Http1AndHttp2 value looks like it covers both, yet it quietly falls back to HTTP/1.1 without TLS, because negotiation depends on ALPN. Container Apps terminates TLS at the ingress, so your container never sees a handshake to negotiate over.

Kestrel says so during startup, in plain language:

HTTP/2 is not enabled for [::]:8080. The endpoint is configured to use
HTTP/1.1 and HTTP/2, but TLS is not enabled. HTTP/2 requires TLS application
protocol negotiation. Connections to this endpoint will use HTTP/1.1.

That line sits in your container logs from the first boot. Nobody reads startup logs when the application starts successfully, so it waits there until a gRPC call fails hours later.

Get one half right, and the other wrong, and the client receives this instead:

upstream connect error or disconnect/reset before headers.
retried and the latest reset reason: remote refused stream reset

That message names the upstream connection, which sends you to inspect ingress, networking, and scale rules. Meanwhile, the container is the one answering in the wrong protocol. So read the container startup logs first when a gRPC call fails at the edge. The answer usually arrives before the question.

The sample settles both halves at once. It listens on port 8080 for HTTP, SSE, and WebSockets over HTTP/1.1, and on port 8081 for gRPC over h2c. It then deploys the same image twice, once with transport: auto pointed at 8080, and once with transport: http2 pointed at 8081. One image, two ports, two ingress configurations, no compromise.

What the edge does to your stream

Here is the failure I keep seeing, and it never appears on the pattern diagram.

You build an SSE endpoint. It works perfectly on your laptop. You then publish it through API Management, and every event stops arriving. Nothing appears for two minutes, and then the whole stream lands at once, or the connection simply times out.

Your code is fine. The gateway buffers the response by default. It collects chunks from the backend, typically in 8 KB increments, and forwards them only when the buffer fills, or the stream ends. That is reasonable behavior for a normal API and fatal for a stream.

The fix is one attribute in the backend policy:

<backend>
<forward-request buffer-response="false" timeout="240" />
</backend>

Set buffer-response to false, and the gateway forwards each chunk as it arrives. Note the hyphen. An underscore looks close enough to survive a code review and fails policy validation.

Meanwhile, three more layers deserve a look before you ship:

  • Kestrel buffers writes unless you call DisableBuffering and flush after each event.
  • Front Door applies an origin response timeout. Standard and Premium default to 60 seconds, while the Classic SKU defaults to 30. You can raise it, but only to 240 seconds. That ceiling matters more than the default. No SSE stream published through Front Door survives past four minutes, so build the reconnect path and honor Last-Event-ID from the start. Otherwise, a silent disconnect looks exactly like a bug in your code.
  • App Service and Functions cap HTTP requests at 230 seconds. The Azure Functions hosting documentation states it plainly, and no setting overrides it. Microsoft’s own recommendation is the asynchronous pattern: return 202 Accepted and let the client poll. In other words, the platform pushes long streams back toward pattern number one.

Notice what those numbers have in common. App Service stops at 230 seconds. Front Door caps at 240. API Management accepts a higher timeout on forward-request, yet values above 240 are discouraged, because the network drops idle connections at that same boundary. All three inherit one limit from the load balancer underneath.

As a result, four minutes is the practical ceiling for a single connection in Azure. No pattern on the diagram changes that. Plan the reconnect instead.

In short, the pattern lives in your code, but the behavior lives in your platform configuration. Test through the full path, not against localhost.

Cost and scale change the answer

Persistent connections change how you scale. A stateless request occupies a replica for milliseconds. A WebSocket occupies one for hours.

Container Apps scales on concurrent requests, and a long-lived connection counts as one. Therefore, set that threshold low. Otherwise, your replicas fill with idle sockets long before CPU tells you anything is wrong.

Managed services price per unit and per connection, so check the current pricing pages before you compare. Still, the service fee is rarely the real cost. Backplane configuration, connection affinity, graceful drain during scale-in, and reconnect storms after each deployment cost far more engineering time than the invoice suggests.

Where this is the wrong answer

Every pattern has a place where it becomes a liability. Here are the ones I argue about most.

WebSockets for a dashboard that changes hourly. You pay for a connection per user to deliver one update, and you now maintain reconnect logic forever. Nobody has ever thanked an architect for putting a WebSocket on a management report.

SSE for mobile clients on unreliable networks. Reconnect works, yet you still need Last-Event-ID and server-side replay to avoid gaps. When the app spends half its day in the background, use push notifications instead.

Webhooks when ordering matters. Event Grid retries, so your receiver sees events twice and out of sequence. If order affects correctness, pull from a log or a session instead.

gRPC streaming to a browser. You need grpc-web plus a translating proxy. As a result, you added infrastructure to solve a problem that SSE already solved.

Long polling in new code. It exists for compatibility with systems you cannot change. For greenfield work, SSE is simply better.

A private socket server to avoid managed service pricing. The service fee is not what hurts. Sticky sessions, drain behavior, and reconnect storms are what hurt.

Real-time at all. If a user reads the number once per hour, batch it. Real-time is a requirement, not a compliment.

Try it yourself

The companion repository runs all six patterns from one image, against one price tick produced every second.

Because the data is identical everywhere, the six-pane test page shows exactly what differs. You watch the request counter climb under polling, the lag column drop under SSE, and the WebSocket pane accept a filter sent back up the same connection. You can also put API Management in front and reproduce the buffering failure in about ten minutes.

Deploy it with one script:

Bash, macOS or Linux:
./infra/deploy.sh
PowerShell on Windows:
./infra/deploy.ps1

Choose the boring one

The original advice still holds. Choose the simplest pattern that meets your latency, reliability, and scalability requirements.

I would add one line to it. Choose the simplest pattern that survives your gateway, your scale rules, and your next deployment. That constraint eliminates more options than the latency budget ever will.

Cosmos DB Agent Kit: Auditing My Own Sample Repo

A bonus post, series-adjacent, running the Cosmos DB Agent Kit against cosmos-agent-memory-lab to see what six posts of hand-checked rules missed.

Every rule in this series hierarchical partition keys, TTL modes, vector index types, the composite ID trick came from reading Microsoft Learn docs and debugging real errors one at a time. The Cosmos DB Agent Kit packages that same category of knowledge differently: 100+ best-practice rules across 12 categories that AI coding agents apply while writing or reviewing Cosmos DB code. This post checks the rules against the actual sample repo behind posts 2 through 4, rule by rule, instead of taking the kit’s word for it.

What the Cosmos DB Agent Kit Actually Ships

It’s not a linter or a static-analysis tool. It’s a set of Markdown rule files, one per practice, each with an incorrect example, a corrected example, and a short explanation of why, following the Agent Skills format, so tools like Claude Code, GitHub Copilot, and Gemini CLI can load them and apply them while generating or reviewing code. npx skills add AzureCosmosDB/cosmosdb-agent-kit installs it. The categories track the same ground this series covered: data modeling and partition keys rank CRITICAL, queries and SDK usage rank HIGH, vector search and full-text search get their own dedicated categories.

Nine Things the Repo Already Gets Right

Cross-checking models.py, schema.py, and checkpointer.py against the kit’s rule files turned up a solid list of matches. The hierarchical partition key orders tenantId before threadId, broad to narrow, exactly as partition-hierarchical recommends. The composite turn ID uses: as a separator — the kit’s model-id-constraints rule calls out #, ?, /, and \ as characters that break Cosmos DB’s REST auth signing, and: sits on its explicit safe list. TTL follows the kit’s own “correct” pattern precisely: container default -1, item-level ttl overriding it per turn. The vector embedding policy and DiskANN index both match vector-embedding-policy and vector-index-type field for field, and the indexing policy excludes the embedding path from the regular range index to avoid double-indexing cost.

The full-text policy uses en-US, case-sensitive, as the kit’s fts-define-policy rule insists, and the content field sits in fullTextIndexes without also cluttering excludedPaths incorrectly. Turns live in their own container, separate from checkpoints, precisely the pattern-langgraph-chat-history-separate rule, which exists because checkpoint blobs make poor chat history. And checkpointer.py reaches for a point read wherever it already has both the id and the partition key, rather than running a query that costs roughly 2.5x more.

Three Things It Would Flag

Not everything cleared. seed.py never normalizes its mock embeddings to unit length vector-normalize-embeddings flags exactly this, and it’s a genuine miss, not a style preference: unnormalized vectors produce inconsistent cosine-similarity scores, and the fix is one line (v / ||v||₂) that never made it into the SHA256-based fake_embedding() function. Second, the checkpointer is a hand-rolled BaseCheckpointSaver implementation against the synchronous azure-cosmos SDK, not the official async CosmosDBSaver from langchain-azure-cosmosdb that sdk-langchain-cosmosdb-saver recommends. That one’s a deliberate tradeoff rather than an oversight — building it by hand is what made post 4’s checkpointing section possible to explain from the inside. Still, a production app should almost certainly reach for the maintained package instead. Third, turn upserts carry no ETag check so that sdk-etag-concurrency would flag the read-modify-write path as vulnerable to lost updates under concurrent writes. The demo’s single-writer pattern never triggers it, but two agents racing to update the same turn would silently drop one of them.

Where a Rulebook Like This Actually Helps

A tool that catches “your mock embeddings aren’t normalized” in seconds, instead of after a confusing test failure, earns its place in a workflow. What it doesn’t replace is the reason each rule exists — Cosmos DB Agent Kit tells you: is a safe ID separator, but this series spent a paragraph on why # breaks HMAC signing specifically in Gateway mode. Both matter: the rulebook for speed, the explanation for judgment calls like the checkpointer tradeoff above, which no automated check can make for you. Six posts of hand-checked Cosmos DB rules and a 100-rule kit landed on the same answers almost everywhere — that convergence is worth more than either source alone.


Sources

Rate Limiting Algorithms on Azure: 6 Compared

Your API works perfectly, until someone hammers it with 10,000 requests in a second. Rate limiting stands between a stable system and an outage, and the usual list of rate limiting algorithms has six entries: fixed window counter, sliding window log, sliding window counter, token bucket, leaky bucket and concurrency limiter. However, few guides show what each one actually lets through, or where it should run. So I built all six in C# on Azure, load-tested them, and wrote down what I learned. The code is in github.com/steefjan1/rate-limiting-azure.

Four of the six are already in .NET

In fact, on .NET you mostly do not implement these. System.Threading.RateLimiting ships the fixed window, sliding window, token bucket and concurrency limiter; ASP.NET Core wires them in with AddRateLimiter and RequireRateLimiting. The whole fixed window:

options.AddPolicy("fixed-window", ctx =>
RateLimitPartition.GetFixedWindowLimiter(PartitionKey(ctx), _ =>
new FixedWindowRateLimiterOptions { PermitLimit = 10, Window = TimeSpan.FromSeconds(1) }));

The missing two, sliding window log and leaky bucket, are small RateLimiter subclasses in the repo. The log keeps a queue of timestamps per client. The leaky bucket, by contrast, keeps one number, the next free drain slot, and either holds the request until then or rejects it when the queue is full, making it the only limiter that adds latency on purpose.

One nuance: .NET’s sliding window is segmented, not the weighted “current plus previous window” version most explanations describe; the repo has the weighted one in Redis.

What the load test shows

Every endpoint gets the same budget: 10 requests per second per client. What matters most is the highest number of requests the backend saw inside any one-second span.

Scenario 1: a burst across the window boundary

First, a burst across a window boundary: one request opens the window, nine more arrive at 900 ms, ten more at 1100 ms.

AlgorithmMax accepted in any 1 sRejected of 20
fixed window190
sliding window log109
sliding window counter1010
token bucket181
leaky bucket110 (the last ones waited 999 ms)

The fixed window result is the textbook flaw, measured: a “10 per second” limit let 19 through in 200 ms. The log and the segmented counter, by contrast, hold at 10.

The token bucket result surprised me. A full bucket plus a couple of refilled tokens should give 12 or 13, not 18. The cause is in .NET: when the bucket is full, TokenBucketRateLimiter.ReplenishInternal returns early without moving its last-replenishment timestamp, and the PartitionedRateLimiter that hosts it refills on elapsed time. So after a quiet spell, the first refill credits all the time the bucket spent full, roughly doubling your burst allowance. The Redis token bucket in the repo, however, updates the timestamp on every call and does not do this.

Scenario 2: sustained overload

Next, sustained overload: 30 requests per second for three seconds. Fixed window and sliding log each accept 30 of 90; the segmented counter accepts 26. The token bucket accepts 39 by design, since capacity sets the burst and refill rate sets the average. Meanwhile, the leaky bucket accepts 40 with a p95 latency of 1002 ms, because it queued ten and drained them at the fixed rate.

Scenario 3: concurrency at a slow backend

Finally, 20 simultaneous calls at a backend that takes 250 ms. Without a limiter, all 20 hit it at once; with the concurrency limiter at 4, four get through and sixteen get a 429 within 2 ms. In other words, a service handling 10,000 cheap requests a minute can still fall over when 500 slow ones run at once.

The decisions that come before choosing a rate limiting algorithm

Which layer

“Combine two or three at different layers” is advice everyone gives and nobody draws. Here is mine.

On Azure, the gateway is API Management, where two of the six already live as policy XML: rate-limit-by-key counts requests per key, and limit-concurrency caps in-flight calls to your backend. Through the deployed gateway, the load test gave the same numbers for 2 to 5 ms of added latency, and APIM rejected nothing, since its policy counts only successful responses. Note, though, that rate-limit-by-key is a sliding window on the classic tiers and a token bucket on the v2 tiers: same XML, different burst behaviour.

Which unit

Ten cheap reads and ten model calls put different pressure on a system, so one threshold everywhere is the wrong shape. Instead, limit against the bottleneck: in-flight calls for a slow dependency, tokens for an AI backend, connections for a database. For example, APIM’s llm-token-limit is the token bucket with tokens as currency, and the RateLimiter API takes a permitCount, so an expensive endpoint can cost five permits while a cheap one costs one.

Where the counter lives

Every in-process limiter counts per instance. Because the sample deploys the API to Container Apps with two replicas on purpose, the same load test shows what that means: the fixed window and the sliding log each accepted 60 of 90 instead of 30, and the concurrency limiter let 8 calls reach the slow backend instead of 4. Every configured number doubled, silently, as one Log Analytics query over the API’s rejection log makes visible:

Each in-process policy rejected about 30 of 90, split across two replicas, while each Redis-backed policy rejected about 60, from one counter. APIM has the same shape: counters are per gateway node, never aggregated. Since the only way to get one counter per client across replicas is shared state, the repo runs each algorithm as one atomic Lua script in Azure Managed Redis: INCR plus PEXPIRE for the fixed window, a sorted set for the log, two weighted keys for the counter, a hash for the token bucket. Against the same two replicas, the Redis versions accepted 30, 30, 29 and 39. It costs one round trip per request, including rejects, but added no measurable latency in the same region.

Shared state, however, brings two failure modes of its own. First, clocks: the scripts take “now” from the calling replica, so drifting clocks disagree about the window; if that matters, read TIME inside the script and let Redis be the clock. Second, network: the limiter fails open when Redis is unreachable, because failing closed would be a worse outage than the one it prevents. That is a choice, and it needs an alert. Also, Azure Cache for Redis closes to new creations on 1 October 2026, so the sample uses Azure Managed Redis, clustered by default with hash tags to keep a client’s keys in one slot.

Who is being limited

Per IP punishes everyone behind a corporate NAT, and per subscription is what APIM gives. Per-user needs a validated token, so the limiter sits after authentication, and unauthenticated floods reach your identity provider. So real systems combine more than one key: IP at the edge, subscription at the gateway, user in the service. Fairness between tenants is policy, not algorithm. The sample uses an X-Client-Id header that APIM sets from the subscription ID, so the gateway and service partition on the same identity.

What the client gets

The client needs a 429 with Retry-After, or it retries at once and your limiter becomes a load generator. Every limiter computes Retry-After from its own state: the window’s remaining time, the log’s oldest entry, the leaky bucket’s next free slot.

The other half is on the caller: honour the header, add jitter, and tune retries with the limits they will hit. Microsoft.Extensions.Http.Resilience does the first two out of the box.

What you watch

Rejections and saturation are different signals. A rising 429 count per policy says a client is over budget, while a concurrency limiter pinned at its permit limit, or a leaky bucket queue that never drains, says the system is at capacity. So the sample tags every rejection with X-RateLimit-Policy and ships APIM gateway logs to Log Analytics. The repo’s docs/kql.md has the queries, starting with the one that says whether the gateway or the service produced a 429 (BackendResponseCode empty versus 429).

Where each one is the wrong answer

A fixed window is wrong when the thing you protect cannot survive 2x for a moment. A sliding log, however, is wrong at high limits: 10,000 per minute means 10,000 timestamps per client. A token bucket is wrong when the downstream needs a smooth rate rather than an average rate. A leaky bucket is wrong at an HTTP edge where clients time out before they drain; it belongs in front of a fragile dependency you own. A concurrency limiter is wrong for fairness between clients, because it says nothing about rate. And every in-process limiter is wrong the moment you have two replicas and still expect the number you configured.

What I would do

Token bucket at the gateway, keyed per client, loose enough that honest bursts pass. Concurrency limiter at the service, in front of the slow thing, sized to what it can take. For more than one replica, use a per-client counter in Redis, and return Retry-After so your own clients respect it. Two rate limiting algorithms are usually enough for a public API on a side project. Then run the load test, because two of the five let through nearly twice what their configuration says.

Foundry Model Ledger: Which Model, at What Price, Until When, and Which of Mine

Every model conversation with a team ends with the same four questions. Is the model available in our region? What can it do? What does it cost per million tokens? And when does Microsoft retire it? Then the platform team asks a fifth: which of our own deployments are affected? Microsoft’s Foundry Model Explorer answers the first question well. However, it does not answer the other four. It cannot, because the answers live in your subscription and in a price list that was never meant to be read by software. So I built the Foundry Model Ledger to read them together. The interesting part is what the data looks like once you do.

What the Foundry Model Explorer already does

Microsoft’s explorer is a region availability reference. It shows one row per model and version, with lifecycle status, deployment SKUs, retirement date, and the list of regions that carry it, exportable to CSV. If your question is “where can I deploy gpt-5.6-terra”, it is the fastest answer there is. However, it is a periodic snapshot rather than a live read. It shows no prices in any currency. And it knows nothing about what you have deployed. Those three gaps are the Foundry Model Ledger.

Three sources, one Foundry Model Ledger

The model catalog lives in Azure Resource Manager. A single call, GET /subscriptions/{id}/providers/Microsoft.CognitiveServices/locations/{region}/models, returns every model and version a region carries. Each entry has its capabilities (chat completion, tool calling, embeddings, context window), its lifecycle status, its deployment SKUs, and its retirement date under deprecation.inference. This is the same data the Foundry portal shows. Moreover, it needs nothing more than Reader on the subscription.

Prices live in the Azure Retail Prices API. It is public, needs no authentication, and returns every meter Azure bills. Filter on serviceName eq 'Foundry Models' and a region, and you get the list prices for tokens, images, and hours, in any currency the API knows.

Your own deployments live under each Cognitive Services or Foundry account. At first I reached for Azure Resource Graph, which is the natural place to list resources across a subscription. It returned zero rows against a subscription with twelve accounts. Resource Graph does not index the accounts/deployments child type. So the Ledger lists the accounts through ARM and then calls each account’s /deployments endpoint, six at a time.

How the Foundry Model Ledger is built

The Foundry Model Ledger is a .NET 8 isolated Azure Function on Flex Consumption with a single-page UI. It deploys with azd up, in the same shape as the [INTERNAL LINK: RAG in 8 Steps on Azure] sample. A user-assigned managed identity with Reader on the subscription reads the first and third source. The second source needs no identity at all.

Pick a region and a currency, and the table shows model, version, lifecycle, capabilities, and retirement date with a days-left bar. Next to those sit the input and output prices per million tokens. Click a row for every SKU with its capacity range, every capability ARM reports, and every price meter the tool matched. The same panel has a button that checks every other region for the same model and version.

Finally, a second tab joins your deployments to the catalog of their own region and sorts them by soonest retirement.

What the data says

Sweden Central, on the day I took these screenshots, carried more than 300 catalog entries from 11 publishers. 216 were generally available, 66 in preview, and 16 marked as deprecating. 70 versions retire within 90 days. Furthermore, 30 are already past their retirement date but still listed. That last group matters. The catalog endpoint tells you what the region knows about, not what you can still deploy. For example, gpt-4o-mini has been closed to new deployments for a long time and still appears. The SKU list and the retirement date are the better signals.

The same model and version often appears twice, once for account kind OpenAI and once for AIServices, each with its own SKU list. The Foundry Model Ledger merges those into one row. If you script against the endpoint yourself, expect the duplicates.

The Retail Prices API returned 1,727 meters for Foundry Models in that one region. None of them carries a model identifier.

Prices are written for invoices

A meter name is a billing label, not a key. The same model, gpt-5.6-sol, is spread across meters such as 5.6 sol ShortCo Inp Std Gl 1M Tokens, 5.6 sol LongCo Cd Wr PP DZ 1M Tokens, and 56sol ShCo Cd Wr Fl Gl 1M Tokens. Older meters read gpt 4.1 nano cached Inp glbl Tokens and bill per 1K tokens; newer ones bill per 1M. Similarly, grok-4.6 appears as 4.6 Inp DZ Tokens under the product Azure Grok Models, with no “grok” in the meter name at all.

The abbreviations are their own dialect. Input is Inp, inpt, or in. Output is Outp, opt, outpt, or out. Cached input is Cd. Global, data zone, and regional are Gl, DZ, and regnl. Batch, priority, and flex tiers have their own tokens.

How the matcher works

So the Foundry Model Ledger has a matcher. First, it narrows meters to the model’s publisher. Then it tokenizes both sides the same way and requires every token of the model name (minus the family prefix) to appear in the meter, in order. It also rejects meters that carry a sibling variant such as mini or pro the model does not have. When a meter carries a date token, o3 0416 or chat-latest 08062026, it pins the version. Next, it classifies direction, deployment type, tier, and context length, and normalizes the unit to a price per million tokens. Finally, it picks a headline: global standard, short context, uncached. Every price in the table carries a confidence label, exact, name, or loose. In addition, the detail panel shows every matched meter, so the headline number is never the only evidence.

What the first live run got wrong

On the first live run it priced 204 of the entries in Sweden Central. The misses were instructive. FLUX image models came out at “40,000 per million” because their meters bill per 1K images, and I had treated every 1K unit as tokens. Qwen sits under product Qwen models, which the publisher map did not know. Moreover, most of its meters are fine-tuning meters that must never become a headline price. And the Anthropic models, plus Cohere rerank and parse, are in the catalog with no Foundry Models meter in the region at all. The model exists; the public price does not.

Two tabs that came from using it

The first version stopped at the tabs above. Two more followed within a day, because the first questions people asked were not the ones I had built for.

Where else is this model available

The region check in the detail panel gives a list of names. That is correct and hard to read. So the Foundry Model Ledger now has an availability map. Every Azure region is plotted from the subscription’s own location metadata, on a world outline from Natural Earth. Type a model name, pick a version, and the regions that carry it light up in the brand blue. Regions without Azure AI services show as a dashed ring, and regions that host the service but not the model stay grey. Zoom presets for Europe, North America, Asia Pacific, the Middle East and Africa, and South America keep the labels readable where regions cluster. In Europe alone there are twenty. The data behind the map is one lightweight catalog call per region, cached, so a check across every region takes a few seconds the first time and is instant after that.

What else does the same job for less

The second question is the one that decides budgets. If we use this model for that workload, what else in the region can do the same job, and what would it cost instead? The Alternatives tab answers it with the data the Ledger already has. Pick the model you use or consider. Its capability flags from the catalog, such as chat completion, tool calling, JSON schema output, and image input, become chips, and every chip is required by default. Click one off if you do not need it and the list widens. Every other model in the region that reports all remaining capabilities is listed with its list price, what your monthly workload would cost on it, and the delta against your model, cheapest first. Enter the workload as millions of input and output tokens per month. Filters restrict the list to the same publisher or to generally available models, and retired versions are excluded. Other versions of the same model are listed last, so they do not pose as alternatives.

The comparison is on capabilities, not on quality. A nano model will always look like a 99 percent saving next to a frontier model, and the table does not know whether it would pass your evaluation set. What it does give you is the shortlist and the price gap in one view, which is the part that used to take an afternoon with the pricing page open in three tabs.

Where the Foundry Model Ledger is the wrong answer

Do not budget on it. These are list prices, matched heuristically, in whatever currency you pick. The matcher will be wrong somewhere the day Microsoft renames a meter. It also cannot price provisioned throughput, which is billed per hour per unit rather than per model. Use it to see the shape of a decision, then confirm on the pricing page.

Do not put it on the internet as is. The Function has no authentication, and it exposes your deployment list to anyone with the URL. Keep it internal, or put Easy Auth with Entra ID in front of it. If you already run model endpoints behind API Management, the same gateway can front the Ledger, with the policies from the APIM for AI Workloads series.

And do not read the catalog as an availability promise. Listed does not mean deployable, and a retirement date is a floor, not a schedule. For a plain “where is it available” question, Microsoft’s explorer remains the quicker tool.

What I would build next

The tab I did not plan is the one people ask about: my deployments, joined to the catalog, sorted by retirement. That is a governance question, not a browsing question. It sits next to the routing questions from Don’t Build Around Today’s Model. Build for the AI Control Plane. Consequently, the natural next steps are all on that side. All subscriptions in a tenant instead of one. An alert when a deployment sits on a version retiring within 90 days. A history, so you can see what appeared and what closed in a region since last month, which is a changelog Microsoft does not publish. And a cost delta for moving each deployment to its successor, which the Alternatives tab now does for one model at a time and should do for the whole deployment list.

The Foundry Model Ledger repo is at github.com/steefjan1/foundry-model-ledger. azd up deploys it; scripts/run-local.ps1 runs it against your az login; node tools/mock-server.mjs runs the UI on a sample snapshot without .NET or Azure. If the matcher misprices a model in your region, the detail panel shows why. An issue with that screenshot is the fastest way to get it fixed.

Microsoft Entra ID under the hood

Recently I saw a Microsoft Entra ID architecture poster on LinkedIn, put together by Srawon Kumar Reddy Mula, and it is a good one: six numbered steps across the top, an architecture panel underneath, security and governance columns down the side. User signs in, request goes to Entra ID, authenticate, evaluate access, issue token, access application. A dashed arrow loops back from the last box to the third and carries the label “token renewal / continuous access evaluation”.

(Source: LinkedIn post by Srawon Kumar Reddy Mula)

And I was interested in how that actually works, because the diagram does not really show it. One dashed arrow covers a lot of ground. How long does that token live? What happens to it when you revoke someone? Does the loop do anything at all if the client never asked for it?

So I went and found out. I built the flow against a real tenant, one working sample per box, and measured the parts the poster fits into an icon. Nothing on the poster is wrong, and none of what follows is a correction. It is the answer to a question a picture that size cannot hold.

Two comments under the original post were asking better versions of the same question. Ernie Prescott pointed out that “Hybrid Identity with Active Directory” was sitting quietly in the highlights list. It was doing far more work than it looked like. Sajeed Mullaji wrote that PIM only solves the activation window, not the payload. Both of them named a boundary the picture compresses, so both of them got a sample.

Six samples, one per box

The repository is at github.com/steefjan1/entra-id-end-to-end. Six samples, one per box, every one of them run against a real tenant.

The dashed arrow is doing the most work on the whole poster

Continuous access evaluation gets one dashed line and one label. In practice it is a narrow feature with a wide reputation.

It covers five critical events: the account is deleted or disabled, the password changes, MFA is enabled for the user, an administrator revokes all refresh tokens, or Identity Protection detects high user risk. It reaches Exchange Online, SharePoint Online, Teams and Microsoft Graph, and even that list needs footnotes. The client matrix marks Teams partially supported in every row. SharePoint Online does not support the user risk event. Azure Resource Manager is not on the list at all. And none of it reaches a client that did not ask for it. Asking means declaring the cp1 capability, so that Entra ID treats the session as CAE aware.

Two clients, one line of configuration apart

That last part is a single line of MSAL configuration, and everything hangs off it. So the first sample in the repository measures the difference rather than describing it. It signs the same user in twice against Graph, once with clientCapabilities: ['cp1'] and once without. Then it revokes the sessions and polls with both tokens until each one stops working.

Here is what that returned on my own tenant this morning. Client A, the one declaring cp1, came back with a token good for 1439 minutes. Client B, identical except for that one line, came back with 65 minutes. Both sit exactly where the documentation says they should: 20 to 28 hours for a continuous access evaluation session, and the randomized 60 to 90 minute band for an ordinary one.

Then I revoked the user’s sessions.

Client A stopped working four seconds later. I know it was four seconds rather than roughly, because the claims challenge carries the timestamp: the nbf value Entra ID sent back decodes to 07:57:38Z and Graph turned the token away at 07:57:42.

Client B kept answering 200 for the rest of its 65 minutes. Revoking sessions invalidates refresh tokens and browser cookies. Microsoft’s own guidance on removing a user’s access says the rest: for applications using access tokens, the user loses access when the access token expires. Nothing in the tenant shortens that window.

The challenge is a timestamp, not a capability

The challenge itself is worth a look, because it is not the one most articles show:

{"access_token":{"nbf":{"essential":true,"value":"1787731058"}}}

Not a capability negotiation. A timestamp. The resource is saying: your token predates the revocation instant, bring me one issued after it. essential: true means the client does not get to negotiate.

Which is the whole of that dashed arrow, drawn out. Notice what the client does with the challenge: it clears its cache, asks again with the claims parameter, and gets a fresh token. The user sees none of it. That is the part worth implementing. It is also why declaring cp1 without handling the challenge is worse than not declaring it at all: the client will loop, retrying a token the resource has already told it to replace.

Run that once against your own tenant and the dashed arrow stops looking like a safety net.

“Evaluate access” is not continuous, and that is the part people get wrong

One passage from Microsoft’s own documentation deserves a place on that poster more than anything already on it: policies targeting roles or groups are evaluated only when a token is issued, and if a user already has a valid token before being added to the role or group, the policy does not apply retroactively.

Read that against how offboarding usually works. You remove someone from a group. The group was the thing granting access. You now believe access is gone. It is not, because the token in the client’s cache does not care about your group change, and continuous access evaluation does not cover group membership. Microsoft documents that replication as taking up to one day. There is an optimization that brings it down to two hours. It applies to policy updates rather than to group membership, and the documentation says plainly that it does not cover all scenarios yet. Their own recommended workaround is to revoke the user’s sessions by hand.

The same applies to a new Conditional Access policy and to a role assignment. Which is a strange thing to discover during an incident.

Where authority actually lives

Ernie’s point deserves the space. In a hybrid tenant, the box marked “authenticate” is not where the answer comes from. It comes from wherever the account state currently lives, and the delay between the two directions is not symmetric.

With password hash synchronization, disabling an account in Active Directory does not immediately end cloud access. Microsoft puts the window at up to 30 minutes, and adds that sync never carries password expiry or account lockout state to Entra ID at all. Pass through authentication and federation enforce those states at sign in, immediately. That is a real architectural difference hiding behind a single line item in a highlights list.

The other direction has no delay because it has no mechanism. Neither Connect Sync nor Cloud Sync provisions a user disable back to Active Directory. Disable the cloud account and Kerberos, NTLM and LDAP access on premises continues exactly as before. Conditional Access does not sit in that path natively. Entra Private Access for domain controllers went GA in January 2026. It is the way to put policy in front of Kerberos, and it works by putting sensors on the domain controllers rather than by extending the token flow.

There is a clock on Connect Sync now

There is a clock attached to this now. Microsoft published a phased transition plan from Connect Sync to Cloud Sync. From July 2026 it starts telling tenants their individual transition windows through the Message Center, Connect Health and targeted email. The tenants where Cloud Sync already covers everything go first. Source of authority conversion for individual users went GA in January 2026. The Active Directory to Entra ID boundary is moving under people while they are still drawing it as one arrow.

Sample 05 in the repository does not fix any of this. It measures it. The report lists every principal by where its authority lives, and flags synced objects whose last sync is stale. It also names the case that should worry you most: a principal whose authority is on premises but whose privilege is in the cloud.

It also ships a coverage matrix that contacts nothing and takes a second to read. Six controls down the side, five surfaces across the top, and the column for on-premises Kerberos, NTLM and LDAP reads no from top to bottom. Conditional Access, MFA, continuous access evaluation, sign-in risk, device compliance, sign-in logs. Not one of them reaches that surface natively. That column is the highlights-list bullet, drawn honestly.

PIM controls the window, not the payload

Sajeed’s comment is the best one line summary of privileged access I have read this year, so I built a script around it.

PIM answers when a role is active. It says nothing about what the role can do while it is. An eligible assignment with a one hour activation limit, MFA on activation and an approval step looks like a strong control on a dashboard. If the role behind it grants four hundred actions including application credential management, you have not reduced the blast radius. You have scheduled it.

So sample 04 prints both numbers on the same row. Per assignment it prints the activation maximum duration, and whether the role requires MFA and approval. Next to that sits the number of resource actions the role definition actually grants. Then a payload risk verdict that flags wildcards, and a list of high impact actions.

What the report found on my own tenant

I ran it on my own lab tenant expecting to demonstrate a point. It found something instead. Two service principals, one of them a monitoring integration, both holding Directory Writers. That role carries microsoft.directory/servicePrincipals/appRoleAssignedTo/update, which grants application permissions to any application in the tenant. It also carries microsoft.directory/groups/members/update. Neither assignment expires. Nobody would think to look at a monitoring identity when auditing privilege. No PIM dashboard would show it as a problem either, because PIM is not involved at all.

That tenant has no P2, which turned out to matter in an instructive way. The report falls back to the plain role assignment endpoint and prints the window columns as unavailable. The action counts do not change, because a role definition grants what it grants however the assignment happened. A tenant with no PIM is the more alarming case, not the less: every assignment is standing, permanent, with no window to shorten.

Writing that fallback also caught a bug in my own scoring. Entra ID writes its largest wildcard in words rather than asterisks. Global Administrator’s payload includes microsoft.directory/allEntities/allProperties/allTasks, and matching wildcards on * alone ranked Global Administrator below a billing role on a naive action count. In a script whose entire purpose is measuring the payload, the measurement inverted the ranking at exactly the row that mattered most.

Sajeed made the point about ERP duties, and he is right that Entra ID cannot see inside the application. That is the honest boundary of this work. Sample 04 reports Entra’s half. The duty separation inside your finance system is a separate project with separate tooling. Pretending PIM covers it is how audits get closed without anything getting safer.

Step one says “user signs in”

In a real Azure tenant, most sign-ins are not users. They are service principals, managed identities, and pipelines.

Almost none of the controls in the diagram apply to them the way people assume. Conditional Access for workload identities covers single tenant service principals registered in your tenant. It does not cover managed identities at all. It does not cover multitenant or Microsoft applications. Its conditions cover location, service principal risk and authentication context, and block is the only grant control it offers. If your plan was to require MFA for a pipeline, there is no such thing. Continuous access evaluation for workload identities reaches Microsoft Graph and nothing else.

What you can do is remove the secret. Sample 06 deploys workload identity federation, so a GitHub Actions workflow reaches Azure with no stored credential at all. It also ships an inventory script. That script ranks every workload identity in the tenant by credential type, days to expiry and the application permissions it holds on Graph.

What 348 service principals look like

Run against my lab tenant it found 348 service principals, 246 of them Microsoft first party, leaving 102 worth inspecting. Twenty seven of those are managed identities. That is the number behind the claim at the top of this section. One person, one lab, and a hundred non human principals authenticating, without anybody ever having thought about them as sign-ins.

The report was also wrong twice, and both mistakes are instructive. Graph returns signInAudience as null for a managed identity, and null !== undefined, so my check labelled all twenty seven of them multitenant. And Azure rotates a managed identity’s certificate itself, leaving stale entries on the service principal, so the report announced five certificates that expired two thousand days ago as high risk. Two categories of confident nonsense, both crowding out the one row in that tenant that genuinely mattered: a multitenant application holding Directory.ReadWrite.All and User.ReadWrite.All.

A managed identity has no credential you rotate and no Conditional Access at all. Its entire blast radius is its permissions and its Azure RBAC scope. I made the same argument from the other end when I looked at Logic Apps agent loop security: once the identity is managed for you, the permissions are the only surface left to get wrong. Which is the argument for reading the permissions column rather than the expiry column. I had to be wrong in public on my own tenant to see it.

Conditional Access deserves the same pipeline as your code

The shield in the middle of the poster is a set of policies someone clicked together in a portal. Conditional Access is the tenant wide half of authorization. The per API half I covered separately in securing AI APIs with authentication and authorization in Azure API Management. Neither one substitutes for the other. No history, no review, no test, and no way to answer “what breaks if I enable this” other than enabling it.

Sample 03 treats it as code. Policies are JSON files with placeholders instead of hard coded object IDs. A guard refuses to ship a policy that targets users, carries a grant control and does not exclude the break glass group, because Conditional Access has no “except the person who wrote it” fallback. Everything deploys in report only unless you opt in twice.

The part I would steal even if you ignore the rest is the test runner. Graph exposes the What If evaluation as an API, so every pull request can check a set of hypothetical sign-ins against your live policies. The most valuable assertion in the file is the one that says no policy of your own catches the break glass account. Run it forever.

Where this is the wrong answer

Not all of this is worth doing.

If you run a small tenant with a handful of policies and one administrator, the portal is fine. A Conditional Access pipeline is machinery you will maintain instead of using. Without Entra ID P2, PIM and access reviews are not a decision you get to make, and standing assignments with a tight scope plus an honest quarterly review beat pretending otherwise. If you are already on Cloud Sync and everything works, the 2026 migration is not a project you need to start this quarter.

And the measurement scripts come with a warning I will repeat here. The revocation stopwatch revokes a real person’s sessions. The propagation watcher requires you to disable a real account. Tell the person first.

This is the same conclusion I keep arriving at, and I wrote up how I got here in my Azure security journey. The one thing I would not skip, whatever size you are, is running the two read only reports. The entitlement report in sample 04 and the sync report in sample 05 need no licence beyond what you already have. They change nothing. And they answer questions about your own tenant that the diagram cannot.

The Microsoft Entra ID samples repository

Everything above ships as working code at github.com/steefjan1/entra-id-end-to-end, MIT licensed. Six samples, each with its own README, its own permissions list, and its own section on where it is the wrong answer. Every write supports a dry run. Nothing deletes anything.

I deployed sample 01 and signed into it. Sample 06 compiles clean and I have not deployed it yet, which is exactly the kind of thing a diagram would let me leave out and a README should not.

Five defects that only showed up on deployment

Deploying it for real was worth more than writing it. Five defects turned up that no amount of template validation would have caught. The repository documents every one, with the actual error text:

  • Entra ID validates preAuthorizedApplications against the scopes that already exist on an app, not the ones you create in the same request. Exposing a scope and pre-authorizing a client for it has to be two calls with a wait in between.
  • az writes to stderr for entirely ordinary things. Windows PowerShell under ErrorActionPreference = 'Stop' treats that as fatal, which killed the same script three times for three unrelated reasons.
  • The execution policy on a normally configured Windows machine blocks an unsigned .ps1 azd hook. I stopped using PowerShell for hooks at all.
  • Node 20 reached end of life in April 2026. App Service quietly substitutes the nearest LTS it actually has in your region, rather than failing.
  • And one of my own. I read an azd timeout warning as a deployment failure, and changed a working deployment strategy on the strength of it. The app had already started successfully two minutes after azd stopped watching. That sample’s README writes it up under a heading telling you not to make it.

The number on my screen while writing this

The token in front of me as I write this carries 75 minutes, the exact average Microsoft documents for the randomized 60 to 90 minute window. It is a small thing to see the number rather than read it. It is also the entire argument of this post in one line.

Run any of it against your own tenant. If it does not behave the way I have described, that is worth an issue rather than a comment thread. The whole point of putting code behind an architecture picture is that the picture can then be wrong out loud.

Cosmos DB Agent Memory Cost: Caching, RU Drivers, and a Pitfalls Roundup

Post 6 of 6 on Cosmos DB agent memory cost: the hard numbers post 1 promised back at the start of this series.

Post 1 opened with a 2017 CloudBrew reviewer calling a Cosmos DB proof of concept “an hour-long marketing pitch,” and my answer then was that cost means nothing without the revenue it enables. Five posts later, that argument still needs the numbers behind it.

Where This Post Picks Up

This post closes the series with them: what actually drives Cosmos DB agent memory cost at the RU level, how semantic caching cuts LLM spend specifically, what to monitor once an agent runs in production, and a pitfalls roundup that pulls every thread from posts 2 through 5 into one list.

Semantic Caching: Reusing What You Already Paid to Compute

An LLM call is almost always the most expensive, highest-latency step in an agent’s request path, far more than any Cosmos DB read or write. A semantic cache cuts that cost by skipping the LLM entirely when a close-enough answer already exists. Instead of matching prompts by exact string, it vectorizes the incoming prompt and runs a similarity search against the prompt-completion pairs already sitting in the cache. The mechanics are the same as the VectorDistance() query from post 3; only the container changes, from memory to cache.

Two details make this different from a normal cache, and both matter for cost control. First, the similarity threshold is a real trade-off, not a default to leave alone: set it too high, and near-identical questions still miss and hit the LLM anyway; set it too low, and the cache starts returning answers that don’t actually match what the user meant. Second, a semantic cache needs the same context window as an LLM.

Cache only the raw prompt, and two different users who each ask “what’s the second largest?” in unrelated conversations get whichever answer the cache stored first, correct for one thread, wrong for the other. Vectorize a slice of the conversation history alongside the latest prompt, the way post 2’s turn-based schema already structures it, and the cache lookup carries the same context the LLM would have used. TTL handles cleanup the same way it does for turns in post 2, with one addition worth considering: a hit-count field that increments on each cache hit lets a pruning pass keep frequently reused entries around longer than questions the cache only ever answered once.

What Actually Drives Cosmos DB Agent Memory Cost

Four decisions drive most of the RU bill for an agent workload, and they’re not evenly weighted. Partition skew usually costs the most: at Cosmos DB Conf 2026, an engineer described a production account running at 100% RU utilization, throttling and retrying under load, where the obvious fix looked like provisioning more throughput. The real cause turned out to be a single logical partition absorbing over 80% of traffic, one automated integration account driving most writes under a partition key that looked reasonable on paper. Fixing the data model, without adding a single RU of throughput, dropped utilization to 20–35% and made the throttling disappear entirely. More throughput would have masked that problem, not fixed it.

Item size and shape matter next, and this series already covered the mechanism in post 2: one document per turn keeps writes small and cheap. At the same time, one-document-per-thread turns every new message into a full-item rewrite that gets steadily more expensive as the thread grows. Vector index choice is the third lever. DiskANN’s sharding and approximate search solve a scale problem post 3 already flagged, and paying for that complexity below roughly ten thousand vectors buys nothing quantizedFlat wasn’t already providing. TTL is the fourth: expired short-term memory that never actually expires, because someone set a container-level default once and never came back to it, quietly inflates storage and index size on data nobody queries anymore.

The Fifth Lever: Consolidation

There’s a fifth lever underneath all four, and it’s the one this whole series has been arguing for since post 1: consolidation. Running a cache, a relational store, and a dedicated vector database as three separate systems means paying for three separate throughput allocations, three separate operational surfaces, and cross-system network cost on every request that touches more than one of them. One Cosmos DB account carrying memory, search, and cache together shares throughput across all three instead of over-provisioning each in isolation; the 2017 CloudBrew critique missed the same argument when it judged the account’s line-item cost without asking what running three systems instead of one would have cost by comparison.

Monitoring: What to Watch Once It’s Running

Three signals catch most problems before they become an incident. Change feed lag matters most for the multi-agent handoffs post 4 covered — a growing lag between a write and the Function that reacts to it means a specialist agent is falling behind the conversation, not just running a little slower. Break RU consumption out per container instead of watching one account-wide total, and it shows which specific workload is driving spend — turns, checkpoints, or the semantic cache — instead of leaving that as a guess. And for catching expensive patterns before they ship at all, the Azure Cosmos DB VS Code extension’s Query Insights and Index Advisor flag cross-partition queries, missing filters, and indexing gaps directly in the editor, well before a query shape becomes production traffic.

Pitfalls Roundup: Every Thread from This Series

  • Unsharded vector index in a multitenant app (post 3) — without a vectorIndexShardKey, semantic search scans every tenant’s vectors, not just the current one.
  • Thread-per-item growth (post 2) — an item that grows by one append per turn gets more expensive to write with every message, and eventually hits a hard size limit.
  • Missing or forgotten TTL (post 2) — short-term memory nobody set an expiration for keeps sitting in the container indefinitely, quietly inflating storage.
  • Cross-tenant memory leakage (posts 3 and 4) — a global vector index or an unscoped checkpoint container lets one tenant’s context bleed into another’s.
  • Treating change feed as globally ordered (post 4) — ordering holds within a partition key, never across the whole container.
  • Confusing Foundry Agent Service Classic and New containers (post 5) — the newest trap in the list, and already the most common source of “why is my thread storage empty” reports.

What I’d Ask the Product Team

Multi-region writes for a globally distributed agent multiply throughput cost by the number of regions. The guidance so far is “add regions only where traffic justifies it,” which is reasonable. Still, it leaves the actual crossover point (how much traffic, at what latency requirement) for each team to work out through trial and error rather than a documented formula. A cost calculator that takes a workload shape and a target latency and outputs a recommended region count would save a lot of that guesswork.

Where This Is the Wrong Answer

Not every agent workload belongs on one account. A workload with one enormous, narrowly specialized vector search needs tens of billions of vectors; nothing else can still get better unit economics from a dedicated vector database that specializes in exactly that shape, rather than a general-purpose store carrying memory, search, and cache together. The unified argument holds for the vast majority of agent workloads this series has covered, not for every workload unconditionally.

Closing the Series

That 2017 reviewer wasn’t wrong that the account cost more than a bare-minimum alternative; the miss was judging that cost without the workload it made possible, the same mistake the Figma AWS costs piece argued against in a completely different context. Six posts and one real production case study later, Cosmos DB agent memory cost comes down to the same handful of decisions this series has covered since post 2: partition key, item shape, index choice, and TTL, with semantic caching and consolidation compounding the savings on top. That’s the whole series in one sentence, and it’s the argument I’d have made at CloudBrew in 2017 if I’d had the RU numbers to back it up yet.


Sources

Cosmos DB Foundry Agent Service: Bring-Your-Own Thread Storage

Post 5 of 6 on Cosmos DB Foundry Agent Service integration, owning the thread store instead of leaving it opaque behind a managed API.

Posts 1 through 4 assumed you manage the Cosmos DB account directly. Foundry Agent Service changes that assumption by default: spin up an agent the basic way, and Microsoft manages the thread store for you, out of reach of a direct query. Standard agent setup flips that around. Cosmos DB Foundry Agent Service integration lets threads, system messages, and agent metadata land in a Cosmos DB account you own, sitting right where the schema, search, and checkpointing patterns from the rest of this series already apply.

Why Bring-Your-Own Thread Storage Matters

Data residency, security review, and auditability all get harder when a vendor holds conversation history in a store you can’t query. Standard setup solves that by provisioning three customer-owned resources instead of one managed black box: Azure Storage for uploaded files, Azure AI Search for the agent’s vector stores, and Azure Cosmos DB for everything: conversational messages, threads, and agent metadata. Cosmos DB carries the load that matters most for this series: it’s where the actual conversation lives.

Inside enterprise_memory: Cosmos DB Foundry Agent Service Containers

Standard setup names the resulting database enterprise_memory, and container names inside it depend entirely on which Foundry Agent Service runtime the agent runs on. Foundry Agent Service (Classic) writes to three containers: thread-message-store for end-user conversation messages, system-thread-message-store for internal system messages, and agent-entity-store for agent metadata like instructions and tools. Foundry Agent Service (New) writes to two different containers instead of agent-definitions-v1 and run-state-v1, and neither runtime reads the other’s containers. Check thread-message-store for an agent running on the New runtime, and it comes back empty, not because BYO thread storage failed, but because the data landed somewhere else entirely.

Provisioning Standard Agent Resources

Provisioning means more than a Cosmos DB account on its own. Standard setup also expects an Azure Storage account, an Azure AI Search resource, and an Azure Key Vault for secrets, alongside a deployed agent-compatible model. Once those exist, Microsoft’s Bicep template accepts the resource IDs of existing accounts and wires up the rest: account and project connections, role assignments, and the capability hosts that tell Foundry where agent state actually lives

Two things catch people off guard here, so it’s worth flagging both before you deploy. First, throughput: your Cosmos DB account needs at least 3,000 RU/s total 1,000 RU/s for each of the three baseline containers and that number scales up with every additional project sharing the account, since each project gets its own container set. Undershoot it, and the deployment doesn’t fail quietly; it throws CapabilityHostProvisioningFailed during the capability host step. Second, roles: the project’s managed identity needs Cosmos DB Operator at the account level to provision containers, plus Cosmos DB Built-in Data Contributor at the database level for enterprise_memory. The database-level scope covers every container inside it, so a single role assignment handles the whole set instead of one per container.

Querying Thread History Directly

This is what bring-your-own thread storage actually buys over the default: a direct line into conversation history that Foundry’s own API doesn’t expose. Once a thread exists in thread-message-store (Classic) or run-state-v1 (New), you can run the same vector, full-text, and hybrid queries from post 3 against it — same RANK RRF(...) syntax, same partition-scoped WHERE clause. The only difference is the container: Foundry manages it instead of your own code.

Treat this as a connection into Foundry Agent Service, though, not a replacement for it. Foundry still owns thread creation, run orchestration, and tool invocation; Cosmos DB Foundry Agent Service integration only changes where the resulting data sits and who can query it directly.

Pitfalls

Checking the wrong container set. The single most common source of “BYO thread storage isn’t working” reports is querying Classic’s containers for an agent running on the New runtime, or vice versa. Confirm which runtime a project uses before assuming a missing thread means a broken connection.

Underprovisioning throughput for multiple projects. The 3,000 RU/s floor covers one project’s container set. Add a second project to the same Cosmos DB account, and the container count and the RU/s requirement under it doubles. CapabilityHostProvisioningFailed almost always traces back to this, not to a misconfigured connection.

Assuming you can edit a capability host after creation. You can’t. Pointing a project capability host at the wrong Cosmos DB resource ID means deleting and recreating the project, not patching the connection — worth getting right on the first deployment rather than treating it as a setting to adjust later.

Next: Cost, Caching, and Production Pitfalls

That settles Cosmos DB Foundry Agent Service integration for readers who need full ownership of thread storage rather than a managed default. The last post in this series pulls back to the practitioner-notes view: RU cost drivers, semantic caching, and a closing pitfalls roundup across everything posts 2 through 5 have covered.


Sources

Five RAG Architectures in Real Azure Code

Over the past few months I kept running into the similar looking infographics, in one form or another: five or six boxes, each a named RAG architecture, arrows showing how a query flows through it. Hybrid RAG. GraphRAG. Agentic RAG. Corrective RAG. Multimodal RAG. They’re useful as vocabulary. They are not implementation guides. None of them show you the part that actually takes the time.

So I built all five RAG architectures, on Azure, against one shared corpus and one shared set of test questions, and measured what came out. This post is the result: what each diagram leaves out, what the equivalent Azure code actually looks like, where the real deployment pitfalls were, and a comparison table built from real runs, not from argument.

The corpus is a fictional Dutch health insurer, Zorgverzekeraar Meridiaan, the same one I’ve used in a couple of other posts in this series. Nine documents: dental and physiotherapy policies, a provider network, an authorization process, a member complaint and the quarterly report that restates it, a stale FAQ sitting next to the current policy, a reimbursement table, and a scanned claim form. Fifteen questions, tagged by which pattern they were designed to stress. All five patterns answer all fifteen questions, so the comparison is apples to apples.

The repo is at github.com/steefjan1/five-rag-patterns if you want to run it yourself.

What the diagram shows vs. what the Azure code does

Hybrid RAG

The diagram draws dense and sparse retrieval as two separate paths that merge into a box labeled Reciprocal Rank Fusion. That box is mostly a non-event on Azure. Azure AI Search’s hybrid query type takes a vector query and a text query together and fuses them server-side. There is no RRF code to write.

What actually takes engineering effort is the index schema: chunk granularity (I chunk by document section, not by a fixed token window, so a retrieval unit is a coherent answer, not an arbitrary slice), which fields are filterable versus searchable versus vector, and whether semantic ranking is worth its cost on top of the fusion you already get for free.

Measured: recall 1.00 across all fifteen questions, the best of any pattern on pure coverage. Precision sits at 0.38, diluted by a fixed top-5 retrieval regardless of how many documents a question actually needs. Cheapest sane baseline in the set: $0.0026 and 2.00 seconds per query.

GraphRAG

The diagram draws one static graph: entities, edges, a subgraph retrieval step, a box for community summaries. What it doesn’t draw is that the graph has a maintenance cost. Community detection (Louvain, via networkx, which runs fine at this corpus’s scale without a dedicated graph database) and community summarization are real compute and real Azure OpenAI spend, paid once at setup and again every time the graph changes enough to shift community boundaries. Nothing about that shows up in the box-and-arrow version.

Entity linking here is a cheap substring match against entity names, not an embedding call, which is part of why this pattern is the cheapest per query in the whole set. Retrieval is a graph walk: two hops, both directions, so a question like “which hospital did this referral come from, and which GP group refers into that hospital” resolves correctly even though no single document states the answer. It’s two separate edges, walked in sequence.

Measured: recall 1.00 on its own three relational questions, the two-hop case included, and 0.21 on the other twelve. No other pattern swings that hard between its own territory and everything else. $0.0013 per query, the cheapest pattern here, in the narrowest lane.

Agentic RAG

The diagram shows a planner routing to tools and a reasoner that loops “until confident.” There is no upper bound drawn anywhere on that loop. Left alone, that is a cost leak, not a reliability feature, so the actual implementation caps it at five iterations and reports hitting the cap as its own outcome rather than quietly forcing an answer and calling it clean.

The other thing worth knowing if you’re building this on Azure: the AI Foundry Agent Service SDK bypasses API Management for its own LLM calls. I found this the hard way on an earlier project in this series. If your governance model depends on APIM, that means routing tool-calling agents through the standard OpenAI SDK pointed at the gateway, not through the framework’s own agent runtime, or every rate limit and kill switch you built stops applying the moment the agent framework makes the call instead of your code.

Two of this pattern’s four tools aren’t retrieval at all. The dental waiting-period and annual-maximum arithmetic is transcribed from the policy documents as plain code, not left for a language model to compute from prose. Insurance eligibility math is exactly the kind of thing an LLM gets subtly wrong under pressure, and exactly the kind of thing code gets right every time.

Measured: precision 1.00, recall 1.00 on its own two questions, and it’s the only pattern that doesn’t collapse elsewhere: recall 0.85 on the other thirteen, because it always has a general search tool as a fallback when nothing more specific fits. That’s the real finding here. It’s not that Agentic RAG is “better,” it’s that it hedges.

Corrective RAG

The diagram shows retrieve, grade, then three branches: answer, rewrite the query and loop back, or fall back to a web search. The rewrite loop has an arrow pointing backward and no stated exit condition. A closed corpus also has no web to fall back to, so “incorrect” here means declining to answer rather than guessing.

The corpus has a document built specifically to test the grading step: an archived FAQ with a plausible, wrong number sitting right next to the current policy with the right one. A pattern with no grading step retrieves both and may cite either. This one grades the retrieval, asks the model to identify which passage is authoritative using published dates and explicit supersession language, and only feeds the authoritative passages to the final answer. What got fetched and what got used are tracked separately on purpose, so a working grader shows zero distractor citations even though the distractor was retrieved.

Measured: precision 1.00, recall 1.00 on the three distractor questions, confirmed live, not just in the design. The more interesting number is that its other twelve questions score better (precision 0.79) than its own target slice (0.67). The grading discipline isn’t just catching the one distractor it was built to catch, it generalizes. That comes at a real cost: 3.97 seconds average latency, roughly double every other pattern, and the highest cost per query in the set, because a full run can mean three model calls instead of one.

Multimodal RAG

The diagram’s box says “shared multimodal embedding model (e.g. CLIP or ColPali),” which means self-hosting an embedding model. That’s a heavier operational commitment than anything else in this comparison needs, and it’s avoidable. This uses caption-then-embed instead: Document Intelligence extracts the actual structure of the reimbursement table (tables are exactly where a vision model hallucinates a plausible-looking row that isn’t in the source, so that step doesn’t get skipped), the vision-capable chat deployment captions the scanned claim form directly, and both captions get embedded with the same text-embedding-3-large deployment every other pattern uses. Same index Hybrid RAG built, two more documents in it, no new index and no new field.

One implementation note that cost real iteration: a first version of the captioning prompt asked for verbatim transcription, which correctly produced the form’s Dutch date format and Dutch status text. A validation step checking the caption against a hand-written ground truth flagged that as a mismatch, because the ground truth expected ISO dates and English. That’s not a captioning bug, it’s a prompt that needed to ask for normalization, not transcription. Worth deciding on purpose, since a shared index with mixed date formats and mixed languages retrieves worse than a normalized one.

Measured: recall 1.00 across the board and groundedness 1.00 on its own two questions, with no degradation on the other thirteen. That composability is the finding: this pattern is Hybrid RAG’s exact retrieve-and-answer loop plus two documents, and the numbers confirm that composition was free.

The comparison table

Same corpus, same fifteen questions, one pass, all five RAG architectures.

PatternAvg latencyCost/queryOverall precisionOverall recallOwn-target recall
Hybrid2.00s$0.00260.381.001.00 (n=5)
GraphRAG2.32s$0.00130.200.371.00 (n=3)
Agentic2.09s$0.00450.450.871.00 (n=2)
Corrective3.97s$0.00510.770.971.00 (n=3)
Multimodal2.18s$0.00270.381.001.00 (n=2)

“Own-target” means the small subset of the fifteen questions each pattern was actually designed to answer (GraphRAG’s two-hop provider questions, Corrective RAG’s stale-document case, and so on). Every pattern hits recall 1.00 in its own lane. What separates them is what happens outside it: GraphRAG falls to 0.21 recall on the other twelve questions, Agentic RAG only falls to 0.85, and Hybrid, Corrective, and Multimodal don’t fall at all, because their retrieval isn’t scoped to a narrow entity set in the first place.

Two honest caveats on this table. Precision across every pattern is capped low by a fixed top-5 retrieval regardless of how many documents a question actually needs, so precision here measures retrieval breadth more than answer quality, read recall and the own-target column as the more meaningful columns. And the cost figures come from a placeholder price table, not a live Azure billing export, useful for comparing patterns against each other, not for a procurement conversation.

Deployment pitfalls

Every one of these was a real failure against a live Azure subscription, not a hypothetical.

Pinned model versions rot. A deployment written against gpt-4o-mini version 2024-07-18 failed eight months later with ServiceModelDeprecated. The fix wasn’t a newer pin, it was to stop pinning: leave the deployment’s model version empty and let Azure resolve the current default, and check az cognitiveservices model list -l <region> -o table before assuming a model name is still offered at all.

The account kind changed. Azure OpenAI is now provisioned through Foundry as kind: 'AIServices', not the older kind: 'OpenAI'. Same deployment mechanism underneath, different account kind and a newer API version. A template written against the old kind fails Cognitive Services preflight validation, not at compile time.

A malformed policy XML fails at ARM validation, not at Bicep build time. An APIM policy embedded as a Bicep string had a raw double-quoted path literal sitting inside an already double-quoted XML attribute. bicep build compiled it clean, because Bicep has no way to know a string is meant to be well-formed XML. The actual break only showed up against the live ARM validation API, after Azure AI Search, Cosmos DB, and the Foundry account had already finished provisioning. A small script that compiles the template and separately parses every embedded policy string as XML catches this before the next azd up, not during one.

RBAC role assignments alone don’t turn on Azure AD authentication. Azure AI Search kept returning a flat 403 on every data-plane call despite two correctly scoped role assignments, because the service still only accepted API-key authentication. Nothing had told it to accept AAD tokens at all. The fix is a separate property, disableLocalAuth: true, on the search service itself. If a resource with roles that look correct still refuses an authenticated caller, check the resource’s own auth settings before re-checking the role assignment.

Where none of this is the answer

None of these five patterns is the right first move for a small, stable knowledge base. Plain vector search, no fusion, no graph, no grading, no agent loop, is the correct answer until you can name the specific failure mode you’re buying insurance against. Every pattern here is a bet against one kind of failure, and every bet has a cost attached whether or not you ever collect on it.

Don’t build all five for one real system either. Pick based on the failure mode your domain actually has. GraphRAG only pays for itself if your questions are genuinely relational, multi-hop, the kind no single document answers. If they’re not, you’re paying setup cost and getting a narrower Hybrid RAG. Agentic RAG’s flexibility costs a planning call before any retrieval happens at all, worth it if your questions genuinely vary in shape, wasted overhead if they don’t. Corrective RAG’s discipline costs roughly double the latency of everything else in this comparison. That’s a fine trade when a wrong answer is expensive and a two-second wait isn’t. It’s a bad trade for a chat widget where speed is the product.

What this actually proves

The infographic’s taxonomy is real. These five RAG architectures are genuinely different, with genuinely different failure modes, and that part of the diagram holds up. What doesn’t hold up is the implication that the hard part is choosing between them. The hard part, in every case, was the piece the diagram didn’t draw: RRF turned out to be free because Azure AI Search already does it, but community detection is not free and has to be redone as the graph changes. An agent loop needs a hard cap or it’s an open-ended bill. A query rewrite loop needs the same cap for the same reason. A self-hosted multimodal embedding model turned out to be avoidable entirely, caption-then-embed onto infrastructure you already have gets you most of the way there.

The single most useful number in this whole exercise might be the smallest one: Corrective RAG’s grading step scored better on questions it wasn’t built for than on the one it was. That’s a pattern worth paying attention to. The things that make a RAG system more disciplined in one specific place often make it more disciplined everywhere, not just in the place you were testing for.

If you want to see the actual failure modes up close rather than the aggregate table, the earlier posts in this series go deeper on two of them: what naive RAG diagrams leave out covers the hybrid retrieval and groundedness gaps in more detail, and choosing between RAG, GraphRAG, and Agentic RAG when auditability is the constraint makes the conceptual case this post backs with numbers.

The full repo, including the corpus, the eval harness, and every pattern’s implementation, is at github.com/steefjan1/five-rag-patterns.

Multi-Agent State and Checkpointing with Cosmos DB

Post 4 of 6 on Cosmos DB multi-agent state coordinating what several agents know about the same conversation, without a separate message bus.

Post 3 settled retrieval for a single agent working alone. This post is about what changes once a second agent enters the picture. Coordinating what several agents know about the same conversation turns out to be a different problem from storing and retrieving one agent’s memory, and Cosmos DB multi-agent state ends up resting on two mechanisms this series already covered: hierarchical partitioning from post 2, and change feed from post 1’s retail monitoring callback.

Shared but Separable: What Changes with Multiple Agents

A single agent needs one memory scope. Moreover, a multi-agent system needs two at once: shared memory that every agent can read and write for coordination, and private memory that lets each agent keep its own persona, prompts, and reasoning history separate from the others. Lose the separation, and agents start bleeding into each other’s context. Lose the sharing, and they can’t coordinate at all.

A triage agent, a product agent, and a specialist agent a common pattern in production multi-agent apps each hold their own scoped state. Still, all three write to the same underlying container, so any of them can pick up where another left off.

LangGraph Checkpointing on Cosmos DB Multi-Agent State

LangGraph’s checkpoint interface persists a graph’s state after every step, and Cosmos DB has more than one implementation of it: langgraph-checkpoint-cosmosdb on PyPI, and the checkpoint saver that ships inside langchain-azure-cosmosdb. Both plug into the same standard LangGraph pattern: compile the graph with a checkpointer, then pass a thread_id on every invocation:

from langgraph. graph import StateGraph
from langgraph_checkpoint_cosmosdb import CosmosDBSaver
checkpointer = CosmosDBSaver(
endpoint=cosmos_endpoint,
key=cosmos_key,
database_name="agentmemory",
container_name="checkpoints",
)
graph = StateGraph(AgentState)
# add_node / add_edge calls wire up triage -> specialist routing here
app = graph.compile(checkpointer=checkpointer)
config = {"configurable": {"thread_id": "contoso:thread-1234"}}
app.invoke({"messages": [...]}, config=config)

Encode tenantId:threadId into the thread_id string, and the checkpointer’s hierarchical partitioning lines up with the [tenantId, threadId] partition key from post 2 — the same pattern manages per-user, per-session state at scale, this time for graph checkpoints instead of turn-based memory items. Microsoft’s own multi-agent-langgraph sample builds a personal-shopper scenario on exactly this foundation: a triage agent routes requests, and a product agent answers them using retrieval-augmented generation against the same Cosmos DB account.

Change Feed as the Handoff Mechanism

Post 1 covered change feed as the primitive behind a 2023 retail monitoring solution, a new record in Cosmos DB firing a Function that could raise an incident. The same primitive coordinates agent handoffs: one agent writes a turn, a Function listening on the container’s change feed picks it up, and it hands the conversation to whichever agent should act next. No polling loop checks for new work; the write itself is the signal.

A minimal handoff trigger, using the turn-based schema from post 2:

python

import azure.functions as func
def main(documents: func.DocumentList) -> None:
for doc in documents:
if doc.get("targetAgent") == "specialist":
notify_specialist_agent(doc["threadId"], doc["turnIndex"])

The triage agent sets targetAgent on the turn it writes; the Function reacts to that write and wakes the specialist agent for that thread.

A Second Worked Example: Spring AI for Java Shops

Python and LangGraph aren’t the only path here. Spring AI 2.0 shipped with a Cosmos DB-backed vector store and memory integration for Java, and Microsoft’s multi-agent-spring-ai sample mirrors the LangGraph pattern in Java: multiple agents, one Cosmos DB account, the same shared-but-separable memory shape. Worth a look if the rest of the stack runs on the JVM rather than Python.

Pitfalls

Shared containers without tenant or session isolation. A checkpoint container that mixes every tenant’s graph state leaks context across customers the moment a partition or vector index goes unsharded; the same isolation failure post 3 flagged for vector search is now showing up in agent state instead of retrieved memories. Apply the same [tenantId, threadId] discipline to checkpoints that post 2 applied to turns.

Treating change feed as globally ordered. Change feed guarantees order within a single partition key, not across the whole container. Moreover, a handoff design that assumes “the Function always sees writes in the exact order they happened across every agent” breaks the moment two agents write to different partitions at close to the same time. Design handoffs so each step only depends on ordering within its own thread’s partition, not on a global sequence that Cosmos DB never promised.

Next: Cosmos DB Inside Microsoft Foundry Agent Service

That covers Cosmos DB multi-agent state when you manage the account directly. Post 5 covers the other path: Microsoft Foundry Agent Service’s bring-your-own thread storage, where Cosmos DB still does the work, but Foundry owns the orchestration layer on top of it.


Sources

Finding the Right Memory: Vector, Full-Text, and Hybrid Search in Cosmos DB

Post 3 of 6 on Cosmos DB agent memory search, because storing memory well doesn’t guarantee you’ll retrieve the right piece.

Post 2 ended with the schema settled and retrieval still open. This post closes that gap: the practical mechanics of Cosmos DB agent memory search, one container, four query patterns. Run the same question four different ways against that container, and it comes back with four different answers, because “find the right memory” isn’t one query pattern; it’s at least three, and knowing which one to reach for is most of the job.

Vector Indexing for Cosmos DB Agent Memory Search

Cosmos DB supports two vector index types, and the right one depends almost entirely on how many vectors you’re searching, not on anything specific to agents.

quantizedFlat compresses each vector and scans the compressed space exactly. It suits smaller workloads (tens of thousands of vectors) and trades a small amount of accuracy for lower RU cost and faster scans. For a single tenant’s short-term memory, this is often enough on its own.

DiskANN, on the other hand, indexes vectors for approximate nearest-neighbor search and scales to hundreds of thousands or billions of embeddings, with dynamic updates and strong recall even at that size. Post 1 already leaned on DiskANN as part of the case for Cosmos DB as a unified store; this is the mechanism behind that claim.

Sharding the Vector Index for Multitenant Isolation

DiskANN doesn’t have to search across every vector in the container. A vectorIndexShardKey partitions the index itself by a property you choose: session, user, or tenant, so a query only searches candidates within that shard instead of the whole container.

That maps directly onto the partition key work from post 2: set the vectorIndexShardKey to tenantId, or to the same [tenantId, threadId] pair you already use as the partition key, and semantic search for one tenant never touches another tenant’s vectors. A global, unsharded index still works and makes searching everything at once simpler, but it’sonly appropriate for a single-tenant app or a genuinely shared knowledge base where cross-tenant recall is the point rather than a leak.

Full-Text Search: When Precision Beats Semantics

Vector search finds what’s semantically similar. Sometimes semantically similar isn’t what you want — a customer asking about “the refund policy” needs the actual refund policy language, not five conceptually related passages about returns in general.

Full-text search on Cosmos DB handles that case through BM25, a statistical ranking function that scores by term frequency and document length. Cosmos DB applies linguistic processing automatically: tokenization, stemming, case normalization, so “running” still matches “run” or “ran.” It’s the right tool whenever exact terms or phrases carry meaning that a vector embedding would blur.

Hybrid Search: Combining Both with RRF

Most agent memory queries don’t need to choose between semantic and lexical relevance; they need a blend of both. That’s what Reciprocal Rank Fusion (RRF) does: it takes the vector-similarity ranking and the BM25 ranking for the same result set and merges them into one combined rank, instead of forcing a pick between the two.

In practice, this shows up as a single ORDER BY RANK RRF(...) clause, which the next section demonstrates directly.

Four Ways to Ask the Same Question

Take the turn-based schema from post 2 — tenantId, threadId, turnIndex, messages, embedding, content and run the same underlying question against it four ways. (content is a flat, denormalized copy of the turn’s text, added specifically because Cosmos DB doesn’t support wildcard array paths like /messages/*/content in a full-text policy or index the full-text and hybrid queries below point at c.content rather than c.messages for exactly that reason.)

Most recent, by recency:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
ORDER BY c.turnIndex DESC

Semantic, by vector similarity:

SELECT TOP 5 c.messages, VectorDistance(c.embedding, @queryVector) AS score
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
ORDER BY VectorDistance(c.embedding, @queryVector)

Hybrid, blending both with RRF:

SELECT TOP 5 c.messages, VectorDistance(c.embedding, @queryVector) AS score
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
ORDER BY VectorDistance(c.embedding, @queryVector)

Keyword, by exact phrase:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = @tenantId AND c.threadId = @threadId
AND FULLTEXTCONTAINS(c.content, @phrase)
ORDER BY c.turnIndex DESC

Run all four against a thread where a customer asked about refunds three times, in different words, across twenty turns, and the differences stop being theoretical fast: recency surfaces whichever turn happened most recently, even if it’s off-topic; semantic search pulls in every conceptually related turn, including the ones that used different words entirely; hybrid balances the two; keyword search returns only the turns that used the customer’s actual phrase, and ranks them by recency underneath that filter.

Running These Queries in Data Explorer

The four queries above use parameterized SQL, the same form search.py, from the companion repo behind this series, sends through the Python SDK, which binds @tenantId, @queryVector, and @phrase properly before the query runs. Paste them as-is into the Azure Portal’s Data Explorer query pane instead, and two things break, neither of which is a schema or code bug:

Data Explorer’s query box doesn’t bind named parameters. A query that leaves @tenantId unresolved either matches nothing and returns “No results” silently, or for VectorDistance() inside ORDER BY and FullTextScore() fails to compile outright, because both functions require their arguments to resolve to literal values at query-compile time rather than at execution time.

Swap every @parameter for a literal value and all four run cleanly. Against the seeded sample data (tenantId = "contoso", threadId = "thread-1234", searching for "refund"):

Recency, with literals:

SELECT TOP 5 c.messages, c.turnIndex
FROM c WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
ORDER BY c.turnIndex DESC

Semantic, with literals:

SELECT TOP 5 c.messages, VectorDistance(c.embedding, [0.8196, 0.6392, -0.2471, 0.1608, -0.8667, 0.4902, -0.2549, 0.2235]) AS score
FROM c
WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
ORDER BY VectorDistance(c.embedding, [0.8196, 0.6392, -0.2471, 0.1608, -0.8667, 0.4902, -0.2549, 0.2235])

Hybrid, with literals:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
ORDER BY RANK RRF(
VectorDistance(c.embedding, [0.8196, 0.6392, -0.2471, 0.1608, -0.8667, 0.4902, -0.2549, 0.2235]),
FullTextScore(c.content, "refund")
)

Keyword, with literals:

SELECT TOP 5 c.messages, c.turnIndex
FROM c
WHERE c.tenantId = "contoso" AND c.threadId = "thread-1234"
AND FULLTEXTCONTAINS(c.content, "refund")
ORDER BY c.turnIndex DESC

Pitfalls

Reaching for DiskANN on a small dataset. DiskANN’s approximate search and sharding options solve a scale problem. Below roughly ten thousand vectors, quantizedFlat gets equivalent recall for less operational complexity and lower RU cost. Default to DiskANN because it sounds like the “serious” choice, and you’ve added index-shard decisions to a workload that never needed them.

A global vector index in a multitenant app. Skip the vectorIndexShardKey, and a semantic query searches every candidate in the entire container, tenant boundaries or not. Nothing stops the query from surfacing another tenant’s conceptually similar memory in the result set unless a WHERE clause happens to filter it back out after the fact, and relying on a filter to catch what the index itself should have scoped is the kind of gap that shows up in an audit, not in testing.

Forgetting WHERE filters still apply. Vector and hybrid queries look like they replace normal filtering, but ORDER BY VectorDistance(...) or ORDER BY RANK RRF(...) still runs inside a WHERE-scoped query, same as any other. Leave the WHERE c.tenantId = @tenantId AND c.threadId = @threadId clause off a semantic query, and it searches everything the container holds, not just the thread the agent is currently in.

Next: Coordinating Multiple Agents

That settles Cosmos DB agent memory search for a single agent working alone. Coordinating what several agents know about the same conversation is a different problem, and it’s where change feed, a mechanism post 1 already covered as a callback to the 2023 retail monitoring work, comes back to tie multi-agent state together. That’s post 4.


Sources