Messages, Events or Streams: Choosing Azure Messaging by What You Send

I’ve worked in the integration space for a long time, and messaging has been at the center of it for most of that time. Queues, topics, publish/subscribe, event routing, streaming: the patterns stay, while the products and their names keep changing. The first question I ask in a design review hasn’t changed either. What is actually on the wire, and what does the sender expect to happen next? When teams ask me to compare Azure messaging services, I answer with that question before I name a single product.

In July I wrote Azure Messaging and Orchestration for Integration Architects as part of the Azure PaaS for Integration Architects series. That post drew one line through messaging: “If losing a message would be a business incident, it belongs on Service Bus. If losing it would just mean a missed notification, Event Grid is fine.” I still stand by it. But it deliberately covered only two of the Azure messaging services, and there are five options for moving data between systems: Storage Queues, Service Bus queues, Service Bus topics, Event Grid and Event Hubs.

Recently I saw a LinkedIn post with an infographic listing all five, and it’s a good moment for me to paint a clearer picture. Cheat sheets like that one usually sort by product. After years of fixing integrations, I look at the payload first. This post makes that distinction explicit, gives a clear rule for when to use what, and backs it with one order flow that uses all five services. You can deploy it yourself from github.com/steefjan1/messaging-choices-azure.

Messages, events and streams

Microsoft Learn’s comparison of the messaging services puts the distinction into two definitions:

  • An event is “a lightweight notification of a condition or state change. The publisher has no expectation about how the event is handled.”
  • A message is “raw data produced by a service to be consumed or stored elsewhere.” And: “A contract exists between publisher and consumer.”

Events split once more. A discrete event reports a state change you can act on. An event series is part of a time-ordered stream that you analyze. One reading from a delivery van tells you nothing; ten thousand of them tell you which van is overheating.

That gives three kinds of payload, and each maps to a family of Azure messaging services. Throughput, cost and the technology stack all matter, but they come second.

Azure messaging services: when to use what

Here’s the short version I give teams, one rule per option.

UseWhenNot when
Storage QueueYou have background work that must get done eventually, and losing a message costs you a rerun, not a customer. Reports, thumbnails, file jobs.You need dead-lettering, ordering, duplicate detection or transactions. You’d end up building them yourself.
Service Bus queueYou’re sending a command with exactly one owner, and the business depends on it: ShipOrder, IssueRefund, PostInvoice.Nobody would notice the loss, or several teams need the same message.
Service Bus topicA business fact has several parties that each owe it something: OrderPlaced goes to payment, inventory and email, each with its own copy, retries and dead-letter queue.The subscribers owe nothing and just want to know. That’s an event.
Event GridSomething happened, and whoever cares may react: a blob landed, a resource changed, a domain event fired. The publisher doesn’t know or care who listens.The handler must process it or the business breaks. Put a queue behind Event Grid for that.
Event HubsYou ingest a continuous stream (telemetry, clickstream, logs) where the value is in the aggregate and several readers need the same data.You need a retry or an acknowledgment per message. A stream has neither.

Two quick tests settle most debates. Would losing one of these be a business incident? Then it’s Service Bus, as in the July post. Would anyone notice a single missing item? If not, and there are thousands per second, it’s a stream.

Why the rules fall where they do

The table is the short version. The reasons come from how each service behaves when something fails.

Storage Queue or Service Bus

Storage Queue is cheap, it’s HTTP, and it scales to a backlog larger than 80 GB. Learn’s side-by-side comparison is blunt about what you give up: messages up to 64 KB, no ordering guarantee, no duplicate detection, no transactions and no dead-letter queue. To find poison messages, “the application examines the DequeueCount property” and moves them itself. If you use Azure Functions, the host does that for you and fills a -poison queue. That’s a Functions feature, though, not a Storage one.

Service Bus is built for work the business depends on. The broker owns a dead-letter queue per entity and records a reason on each message. You get sessions for FIFO per key, duplicate detection on MessageId, and transactions. Standard tier caps messages at 256 KB; Premium goes to 100 MB and gives you isolated capacity. One tier catch: Basic supports queues only, with no topics, sessions or duplicate detection.

Queue or topic comes down to ownership: a command has one owner, a business fact has several. OrderPlaced sounds like an event, which is why people reach for Event Grid. By Learn’s definition it isn’t one. The publisher very much expects someone to take the money. That expectation is a contract, and a contract belongs on Service Bus.

Event Grid routes, it doesn’t queue

Event Grid pushes to Functions, webhooks, Logic Apps and queues, and namespace topics add pull delivery and an MQTT broker. Its delivery behavior tells you what it’s built for. The default retry policy is 30 attempts within 1,440 minutes, backing off from 10 seconds to 12 hours. “Event Grid doesn’t guarantee order for event delivery.” And “by default, Event Grid doesn’t turn on dead-lettering”: an event that runs out of retries is simply dropped unless you configured a storage container for it.

That’s fine for a notification, not for work. So the pattern that holds up in production puts a queue behind Event Grid: the event says “a blob arrived,” and a queue holds the job of processing it, with its own pace, retries and poison handling.

Event Hubs is a log, not a queue

Event Hubs is the one service in the set where reading a message doesn’t remove it. It’s a partitioned, append-only log. Consumers track their own position with checkpoints, and “checkpointing is the consumer’s responsibility.” Events stay until the retention period expires: up to 7 days on Standard and 90 on Premium and Dedicated.

Three things follow from that. Ordering holds only within a partition, so you pick a partition key (the van id, the device id) to keep related events in order. Consumer groups give independent readers of the same stream, each with its own checkpoint. And there’s no per-message acknowledgment and no dead-letter queue. A consumer that can’t handle event 4,711 has to decide for itself what to do and then move on. That’s why it’s excellent for telemetry and wrong for order processing, however high the volume.

One order flow, all five services

The companion repo puts this into one e-commerce flow. It runs on a single Azure Functions app (.NET 8 isolated, Flex Consumption) and deploys with azd up to Sweden Central. All connections use the app’s managed identity, and local auth is disabled on Service Bus, Event Hubs and both storage accounts.

Orders: messages on Service Bus

POST /api/orders publishes OrderPlaced to a Service Bus topic with four subscriptions. The fraud-review subscription uses a SQL rule, total >= 1000, so it only sees expensive orders. Rules evaluate application properties, not the body, and that’s why the publisher sets them explicitly:

var message = new ServiceBusMessage(BinaryData.FromObjectAsJson(order, Json))
{
MessageId = order.OrderId, // duplicate detection keys on this
Subject = nameof(OrderPlaced),
};
message.ApplicationProperties["total"] = (double)order.Total; // the SQL rule reads this

The payment handler settles its own messages. It dead-letters an order it can never charge with reason InvalidTotal. If it hits a failure that looks transient, it throws and lets the broker retry, and after three deliveries the broker dead-letters the message with MaxDeliveryCountExceeded. On success, it sends a ShipOrder command to a session-enabled queue with SessionId = customerId.

One detail is easy to miss: sending the command and completing the incoming message are two operations, not one transaction. After a crash between them, the order is redelivered and the command sent again. Duplicate detection on shipments drops the second copy, because the command’s MessageId is derived from the order id. At-least-once delivery means idempotent handlers, and broker features like this take some of that work off your hands.

Product images: an event, then work

A blob upload to product-images raises BlobCreated on an Event Grid system topic. The subscription is set to 30 attempts within 24 hours and has dead-lettering turned on, since it’s off by default. OnImageUploaded doesn’t process the image. It drops an ImageJob on a Storage queue and returns. ImageWorker does the work, and anything that fails three dequeues lands in image-jobs-poison. The uploader knows nothing about any of this, which is the point of an event.

Telemetry: a stream on Event Hubs

POST /api/telemetry simulates delivery vans and sends batches to an event hub with four partitions, keyed by van id. Two functions read the same events through two consumer groups: one computes averages, the other flags overheating engines. Take the alerts function down for an hour and it resumes from its checkpoint, and the aggregator never notices.

What the demo run showed

scripts/demo.ps1 drives every scenario, including a duplicate order, a poison order and a corrupt image, and then peeks at where the failures ended up. On my run against Sweden Central, the output looked like this:

WhereMessageWhat the platform recorded
orders/payment dead-letter queuezero-total orderreason InvalidTotal, description “Order total 0 is not payable.”, delivery count 0
orders/payment dead-letter queuepoison customerreason MaxDeliveryCountExceeded, “could not be consumed after 3 delivery attempts”, delivery count 3
shipments dead-letter queuenoneempty: every paid order shipped
image-jobs-poison Storage queuecorrupt-104734.pngthe job body and a dequeue count of 0

That table is the lesson in one picture, and it’s where Azure messaging services differ most. The Service Bus dead-letter queue belongs to the broker, and every entry tells you why it’s there and how often it was tried. The dead-lettered order also never burned a retry, because the handler knew it could never succeed. The Storage poison queue only exists because the Functions host moved the message there, and it arrives with no reason and a dequeue count reset to 0. The history of three failed attempts is gone. If a Storage queue carries work you’ll have to explain in an incident review, you’ll be rebuilding that history yourself.

The publisher also can’t tell a duplicate from a first send. Both POSTs of the same order came back accepted, and the API logged “published” twice. The broker drops the second copy silently, inside its 10-minute duplicate detection window.

Where this framing is the wrong answer

No rule for picking Azure messaging services survives every context. These are the ones I’ve seen bend it.

  • When the volume argument is real. A few thousand orders per second is still a message workload. Before you move it to Event Hubs and rebuild dead-lettering yourself, look at Service Bus Premium with more messaging units.
  • When the devices need identity or commands back. Real vans don’t write straight to Event Hubs. If you need per-device authentication or cloud-to-device messages, look at IoT Hub or the Event Grid MQTT broker in front of the stream.
  • When you don’t need a broker at all. A synchronous call with a retry policy is simpler than a queue with one producer and one consumer in the same deployment. Add a queue when you need to decouple, not by default.
  • When the event crosses an organizational boundary. Service Bus subscriptions are cheap inside one team. Across teams or companies, Event Grid with CloudEvents and filtered subscriptions is usually the looser and better contract.
  • When cost dominates. Event Hubs Standard bills per throughput unit whether you send anything or not, and Service Bus Standard has a base charge. Run azd down --purge after the demo.

Takeaway

Choosing between Azure messaging services doesn’t start with “which Azure service?” It starts with “what am I sending, and what does the sender expect?” A command goes to a Service Bus queue, a business fact with several owners to a topic, and cheap background work to a Storage queue. A notification without expectations goes to Event Grid, with a queue behind it for the actual work. A stream goes to Event Hubs. As I wrote in July, the craft isn’t picking a winner. Most real systems use several of these, and the interesting design work happens at the seams between them.

The repo, diagrams and demo script are at github.com/steefjan1/messaging-choices-azure.

Sources

When Agentic Workloads Break the PaaS Assumptions

This series started with a map and grew into seven pieces. Five layers came first: compute, where load shape picks the service; messaging and orchestration, where two questions replace four product choices; data patterns, where idempotency and the outbox keep a platform correct; governance and identity, where policy and audit become the compliance posture; and observability and FinOps, where behaviour and cost become visible. Then came the lens: from design to demonstrable operation, the shift from “is it built?” to “can we operate it responsibly?”

Every one of those pieces rests on a shared assumption. The system does what you told it to do. You wrote the workflow, you defined the routes, you set the policies, and the platform executes them. That assumption has held for every integration platform I’ve built. Agentic workloads break it. So this capstone asks what changes when the thing making decisions inside your platform is a model, not your code, and why each layer, plus the readiness lens itself, deserves a second look because of it.

The assumption agentic workloads break

Conventional integration is deterministic. A message arrives, a workflow runs its defined steps, a router sends it where the rules say. You can read the code and know what will happen. You can test every path. When something fails, you trace it to a step you wrote.

Agentic workloads replace part of that determinism with a model that decides at runtime. The agent reads context, picks a tool, interprets the result, and chooses the next action. Moreover, it does so differently depending on inputs you didn’t fully anticipate. That’s the point of it: the flexibility is the feature. But it means you can no longer read the code and know what will happen. So the ground under every layer shifts: behavior is no longer exactly what you specified.

None of this argues against agentic workloads. It argues for revisiting each layer with the shift named explicitly. Let’s do that.

Compute: the loop changes the shape of the work

The compute layer sorted workloads by load shape: steady request traffic to App Service, event-driven bursts to Functions or Container Apps. Agentic workloads add a shape that sorting didn’t account for: the loop.

An agent doesn’t process a request and return. It reasons, calls a tool, waits, observes, and reasons again, sometimes for many cycles, before it finishes. That’s neither a clean request-response nor a discrete event. Instead, it’s a long-running loop of unpredictable duration with external calls in the middle. So the compute question changes. You’re no longer asking “steady or bursty” alone. You’re asking how to host something that runs for seconds or minutes, holds state across tool calls, and scales on a dimension concurrent reasoning loops that CPU-and-memory autoscale captures poorly. Container Apps with event-driven scaling often fit better here than App Service, and the orchestration frequently belongs in a workflow engine rather than raw compute.

Messaging and orchestration: the agent is a non-deterministic router

The messaging layer drew a clean line. Deterministic routing rules sent messages where the logic dictated. An agent orchestrating tool calls is, in effect, a router too, but a non-deterministic one. It decides which tool to call from its reading of the context, not from a rule you wrote.

The reliability consequences are real. Delivery guarantees still matter; an agent that triggers a business action still needs that action to occur exactly once, so Service Bus and the idempotency store in the data layer remain as relevant as ever. What changes is predictability. You can’t fully anticipate which actions the agent will trigger, or in what order. Therefore, the orchestration has to stay correct under sequences you didn’t design for. One practical lesson from building these loops applies directly: agent outputs rarely arrive as the clean structures a deterministic step would emit, so you build explicit bridges between agent actions rather than assuming shape.

Data: state and correctness under non-determinism

The data patterns held a platform correct when systems it didn’t control misbehaved. Agentic workloads make those patterns more necessary, not less, and they add one more.

Idempotency matters more because an agent may retry a tool call or repeat an action as it reasons, so the dedup store carries a heavier load. The outbox matters just as much, because an agent-triggered write still has to propagate reliably. Workflow state matters more too, since the reasoning loop is exactly the kind of long-running, restart-surviving process that needs durable state and a correlation ID. And then the new one: conversation and context state. An agent carries context across turns, and that context has to live somewhere durable and queryable, which explains why a flexible document store keeps showing up as the default for agentic conversation state. The access pattern points at the store. Same principle as the map, applied to a new kind of state.

Governance and identity: where the assumptions break hardest

This layer changes most, and I’d insist any integration architect think it through before shipping an agentic workload.

The governance layer secured a deterministic platform. Identity answered who the caller was; policy constrained what the platform could be. Both still matter. However, agentic workloads open a gap that neither fully closes. Identity secures who the agent is. It does not touch what a poisoned tool result or a manipulated retrieved document makes the agent do. Prompt injection rides in through the data the agent requested inside the reasoning loop, downstream of the perimeter check everyone assumes protects them.

So the governance layer needs additions a deterministic platform never required:

  • Authorization moves per-action: A validated identity at the edge isn’t enough. Each tool call the agent makes needs its own check: is this specific action allowed for this tenant right now? The perimeter check happens once; the risk recurs on every call inside the loop.
  • Recovery means compensation, not retry: Agent actions have side effects across systems. A failed sequence three actions deep can’t restart from the top; it needs compensating actions to undo what already happened. That’s saga-style thinking, and you design it; it doesn’t emerge.
  • Containment has to be possible: When an agent misbehaves, you stop it fast, and at more than one layer. Layered containment, from a single configuration flip-up to a full block, turns “contain the agent” from an incident-call debate into a seconds-long operation.
  • Evaluation becomes a first-class layer: Operational observability tells you the agent is running. It doesn’t tell you the agent’s outputs are quietly degrading. Under the EU AI Act’s oversight and transparency duties, that stops being optional polish and becomes evidence you’re meeting an obligation.

The readiness lens, asked again

The design-to-operation post posed the question that decides go-live: not “is it built?” but “can we operate it safely, recoverably, auditably, and predictably?” Agentic workloads sharpen every word of that sentence.

Safely now includes per-action authorization and containment, because the threat walks in as data. Recoverably now means compensation and sagas, because retry alone can’t undo side effects. Auditably now covers what the agent accessed, which tool it called, why it acted, and what policy constrained it evidence the EU AI Act increasingly expects. And predictably is precisely the property the agent gave up, which is why the surrounding architecture has to supply it instead. The production baseline, the demonstrable-versus-designed test, the three moments of readiness all of it still applies. Each bar sits higher.

The revised framework

The map closes with five questions. For agentic workloads, they hold, and each gains a harder edge. Before an agentic workload goes near a real system, I’d add these:

Can I host a long-running reasoning loop, not just a request or an event? Can my orchestration stay correct when I can’t predict the action sequence? Does my data layer hold conversation state as well as business state, with idempotency doing heavier duty? Is authorisation per-action, not just per-identity? Can I contain a misbehaving agent in seconds? And can I evidence what the agent did, why, and whether its quality held?

Those aren’t different questions from the series. They’re the same layers, asked again under non-determinism, and then held up against the readiness lens one more time.

The shape of it

Agentic workloads don’t replace the Azure PaaS foundation an integration architect builds on. They stress it. Every layer in this series still includes compute, messaging, data, governance, and observability, but each one now supports a workload that decides for itself at runtime. The compute layer meets the loop. The messaging layer meets a non-deterministic router. The data layer meets conversation state and heavier idempotency. The governance layer meets a threat that walks in through the front door as data. And the readiness lens meets a workload that surrendered predictability, so the architecture has to supply it.

The through-line of the whole series holds here too. The model is the least differentiated part of a production agent. What separates a demo from something you can run against real systems in a regulated industry is the architecture around it: the same layers, asked harder, and proven in operation rather than promised in design. So the foundation was never wasted. It’s exactly what agentic workloads need, applied with the assumptions made explicit.

That’s the series. Start at the Azure PaaS map for the layer-by-layer foundation, take the design-to-operation lens with you as the test, and come back here for what changes when the workload thinks for itself.

Production Readiness: Closing the Execution Gap

Every post in this series so far has covered a layer: compute, messaging and orchestration, data patterns, governance and identity, observability and FinOps. This one isn’t a layer. It’s a question that cuts across all of them: when is the platform actually ready for production?

That question is harder than it looks, because “we built it” and “we can run it” are different claims. I’ve watched more than one integration platform pass every technical checkpoint and still fall short of production-ready, not because the design missed anything, but because nobody had turned that design into something enforceable, operable, and provable. So this post is about the gap between those two states, and how you close it. It’s the through-line under every layer, and it’s the thing that turns a strong architecture into a platform you can responsibly put load on.

The shift: from “is it built?” to “can we operate it?”

Here’s the single most useful reframe I know for this stage. Stop asking “is the platform technically built?” and start asking “can we operate it safely, recoverably, auditably, and predictably?”

Those are not the same question. The first is about whether the components exist and connect. The second is about whether, when something goes wrong at 2 am, someone can see what happened, understand it, recover from it, and prove afterward that they handled it correctly. A platform can pass the first test comfortably and fail the second completely. And the second test is the one that actually determines whether you should go live. So the moment you catch yourself saying “it works,” push on: does it work in a demo, or does it work under a failure you didn’t plan for?

The core move: make the implicit explicit

Most integration platforms at this stage share the same shape. The design is good, and someone has largely written it down. But a lot of what matters lives in documentation, in code that’s still evolving, or in the heads of the people who built it. That works fine while a small team builds the foundation. It stops working the moment the first real production use case lands, because implicit choices become whatever the first team happens to decide.

The fix is a production baseline: an explicit, enforceable statement of what production use demands. It names the mandatory components and patterns, pins down the environment profiles, fixes the security controls and monitoring standards, and settles the recovery agreements and release criteria. Its job is to pull those decisions out of documents and habits and into something the platform enforces, so the first production use case inherits the decisions rather than reinventing them.

Without that baseline, every early integration is free to make its own choices, and you accumulate inconsistency, technical debt, and a future re-platforming you didn’t budget for.

Design versus demonstrable: the recurring gap

The same gap shows up in every layer, and once you see it, you can’t unsee it. On security, the design names Zero Trust, least privilege, and pipeline-driven change, yet permanent broad access rights sit in the environment, quietly contradicting it. Observability design specifies OpenTelemetry and required fields, but the alert rules and dashboards never make it into the infrastructure. And the CI/CD design describes a full release chain with quality gates, while the pipelines really only cover the dev environment.

In each case the design is right and the demonstrable working is missing. That’s the pattern to hunt for when you assess readiness: not “did someone design this?” but “does the platform actually enforce this design somewhere it can prove?” Anything that lives only as intent is a gap, however good the intent.

What operational readiness actually covers

Technical function is necessary but not sufficient. Operational readiness asks a distinct set of questions, and that set decides go-live. In practice, it comes down to whether you can demonstrably answer these:

Can you see what’s happening: monitoring, chain-level tracing, message-level insight? When an incident hits, does someone triage it, and do they have the information they need to do so? For recovery, can the platform handle errors, replay from a known point, and fall back on a real Business Continuity and Disaster Recovery plan rather than an RTO written in a document? Afterward, can you prove what happened through an audit trail, an access log, and change history? On access control, do you separate permanent from elevated rights, and dev from production? For cost, can you attribute it, budget against it, and tier retention? And finally, can you hand it over — have you defined operational ownership, or does every incident route back to the people who built it?

If any of those answers is “only in the design,” please fix it before go-live, not after the first incident teaches you the hard way.

The boundary that decides scale: central versus decentral

The other thing production-readiness has to settle is a boundary, not just a checklist. Most modern integration platforms want value-stream teams to deliver independently within central guardrails. That’s the right ambition. But it only works when you draw the line between central platform ownership and team autonomy deliberately, because that line runs through every layer: API governance, messaging configuration, RBAC, pipeline use, monitoring, error handling, lifecycle, support.

Get the line wrong and you recreate the exact problem the platform set out to solve. The platform team becomes the bottleneck again — the single point through which every change, approval, and incident has to pass. So team autonomy isn’t just a tooling question. It needs explicit ownership agreements, release paths, access models, quality controls, and operational responsibilities. The tooling enables autonomy; the agreements make it safe.

Shared components are where this bites hardest. A shared API gateway, a shared message broker, a shared logging workspace these touch technology, security, governance, operations, cost, and autonomy all at once. They make the platform economical, and they carry the biggest scaling risk. For each one, decide deliberately: why does it stay shared rather than isolated, who owns it, who may change it, how do you monitor it, and how do you attribute its cost? Leave those implicit and the shared component quietly becomes everyone’s dependency and no one’s responsibility.

Readiness isn’t one moment — it’s three

The last reframe worth making: “ready” isn’t a single bar. It’s three different bars at three different moments, and conflating them is how platforms either over-build early or under-prepare for scale.

First go-live. The bar here is the minimum production baseline and demonstrable operational readiness. Not every capability has to be complete, but the ones that are preconditions for running safely in production do. This is where the baseline, the recovery plan, and the release criteria have to be real.

First team onboarding. The bar shifts to whether the federated model actually works in practice. Can one real team deliver independently within the central guardrails, without quality, security, or consistency buckling under the first real use? This is a practice test, and it’s better, even, to let some things get concrete here rather than designing them fully in the abstract.

Scaling to many teams. Now the bar is repeatability. Anything that worked at one team through direct conversation now has to become standardised, documented, and reproducible: lifecycle policy, cost allocation, onboarding, support model, versioning, exception handling. Direct alignment doesn’t scale; product-steering does.

Naming which moment a given concern belongs to is half the battle. It stops you from demanding scale-grade rigour before first go-live, and from discovering at team five that nobody built the repeatable version.

Where this thinking gets over-applied

Consistent with the series, the honesty section. “Production baseline” thinking is right, but it can tip into paralysis.

Not everything has to be complete before first go-live. The three-moments split exists precisely so you don’t. Demanding full lifecycle policy, mature FinOps, and a complete federation model before a single use case runs is how a platform never ships. Match the rigour to the moment.

A baseline that only flags is a baseline that gets ignored. The whole point of the production baseline is that the platform enforces it. A pile of documented-but-unenforced standards manufactures the appearance of readiness without the substance, which is more dangerous than an honest gap, because it invites false confidence.

You can make a decision deliberately without making it central. Drawing the central-versus-decentral line carefully doesn’t mean pulling everything central. Sometimes the deliberate call is “this is team-owned,” and recording that reasoning is the point, not the direction.

The shape of it

For an integration architect, production-readiness isn’t a technical checkpoint; it’s the shift from a platform that’s built to one you can operate responsibly. Make the implicit explicit in a production baseline. Hunt the gap between what’s designed and what’s demonstrable. Answer the operational-readiness questions before go-live, not after. Draw the central-versus-decentral line deliberately, especially for shared components. And treat “ready” as three moments, not one. Do that, and the layers from this series stop being a good architecture on paper and become a platform you can actually run.

This is the lens that ties the series together. The Azure PaaS map has the layer-by-layer foundation; this post is the question you hold every layer up against before you put production load on it.

Azure Functions vs. Logic Apps vs. Power Automate: When to Use What

If you work anywhere near the Microsoft ecosystem, you have probably run into all three of these services. You have probably also run into the confusion around them. They all “automate” something. They all show up in architecture conversations. On the surface, their marketing pages sound almost interchangeable.

This post continues the Cloud Perspectives Azure PaaS series. It follows recent entries on Azure Functions as a serverless agents runtime and managed identity in Logic Apps Standard. Those posts went deep on one service. This one steps back and compares all three. The usual shorthand, Functions for code, Logic Apps for integration, Power Automate for business users, is cleaner than reality.

In practice, Functions is a capable integration tool in its own right. Logic Apps’ headline B2B/EDI capability comes bundled with an extra resource and its own bill. Power Automate is not even an Azure product. Picking the wrong tool does not just produce a clunkier solution either. It can mean months of maintenance pain, licensing costs nobody budgeted for, or a workflow that cannot scale. Here is a more honest breakdown of how the three differ, and when each one earns its place in your architecture.

Where each service actually lives

Before comparing capabilities, it helps to see where these tools sit organizationally. That placement drives billing, governance, and who owns the resource day to day.

Azure Functions and Logic Apps are Azure resources. You provision them in the Azure portal, under an Azure subscription, next to your virtual machines and storage accounts. Platform teams building governance models, like the ones described in Azure governance and identity for integration architects, treat them accordingly.

Power Automate is different. It is licensed through Microsoft 365 and the Power Platform admin center. That difference is not a footnote. It determines which admin center you open when something breaks, and which budget line absorbs the cost.

Azure Functions: built for developers who want full control

What it is

Azure Functions is Microsoft’s serverless compute service. You write code in C#, Python, JavaScript, TypeScript, Java, PowerShell, and more. It runs in response to an event. That event might be an HTTP request, a new file landing in Blob Storage, a message hitting a queue, or a timer firing. You never manage the underlying servers. You simply ship functions and let Azure handle the scaling.

It is also a first-class integration tool

Functions gets typecast as “the compute one” while Logic Apps gets credited with “integration.” That framing sells Functions short. Its HTTP triggers and rich set of bindings include Service Bus, Event Grid, Cosmos DB, and Blob Storage. Those let a function sit in the middle of a system-to-system exchange just as naturally as a Logic App can.

For stateful, long-running orchestration across multiple systems, the exact scenario people usually reach for Logic Apps for, Durable Functions provides that same pattern in code. You get full testability and source control with it. For developers building the kind of serverless orchestration covered in Azure Functions as a serverless agents runtime, Functions is a legitimate path for integration work, not a fallback.

When to use it

Reach for Functions when you need custom logic, complex calculations, or heavy data transformation that a visual designer cannot express cleanly. Reach for it when you want total control over code, dependencies, and third-party libraries. It also fits when you are building microservices and APIs that must perform well under load. It also fits when you are doing systems integration and would rather express it in code than in a visual designer.

Typical use cases: processing images uploaded to Blob Storage, powering a custom REST API for a mobile or web app, running scheduled jobs, handling real-time IoT telemetry, and orchestrating multi-step integration workflows through Durable Functions.

Trade-offs

Functions requires real programming skills. Cold starts on the Consumption plan can add latency to infrequent workloads. You also own more of the maintenance and security surface than a managed workflow tool would give you.

On the upside, Functions is usually the cheapest of the three at scale. The Consumption plan includes a substantial monthly free grant, around one million executions and 400,000 GB-seconds. You only pay for what you actually run.

Logic Apps: built for enterprise-grade integration

What it is

Logic Apps is Azure’s platform-as-a-service for orchestrating workflows across systems. Think of it as the enterprise integration layer. It gives you a visual designer, so it sits at a lower code level than Functions. Even so, it targets IT and integration teams rather than casual business users. Logic Apps shines when you need to connect many systems reliably, at scale, with proper DevOps practices wrapped around it. That is the kind of governance discussed in Azure data patterns for integration architects.

When to use it

Reach for Logic Apps when you need to integrate multiple systems, spanning cloud, on-premises, and SaaS, with formal reliability and monitoring requirements. It also fits B2B or EDI-style exchanges over AS2, X12, or EDIFACT. And it fits any scenario where you want a pay-per-execution model that supports high-volume enterprise workflows without managing infrastructure yourself.

Typical use cases: syncing a CRM like Salesforce with an on-premises SQL database, exchanging invoices with trading partners using industry-standard protocols, and orchestrating responses to Azure alerts or resource deployments.

The Integration Account catch

One nuance is easy to gloss over. The B2B/EDI capability is not something a plain Logic App gives you out of the box. It requires provisioning a separate resource, an Integration Account, to store trading partners, agreements, schemas, and certificates. You then link that account to your Logic App.

Integration Accounts carry their own tiered pricing across Free, Basic, Standard, and Premium levels. In other words, “use Logic Apps for B2B/EDI” really means “use Logic Apps plus an Integration Account.” That adds both cost and an extra resource to manage, something a lot of comparisons leave out entirely.

Trade-offs

Logic Apps requires an active Azure subscription and has a steeper learning curve than Power Automate. Unlike Power Automate, it also has no built-in desktop or RPA automation. Billing runs on trigger, action, and connector executions, which can add up faster than an equivalent Functions workload.

Still, you get native Visual Studio and Git integration, strong monitoring, and no hard execution limits. All of that matters at enterprise scale.

Power Automate: built for business users who need speed

What it is

Power Automate is the low-code, no-code member of the trio, built for productivity rather than infrastructure. It lets business users, not developers, automate day-to-day tasks such as approvals, notifications, report generation, and data syncing between Microsoft 365 apps.

It is not actually Azure

Here is a distinction worth making explicit, since the three tools so often get lumped together as “Azure services.” Power Automate is not an Azure service. It belongs to the Microsoft Power Platform, licensed alongside Power Apps, Power BI, and Copilot Studio. That licensing typically runs through Microsoft 365 plans, standalone per-user or per-flow licenses, or a free tier, not an Azure subscription.

The one exception is pay-as-you-go licensing, which lets you bill Power Automate usage against an Azure subscription as an alternative payment mechanism. Even so, that is a billing convenience, not evidence that Power Automate lives in Azure.

Functions and Logic Apps are Azure resources you provision in the Azure portal, under an Azure subscription, alongside your VMs and storage accounts. Power Automate, by contrast, is a Microsoft 365 offering that happens to interoperate with Azure resources through connectors. That is a real architectural distinction, not just a licensing footnote. It affects who owns the resource, where governance sits, and which admin center you troubleshoot in when something breaks.

When to use it

Reach for Power Automate when you want to boost individual or team productivity without writing code. It also fits when you need to automate UI-based tasks on legacy software through desktop flows, essentially RPA. It fits too when your workflow lives mostly inside Microsoft 365, across Outlook, Teams, SharePoint, and similar apps.

Typical use cases: routing document approvals through Teams and Outlook, using desktop flows to pull data out of an old legacy application, and automatically saving email attachments to a SharePoint folder.

Trade-offs

Power Automate is the fastest option to deploy for non-developers, and it integrates tightly with the Power Platform. That said, licensing can get expensive at scale, debugging and version control stay limited compared to code-first tools, and performance throttles kick in on high-volume runs.

A quick rule of thumb, caveats included

  • If the job needs raw coding power, fine-grained control, or code-first integration, reach for Azure Functions. That includes integration work. Do not rule it out just because Logic Apps carries the “integration” label.
  • If the job needs visual, governed workflow orchestration across systems, reach for Logic Apps. Budget for an Integration Account on top if EDI or B2B is involved.
  • If the job needs a fast, no-code fix for a Microsoft 365-centric business process, reach for Power Automate. Account for it as Power Platform or M365 licensing rather than an Azure cost.

The real power comes from combining them

These three are not really competitors. They are layers. A common pattern looks like this: Power Automate handles the front-end business process, say a Teams approval flow. That triggers a Logic App to orchestrate the broader integration across systems. The Logic App, in turn, calls an Azure Function to run the custom logic or heavy computation that neither low-code tool expresses well.

Used this way, each tool does the part it is actually good at. Power Automate handles speed and accessibility. Logic Apps handles governed integration at scale. Azure Functions handles anything that needs real code. The mistake is not choosing one of these. It is assuming you have to choose only one.

Azure Observability and FinOps for Integration Architects

In the Azure PaaS map post, observability was folded into the governance layer, with a note that Application Insights and Azure Monitor are non-negotiable. That was true, but it undersold the topic. For an integration platform specifically, observability isn’t a sub-bullet of governance. It’s the layer that decides whether you can actually run the thing in production.

So this post pulls observability out and gives it room. And it brings FinOps along, because the two share a root: you can’t manage what you can’t see. One makes system behaviour visible; the other makes cost visible. Both turn a platform from “it runs” into “we can run it responsibly.” Azure observability and FinOps, treated together, are what separate a platform that works in a demo from one you can operate under real load.

The gap between design and demonstrable operation

Here’s the pattern I see most often on integration platforms. The observability design is excellent. There’s a logging standard, a tracing approach, a set of required fields. Then you look at the actual infrastructure, and none of it is enforced. The alert rules aren’t there. The dashboards aren’t built. The diagnostic settings aren’t wired. The design lives in a document; the platform doesn’t know about it.

That gap matters more than it sounds. A monitoring standard that depends on discipline and review isn’t a platform capability; it’s a hope. The moment a team ships an integration without the dashboards, the standard quietly failed. So the real work in this layer isn’t designing observability. It’s making observability demonstrable: wired into the infrastructure, enforced in the pipeline, and impossible to skip.

Let’s walk what that means in practice.

OpenTelemetry as a platform contract, not a suggestion

Most mature integration platforms land on OpenTelemetry as the instrumentation standard. That’s the right call. W3C Trace Context propagates a trace across services, traces and metrics and logs share a model, and you avoid inventing your own correlation scheme. So far, so good.

The catch is that “we use OpenTelemetry” is a design statement, not an enforced one. For it to be a contract, three things must be true. First, the required fields, resource attributes, trace fields, and domain identifiers have to be defined explicitly, not left to each team’s judgment. Second, that definition has to be validated somewhere automatically, ideally at pull request. Third, the platform components themselves have to emit the standard, so a trace actually runs unbroken from the API gateway through messaging to the backend. Miss any of those, and you have telemetry that mostly correlates, which is worse than none, because it looks trustworthy right up until the incident where it isn’t.

Tracing the chain, not just the components

Azure gives you per-resource monitoring for free. You can see API Management’s metrics, Service Bus’s queue depth, and a Function’s execution count. That’s component monitoring, and it’s necessary but not sufficient. An integration platform’s job is to move a message across those components, so the question that matters is whether you can follow a single message or transaction through the entire chain.

That end-to-end view has to map onto the layers of your integration architecture, because each layer asks a different question. The consumer-facing layer cares about availability, latency, error rates, and throttling per channel. The process layer cares about routing, transformations, retries, and failures in async steps. The system-facing layer cares about dependencies on backends’ response times, timeouts, and contract breaks. Without that layered, chain-aware view, you get plenty of technical detail per Azure resource and almost no ability to reason about the integration as a whole.

Message-level insight and the async recovery problem

Component metrics tell you the platform is busy. They don’t help the person who has to answer “what happened to order 47821?” For that, an operator needs message-level insight: business identifiers, error categories, chain status, the last successful step, and retry state. Structured logging with domain attributes a flow ID, a message ID, a route key, and an error category is what makes that possible. And it has to come with explicit data classification, masking, retention, and access rules, because business identifiers in logs are exactly the kind of data a regulator asks about.

Then there’s recovery, which is where the compute choice comes back to bite. Async, message-driven processing needs a replay story: when something fails partway through, you need to know how far it got and re-drive it from there. A workflow engine often gives you some of this out of the box. Raw compute like Functions doesn’t, so you have to design the replay mechanism yourself, as part of the integration pattern rather than an afterthought.

The pattern that works: treat the message on the bus as a reference, not the full payload. Pair it with the claim-check pattern, in which the bus carries technical and functional metadata: trace ID, flow ID, message ID, route key, error category, retry count, and a pointer to the payload, safely stored in storage. Define checkpoints along the flow. Then, on failure, you can determine where processing succeeded and re-drive from the right point, with idempotency (from the data patterns post) making the re-drive safe. For fully synchronous request-response, re-driving belongs with the caller; the platform’s job there is clear error codes and traceability.

Monitoring as a Definition of Done

The single highest-leverage move in this layer costs almost nothing: make monitoring a Definition of Done for every integration. No integration ships without its dashboard, its alerts, its trace-context propagation, its required log fields, its retention setting. And this is the part that turns it from aspiration into capability: the checklist runs as a quality gate in the pipeline, not as a line in a review someone might skip.

That one change moves observability from “depends on the discipline of whoever built it” to “the platform won’t let you skip it.” It’s the difference between a standard and an enforced standard, and it’s the cheapest high-value thing on this entire list.

FinOps: cost is just another signal you can’t yet see

Everything above is about making system behavior visible. FinOps is the same discipline applied to cost. On an integration platform, it fails in the same way because cost visibility typically ends at the subscription or resource group boundary. That’s too coarse. It can’t tell you what an individual integration costs, or an API, or a queue, or a team’s share of a shared component.

Three FinOps problems come up on every integration platform:

  • Attribution needs a taxonomy. Without a consistent tagging scheme for value stream, team, environment, integration, API, owner, and cost category, cost remains a lump sum. With one, you can steer on cost per integration product rather than cost per subscription. This is the foundation; nothing else works without it.
  • Shared components are the hard part. Compute is easy to attribute when each team runs its own. But a shared API Management instance, a shared Service Bus namespace, a shared Log Analytics workspace those get used by everyone and billed centrally, and if you never build a distribution model, nobody owns the cost. The shared components that make the platform economical are exactly the ones whose cost is hardest to place. That’s not a reason to isolate everything; it’s a reason to deliberately decide the split.
  • Storage and retention are FinOps levers hiding within a compliance requirement. Observability generates data logs, traces, payloads held for replay, and dead-lettered messages. Compliance dictates how long you keep it. But retention length and storage tier are separate decisions. Data you must keep for audit doesn’t have to sit in a hot, queryable tier the whole time. Tie retention to data classification, then move cold data to cheaper tiers. The requirement is “keep it”; the FinOps move is “keep it cheaply.”

The through-line: FinOps on an integration platform isn’t financial reporting after the fact. It’s a design and governance concern, sitting right next to observability, because both are about seeing what the platform is actually doing.

Where this layer gets over-applied

Consistent with the series, the honesty section. Observability and cost control both have a failure mode of doing too much.

Not every signal deserves an alert. An alert that fires on something nobody acts on trains people to ignore alerts. Alert on what changes a decision; leave the rest on a dashboard. Alert fatigue is a real operational risk, not a sign of thoroughness.

Not every message needs full payload logging. Metadata-first is the right default. Payload logging belongs where there’s functional need and explicit consent, with masking and retention — not everywhere, because “log everything” is how sensitive data ends up somewhere it shouldn’t, and how your storage bill quietly triples.

Not every cost needs fine-grained attribution. Building per-message cost tracking for a low-volume internal integration spends more effort than the insight is worth. Match the granularity of attribution to the scale of the spend.

The shape of it

For an integration architect, observability and FinOps answer the same question in two currencies: what is the platform actually doing, and what is it actually costing? Wire OpenTelemetry in as an enforced contract. Trace the chain, not just the components. Give operators message-level insight and a real replay story. Make monitoring a Definition of Done the pipeline enforces. Then apply the same visibility to cost: a tagging taxonomy, a distribution model for shared components, and retention tiered by classification. Get both right, and the platform stops being a black box you hope is behaving and becomes one you can actually operate.

Want the layer this sits inside? The Azure PaaS map puts observability and governance in context against compute, integration, and data, and walks the five-question framework across all of them.

Azure Functions Hosted Skills: When to Use It

Editor’s note (September 2026): Microsoft has renamed this feature from “Azure Functions serverless agents runtime” to “Azure Functions hosted skills.” The content below has been updated to reflect the new name. The URL remains unchanged.

Most conversations about agent runtimes on Azure land on Azure AI Foundry Agent Service. That is the right starting point for teams that want a fully managed, enterprise-grade agent host. But it is not the only option and for event-driven scenarios, it is often not the best one.

The Azure Functions hosted skills is a preview programming model that lets you define agents as function apps. Agents are triggered by events, schedules, messages, or HTTP requests. They run on Flex Consumption with scale-to-zero, managed identity, and Application Insights. And they are deployed with azd like any other function app.

This post explains what the runtime actually is, how its three configuration files work together, and where it fits relative to Foundry Agent Service and Durable Functions. If you are new to Azure Functions as an AI platform, start with Azure Functions AI Integration: The Quiet Powerhouse which maps all four AI-enabled patterns. If you are looking specifically at MCP server hosting, Hosting MCP Servers on Azure Functions covers the three hosting options in detail.

How Azure Functions hosted skills work

The runtime is a programming model built on top of Azure Functions. When an event fires — a timer, an HTTP request, a queue message — the runtime starts the agent, runs it through Microsoft Agent Framework, and handles the trigger registration and endpoint wiring automatically.

You do not write trigger code, and do not implement an agent loop, and define three files, deploy a function app, and the runtime does the rest.

The core project files are:

  • .agent.md — defines the hosted skill: its instructions, its trigger, and the tools it can use
  • agents.config.yaml — app-wide runtime defaults, including the model deployment and any shared infrastructure (such as an Azure Container Apps dynamic session pool for sandboxed code execution)
  • mcp.json — lists the remote MCP servers available to hosted skills in the app
  • tools/ — optional folder containing custom Python tool functions decorated with @tool for app-specific logic
  • skills/ — optional folder containing reusable SKILL.md prompt assets that hosted skills can load on demand

The runtime discovers these files at startup, registers the required triggers and endpoints, and wires the agent to Microsoft Agent Framework. You can have multiple hosted skills in a single function app each defined in its own .agent.md file, all sharing the app-wide configuration.

What Azure Functions hosted skills deploy

The Microsoft Learn quickstart deploys two agents from a single function app:

Chat agent (main.agent.md) — an HTTP-triggered agent that exposes a debug chat UI in the browser. It can execute sandboxed Python code via an Azure Container Apps dynamic session pool and browse the web. No email tooling.

Blog summary agent (daily_microsoft_blog_summary.agent.md) — a timer-triggered agent. The YAML front matter in the file declares the schedule; the markdown body contains the agent instructions. On each timer fire, the agent gathers recent Microsoft blog posts, summarises them, and emails the digest via a managed MCP server connected to Microsoft 365 Outlook.

What gets provisioned by azd up for this template:

ResourcePurpose
Flex Consumption function appHosts the agents
Azure AI Foundry project + model deploymentLLM for agent reasoning
Azure Container Apps dynamic session poolSandboxed Python code execution
Storage accountFunction app state
Application InsightsMonitoring
Connector Namespace + M365 Outlook connectionEmail delivery (optional)
Managed MCP serverExposes the Outlook connector to agents

The provisioning is handled entirely by Bicep via azd you do not configure any of this manually.

How the agent definition files work

.agent.md is a markdown file with YAML front matter. The front matter declares the trigger and any agent-level configuration. The markdown body is the system prompt — the instructions the agent follows when it runs.

The timer-triggered blog summary agent front matter looks roughly like:

---
trigger:
type: timer
schedule: "0 0 8 * * *" # timer_trigger uses underscore in type field
tools:
- mcp_server: outlook
---

The markdown body below that is the agent’s instruction set what to gather, how to summarise it, how to format the email. You write it in plain English.

agents.config.yaml sets defaults that apply across all agents in the app. The model deployment is set here, so every agent uses the same Azure AI Foundry model unless overridden. The session pool endpoint for sandboxed code execution is also set here.

mcp.json lists the remote MCP servers the agents can call. In the quickstart template, this includes the managed MCP server for the Microsoft 365 Outlook connector when email delivery is enabled. The runtime reads this file at startup and makes those servers available to all hosted skills in the app.

How Azure Functions hosted skills differ from Foundry Agent Service

The distinction matters for architecture decisions.

Foundry Agent Service is a fully managed service. Microsoft operates the agent host. You configure agents through the Foundry portal or SDK, connect tools, and the service handles orchestration, state, and scaling. It has enterprise SLAs, built-in tooling, and a managed lifecycle.

Azure Functions hosted skills is a programming model you deploy yourself. You own the function app. You manage the deployment, the model connection, and the infrastructure. In return, you get the full Azure Functions hosting model — event-driven triggers, Flex Consumption billing, VNet integration, managed identity, and azd-based deployment pipelines.

The decision table:

SituationUse
Event-driven triggers, AI reasoning, scale-to-zero, VNet integrationAzure Functions hosted skills
Expose deterministic functions as tools for another AI clientAzure Functions MCP extension
Long-running, multi-step orchestration with human-in-the-loop approvalsDurable Functions
Create and manage AI agents without custom hostingMicrosoft Foundry or Copilot Studio

These are not mutually exclusive. Azure Functions hosted skills can call tools hosted in Foundry via MCP servers, and Foundry agents can call tools hosted in Azure Functions. The two can coexist in the same architecture.

How Azure Functions hosted skills differ from Durable Functions

Post 4 in this series covers Durable Functions for directed agentic workflows in detail, but the short version is:

Durable Functions is for directed, deterministic workflows — you define the steps, the model executes them in order, and Durable Functions handles state, retry, and fault tolerance. The workflow is predictable.

Azure Functions hosted skills is for autonomous agents — you give the agent instructions and tools, and Microsoft Agent Framework determines how to use them to accomplish the goal. The execution path is not predetermined.

If your AI-driven process has fixed, ordered steps and you need auditability, use Durable Functions. If you want the agent to figure out the steps, use Azure Functions hosted skills.

What to know before you build

It is preview. The programming model, file format, and configuration details are subject to change. Do not build production-critical workloads on this today without a plan for the preview-to-GA migration.

It requires a Foundry project and model deployment. The azd template provisions both automatically, but you need an Azure subscription with permissions to create Foundry resources and model deployments. Some organisations have restrictions on which model deployments are permitted.

The azd template provisions real Azure resources with real costs. The Flex Consumption plan itself is very low cost for low-traffic agents, but the Foundry model deployment, Container Apps session pool, and Connector Namespace resources all have associated costs. Review the Bicep templates in infra/ before running azd up in a production subscription.

Custom Python tools are how you add app-specific logic. Azure Functions hosted skills provides the agent loop and the MCP connections. For anything that requires your own code — calling internal APIs, reading proprietary data sources, applying business rules — you write Python tool functions and register them in the agent definition.

Getting started

The quickstart template is the right starting point:

azd init --template Azure-Samples/functions-quickstart-serverless-agents-azd -e serverless-agents
azd env set TO_EMAIL <your-email>
azd up

Review the three configuration files in src/ before deploying. They are short and readable, and understanding them before the first deployment saves debugging time later.

The email delivery step (setting TO_EMAIL and authorizing the Microsoft 365 Outlook connection) is optional. If you skip it, the timer agent still runs and returns its digest in the final response, which you can verify in Application Insights logs.

Try it with a weather sample.

If you want to see the runtime in action with a minimal, self-contained example before committing to the full quickstart, I built a companion sample: a weather chat agent that fetches live conditions and 3-day forecasts for any location using Open-Meteo, no API key, no M365 connector, no email setup required.

The agent is defined in a single main.agent.md file. It uses Python code execution via the Container Apps session pool to call the Open-Meteo API and returns structured weather data in the chat UI. Deploy it in three commands:

git clone https://github.com/steefjan1/weather-agents
cd weather-agents
azd up

Select Central US when prompted for location — the runtime is in preview and region availability is limited. The chat UI is at https://<function-app-name>.azurewebsites.net/api/agents/main/ once deployment completes.

The README documents two known issues you will hit if you try to build from scratch rather than the official quickstart: a broken transitive dependency in azurefunctions-agents-runtime that pins a yanked version of github-copilot-sdk, and the region constraint. Both are worth knowing before you invest time in a custom deployment.

Up next: Durable Functions as the Orchestration Layer for Directed Agentic Workflows

Citadel Grew Up: What the citadel-v1 Release Means If You Built on the AI Hub Gateway.

The Citadel Governance Hub accelerator that sits underneath my entire five-part Citadel Platform series just had a significant release. In addition, the citadel-v1 branch of the AI Hub Gateway Solution Accelerator repositions the project from “a solid APIM gateway pattern” to the official reference implementation of Layer 1 in Microsoft’s AI Citadel Blueprint.

I cloned the branch and went through it with one question in mind: what does this change for anyone who, like me, deployed and built on the earlier iteration? The answer starts with one finding. As a practitioner, the whole point of this blog is honesty: the API surface I used throughout the series is now explicitly labeled legacy.

Note that citadel-v1 has not yet been merged to main; if you deployed from the main branch without specifying --branch citadel-v1, you are on the earlier architecture.

Let’s start with the bigger picture, then get to that.

The 4-layer AI Citadel Blueprint

The README now frames the accelerator as one layer of a larger architecture. The AI Citadel Blueprint describes four interlocking layers, each with its own responsibility and implementation:

  • Layer 1, the Governance Hub, is this accelerator: runtime enforcement through a unified AI gateway, policy-as-code, identity validation, token rate limiting, content filtering, and cost attribution—everything my series built and tested lives in this layer.
  • Next, layer 2, AI Control Plane, covers the agent runtime, observability, and compliance: agent traces, AI evaluations, and fleet operations, implemented through the Microsoft Foundry control plane.
  • Subsequently, layer 3, Agent Identity, handles agent identity and lifecycle governance through Agent 365: unique agent identities, blueprints, shadow agent detection, and a sponsorship model.
  • And finally, layer 4, the Security Fabric, provides unified protection through Microsoft Defender for AI threat intelligence, Purview for data governance, and Entra for authentication and authorization.

Looking back at the series through this lens, my five posts covered Layer 1 thoroughly, and the registry work with Azure API Center reached into Layer 2 territory before the layer had that name. The kill switch from Part 4 sits squarely in Layer 1 as runtime enforcement. What the series never touched, and what I now have vocabulary for, is Layers 3 and 4. That’s useful: it turns “what’s missing from my platform” from a vague feeling into a named checklist.

What citadel-v1 Changes in the AI Hub Gateway: The API Surface

Here’s the finding that matters most if you followed the series. The new LLM Access Guide defines three API surfaces on the gateway, and it’s blunt about which one you should use.

The Azure OpenAI API surface, at /openai/deployments/{deployment-id}/*, preserves the exact URL shape the Azure OpenAI SDK expects. This is what Part 2 of my series wired the weather agent against, and it’s what every code sample in the series uses. The guide now labels it “legacy integration only,” for existing code that pins that URL shape. Not the target state for new work.

The Universal LLM API, at /models/*, exposes a clean OpenAI v1-compatible surface across many models and providers through a single stable path.

The Unified AI API, at /unified-ai/*, is the recommended surface: a single wildcard endpoint that serves OpenAI-compatible calls and every provider-native pattern with dynamic routing behind it.

LLM ACCESS guide

The citadel-v1 branch documents these three surfaces in its LLM access guide. Check the guides folder in the branch for the current filename, as the documentation is actively evolving.

I want to be precise about what this does and doesn’t mean. Nothing broke. Code targeting /openai/deployments/... keeps working, and the surface exists precisely because migrations take time. But the arrow points one way: new integrations should target /unified-ai/v1/*, and my series should be read with that footnote attached. If I started the series today, Part 2 would look different.

This doesn’t invalidate the architectural argument, and I’d argue it strengthens it. The reason the series routed the standard OpenAI SDK through APIM was to keep every call on a governed path. The Unified AI API is that same principle with a better front door: one endpoint, every provider, every pattern, all governed. The lesson survived the release; only the URL changed. Citadel is evolving fast, and this is what evolving looks like from the inside.

Contract-driven everything

The second big theme in citadel-v1 is contracts, and if you read my registry post about the AI Publish Contract, this will feel familiar in the best way.

The accelerator now ships a Citadel Access Contract package: declarative, version-controlled .bicepparam files that onboard an AI use case end-to-end. One contract deployment creates the APIM product (with naming like LLM-Healthcare-PatientAssistant-DEV), the subscription with its key, optional Key Vault secret storage, and optionally an APIM connection for Microsoft Foundry agents. I described this as a pattern worth building in the registry post, and the access contract is now live while the publish contract remains upcoming in the current release.

Alongside it sits a backend onboarding contract (llmBackendConfig) for declaratively registering LLM backends, and the whole thing is versioned through a release.json manifest at the repository root. That manifest is worth a moment of appreciation: instead of one monolithic version number, it tracks independent, component-scoped versions for the routing logic, the backend contract shape, the access contract shape, and the usage ingestion pipeline. A change to routing doesn’t force a re-version of contracts that didn’t change. That’s a small design decision that signals the project expects to be operated, not just deployed once.

The parallel to the AI Publish Contract from my registry post is direct. Both encode the same conviction: onboarding an AI workload should be a reviewed, versioned artifact in a repository, not a sequence of portal clicks someone half-remembers. The access contract governs how a workload reaches the gateway. The publish contract governs registration and description. A mature platform wants both.

Multi-provider routing, briefly

The gateway is no longer an Azure OpenAI front door with ambitions. AWS Bedrock, Google Gemini, and Anthropic Claude are first-class citizens, each available through OpenAI-compatible access, provider-native access, or both.

The design that makes this work without chaos is a fragment-based routing architecture, and one detail from the onboarding guide shows how much operational scar tissue is encoded in it. Every API type declares its own compatible pool types, and the Universal LLM API restricts pool selection to OpenAI-compatible pools before backend selection runs. Why? Because if the same model ID is registered against both a native Bedrock pool and an OpenAI-compatible one, a naive router could send an unrewritten OpenAI-shaped path to the native provider, which answers with something as friendly as com.amazon.coral.service#UnknownOperationException. The guide documents the failure mode by name. Someone hit that error so you don’t have to, which is exactly what a good accelerator encodes.

For a platform team, the practical consequence is real: model choice becomes a routing decision instead of an architecture decision. Adding Claude or Gemini to an estate governed by the hub doesn’t create a second governance perimeter. It adds a backend behind the one you already operate.

What I’d do differently starting today

Distilling this into advice for anyone deploying now:

Target the Unified AI API from day one. Start at /unified-ai/v1/* with the standard OpenAI SDK. You get the same governed path my series argued for, plus provider reach and a native-access upgrade path you’ll eventually want.

  • Adopt the access contract instead of hand-rolling onboarding. The .bicepparam contract per use case gives you reviewable, repeatable onboarding with product, subscription, and secrets in one deployment. I built a weaker version of this by hand during the series; you don’t have to.
  • Pin your contract versions consciously. release.json gives you independent version tracks. Treat contract shape changes as reviewable events in your own repo, the same way you’d treat an API schema change.
  • Look at the PII blocking mode. The PII framework now supports managed identity authentication to the Language Services, regex pre-processing before NLP detection, and a strict mode that rejects requests containing PII with a 400 instead of masking. For regulated industries, that hard-fail option changes the compliance conversation: some data should never reach the model, masked or not.

What’s next

The obvious follow-up experiment: migrating the weather agent from the legacy /openai/deployments/... path to the Unified AI API, documenting whatever breaks along the way. If the routing architecture delivers on its promise, that migration should be a base-URL change. If it isn’t, that’s a post worth writing too.

The accelerator that started this series as a useful pattern is now the reference implementation of a named layer in a published blueprint, with contracts, multi-provider routing, and a defined seam toward agent-runtime governance. Preview or not, the direction is clear, and it’s the direction the series has been arguing for all along: one governed front door, everything registered, nothing invisible.

If you’ve deployed citadel-v1 or migrated from the earlier iteration, I’d like to hear what surprised you.

Azure Data Patterns for Integration Architects

In the Azure PaaS map post, the data layer got one paragraph and a rule: pick by access pattern, not by which service feels modern. That rule holds. But it’s also where most write-ups stop: SQL for relational integrity, Cosmos DB for scale, and Redis in front; for an integration architect, that’s the least interesting part of the story.

The interesting part is what the data layer has to do that’s specific to integration. Messages arrive twice. Workflows run for hours and need somewhere to keep their state. A write to your database and a publish to a queue have to succeed or fail together. So this post skips the service comparison and covers the patterns instead. Azure data patterns for integration are less about which store you pick and more about how you use it.

Why integration data is different

A typical application owns its data. It writes, it reads, it controls the whole path. Integration doesn’t work that way. Instead, integration sits between systems it doesn’t own, reacting to events it didn’t originate, and it has to stay correct when those systems misbehave.

That changes what the data layer is for. It’s no longer just persistence. It becomes the place where you enforce correctness that the messaging layer can’t guarantee on its own. Three patterns come up again and again. Let’s take them in turn.

Pattern 1: Idempotency stores

Here’s the problem. At-least-once delivery is the norm for most messaging systems, including Service Bus. So the same message can arrive twice after a retry, a redelivery, or a consumer crash-and-restart. Process it twice, and you’ve charged the card twice or created two orders. That’s not a rare edge case. In a busy integration platform, it’s a Tuesday.

The fix is an idempotency store. Before you process a message, you check whether you’ve seen its ID before. If you have, you skip it. If you haven’t, you record the ID and proceed. As a result, duplicate deliveries become harmless.

The design questions that matter:

  • Where does the key come from? Ideally, the source system supplies a stable business key, an order ID, and a transaction reference. Failing that, a hash of the message content works, though it’s more fragile.
  • Where do you store it? This is a high-frequency, low-latency lookup on a single key. Therefore Cosmos DB or Redis fit well, and a relational table works too if the volume is modest. The access pattern points at the store, exactly as the map post argued.
  • How long do you keep it? Retention has to outlast the longest possible redelivery window. Too short, and a late duplicate slips through. So set a TTL that comfortably exceeds your retry and dead-letter timelines, then expire old keys automatically.

The honest note: idempotency at the store isn’t the same as an idempotent operation. If the downstream side effect isn’t itself safe to repeat, the store only narrows the window; it doesn’t close it. Design the operation to tolerate retries wherever you can.

Pattern 2: The outbox pattern

This one solves the dual-write problem, and the dual-write problem is subtle enough that plenty of teams ship it broken.

Picture a handler that does two things. It writes a record to the database, and it publishes an event to a queue. Both must happen, or neither. But they’re two separate systems, so there’s no shared transaction. Write succeeds, publish fails; now the database and the downstream world disagree. Publish succeeds, write fails; now you’ve announced something that didn’t happen.

The outbox pattern closes the gap. Instead of publishing directly, you write the event into an “outbox” table in the same database transaction as your business record. Because they share one transaction, they commit together or not at all. Then a separate process reads the outbox and publishes the events, marking each one done as it goes.

A few things fall out of this design:

  • The database becomes the source of truth for what should be published. If the publisher crashes mid-run, it restarts and picks up where it left off. Nothing is lost, because nothing left the database until it was safely committed.
  • Publishing becomes at-least-once. The publisher might send an event, crash before marking it done, and send it again on restart. So the consumer on the other end needs, you guessed it, an idempotency store. The two patterns work together.
  • A change-feed makes it cleaner. Cosmos DB’s change feed, or a similar mechanism, lets the publisher tail committed changes rather than poll a table. That reduces latency and load, though a simple polling publisher is perfectly fine to start.

The trade-off is honest latency. The outbox adds a hop between commit and publish. For most integration workloads that’s a few seconds at most, and well worth it for the correctness guarantee. But if you need genuinely instant propagation, the outbox isn’t your pattern.

Pattern 3: State for long-running workflows

Synchronous request-response keeps its state in memory for the length of a call. Integration workflows don’t have that luxury. A process can span minutes, hours, or days while waiting for approval, a batch window, or an external callback. That state has to live somewhere durable, because the compute running it will scale, restart, and move underneath it.

So where does workflow state go? It depends on who’s orchestrating.

  • Logic Apps and Durable Functions manage their own state: Both persist workflow state for you; that’s a large part of why they exist. Durable Functions keeps it in a storage backend; Standard Logic Apps keeps it in its own runtime store. In these cases, you rarely touch the state directly, but you should know it’s there and know that it’s what makes the workflow survive a restart.
  • Hand-rolled orchestration needs an explicit store: When you’re coordinating steps in your own code rather than a workflow engine, you own the state. A document store like Cosmos DB fits well here: one document per workflow instance, updated as the process advances through its steps. The flexible schema helps, because a workflow’s state shape often evolves as you add steps.
  • Correlation is the piece people forget: Long-running workflows wait for things to come back, and when a callback arrives, you have to match it to the right in-flight instance. That means a correlation ID, stored with the instance and carried on every outbound call. Without it, you have durable state you can’t reconnect to the event that needs it.

Where these patterns are the wrong answer

Consistent with the rest of the series, the honesty section. Patterns solve problems, and applying them where the problem doesn’t exist adds cost.

  • Skip the idempotency store when the operation is naturally idempotent: Setting a status to “shipped” twice changes nothing. If every side effect is already safe to repeat, a dedup store is machinery you don’t need.
  • Skip the outbox when you don’t dual-write: If a handler only writes to the database, or only publishes, there’s no gap to close. The outbox earns its keep specifically when one commit must produce one publish.
  • Skip explicit state stores when a workflow engine already owns the state: Standing up your own Cosmos-backed state store next to Durable Functions duplicates what the runtime already gives you. Reach for the explicit store only when you’re orchestrating by hand.

The shape of it

For an integration architect, the data layer isn’t mainly a choice between SQL and Cosmos. That choice matters, but the access pattern usually makes it for you. The real work is the patterns that keep an integration platform correct when systems it doesn’t control misbehave. So an idempotency store absorbs duplicate deliveries. An outbox makes both a write and a publish succeed. A durable state store lets a workflow outlive the compute running it. Get those right, and the underlying store SQL, Cosmos, and Redis become implementation details rather than the headline.

Want the layer this sits inside? The Azure PaaS map puts data in context against compute, integration, and governance, and walks the five-question framework across all of them. And the messaging and orchestration post covers the delivery guarantees these patterns lean on.

Token Economics in Practice: What the Citadel Cost Attribution Policy Actually Meters

The FinOps Foundation published a piece called Token Economics: The Atomic Unit of AI Value, and it’s one of the better attempts I’ve seen at giving AI cost management a real vocabulary. Tokens as the atomic unit of cost, goodput instead of raw throughput, and a warning that the token meter is increasingly hidden inside SaaS subscriptions you don’t control.

Most writing on this topic stays theoretical. I have something to test it against. The Citadel Platform series on this blog built cost attribution and semantic caching into a real APIM gateway, running in Sweden Central, metering a real agent. So instead of summarizing the FinOps article, this post uses it as a checklist. Where does the Citadel implementation already deliver on token economics, and where does it fall short?

The honest answer: it holds up well on attribution and caching, and it has clear gaps on goodput, yield, and routing. Let’s go through it.

Three token economics ideas worth carrying forward

I’ll paraphrase the three concepts I’m testing against, and you should read the original for the full argument.

Goodput, not throughput. Raw token volume tells you what you spent, not what you got. Goodput asks how many of those tokens produced useful output within acceptable latency. A retry storm and a productive session can burn the same token count.

The cost stack extends beyond the token. Tokens are the atomic unit, but the bill includes orchestration overhead, retries, tool-calling scaffolding, and increasingly, SaaS subscriptions that embed token consumption behind a flat price. When the meter sits inside someone else’s product, you lose the visibility FinOps depends on.

Engineering levers matter more than procurement levers. Model routing, semantic caching, and compressing tool-calling overhead move the cost curve more than negotiating a discount does. The FinOps article cites Cloudflare’s Code Mode work, which cut MCP tool-schema token overhead dramatically by changing how tools present themselves to the model.

Now let’s hold the Citadel Hub against those three ideas.

What the Citadel Hub already meters

Every call the weather agent makes flows through apim-wpvlimv4ngkns, and two of the five governance policies from earlier in the series do the token economics work.

The cost attribution policy emits token metrics per call, dimensioned by subscription and agent:

xml

<azure-openai-emit-token-metric namespace="citadel">
<dimension name="Subscription ID" />
<dimension name="Agent ID" value="@(context.Request.Headers.GetValueOrDefault("X-Agent-Id", "unknown"))" />
<dimension name="API ID" />
</azure-openai-emit-token-metric>

That gives us prompt tokens, completion tokens, and total tokens per agent, per subscription, queryable in Application Insights. When someone asks what the weather agent cost last week, the answer is a query, not an estimate.

The semantic caching policy sits in front of the model and short-circuits repeat questions:

xml

<azure-openai-semantic-cache-lookup
score-threshold="0.85"
embeddings-backend-id="embeddings-backend"
embeddings-backend-auth="system-assigned">
<vary-by>@(context.Request.Headers.GetValueOrDefault("X-Agent-Id", "unknown"))</vary-by>
</azure-openai-semantic-cache-lookup>

A cache hit costs an embedding call instead of a full completion. For an agent that answers weather questions, where “what’s the weather in Amsterdam” arrives in twenty phrasings, that’s not a rounding error.

In FinOps for AI terms, the first policy lives in the Understand Usage and Cost domain, and the second in Optimize Usage and Cost. So far, the framework and the implementation agree.

Where the mapping holds up

Two places, and one of them matters more than I expected before reading the article.

Semantic caching is a named lever. The FinOps article lists it explicitly as an engineering-side optimization, and the Citadel implementation has it running in production policy XML, not on a roadmap slide. Score threshold tuning is real work (0.85 took iterations, and I documented the false-positive risk in the original policy deep dive), but the lever exists, and it’s been pulled.

The gateway is the anti-aggregator, and the FinOps article’s sharpest warning is that token consumption is being hidden in SaaS subscriptions, where a flat monthly price hides a metered reality beneath the surface. The hub-and-spoke model is the architectural inverse of that problem. Nothing reaches a model without crossing the gateway, so nothing consumes tokens invisibly. The whole point of Part 2 in the Citadel series was to refuse the path that bypasses the meter when the Agent Service SDK tries to call the model directly.

I’d go one step further than the FinOps article does. Centralized metering isn’t just a FinOps convenience. It’s the same choke point that enforces content safety and the kill switch. Cost visibility and governance aren’t two systems in this architecture, they’re one policy pipeline.

Where Citadel’s token economics fall short

This is the useful part, because the gaps are specific.

  • No goodput tracking: The Hub knows how many tokens the agent consumed. It does not know how many of them were worth consuming. Time-to-first-token and tokens-per-second aren’t captured as dimensions, and nothing distinguishes a completion the user acted on from one that got regenerated three times. By the article’s standard, Citadel measures throughput and calls it a day.
  • No token yield rate: Closely related, but distinct. Yield asks: cost per successful outcome, not per call. The weather agent writes every conversation to Cosmos DB (Part 3 of the series), so the raw material for outcome tagging exists. Nothing joins it to the token metrics yet. That’s a gap in instrumentation, not in data.
  • No model routing: Every query hits the same deployment, whether it’s “weather in Ede” or a multi-step tool-calling chain. The article’s Pareto framing (bulk tokens, mid-tier tokens, premium low-latency tokens, reasoning tokens, drawn from SemiAnalysis’s InferenceX benchmarking) implies a cascade: cheap model first, escalate on need. APIM can express this with backend pools and routing policy. Citadel doesn’t, yet.
  • Tool-schema overhead is unmeasured: Every tool-calling request carries the Open-Meteo tool definition in the payload, on every single call. One tool, so the overhead is small. But the Cloudflare finding the article cites is a warning about what happens at ten or twenty tools, and I have no metric today that would even show me the problem growing.

Why this bites harder on agentic workloads

There’s a compounding effect the article touches on that I can back with a documented example. Orchestration overhead isn’t a fixed tax, it multiplies through agent chains.

In the Logic Apps Agent Loop series, I found that sequential agents don’t pass plain strings between each other. Each agent action returns a structured JSON messages array, and you need a Compose action to bridge it into the next agent. Every one of those bridged payloads is tokens. Single-agent token math is linear. Multi-agent token math is not, and that’s where token economics stops being a dashboard exercise. If your metering only captures totals per call, the orchestration overhead hides inside numbers that look individually reasonable.

What I’d add to the Citadel Hub next

In order of effort against payoff:

  • Outcome tagging first: The conversations container already holds every run. Adding a resolution field (answered, retried, abandoned) and joining it against the token metrics in Application Insights gets me a real token yield rate with no new infrastructure. This is the cheapest gap to close and the one that changes the conversation from “what did we spend” to “what did we get.”
  • Latency dimensions second: Emitting time-to-first-token and total duration alongside the existing token dimensions turns the same App Insights workspace into a goodput dashboard. APIM sees the timing already, it just doesn’t emit it.
  • A routing experiment third: The weather agent is a good candidate for a two-tier cascade precisely because it’s boring. Simple lookups go to a small model, tool-calling chains escalate. If the cascade breaks the agent, it breaks it cheaply, and I’ll write up whatever goes wrong.

Tool-schema compression stays on the watch list rather than the to-do list. With one tool, measuring it first beats optimizing it blind.

Pitfalls

Adopting the vocabulary without the substance is the most common trap. It’s easy to say ‘we do token economics’ because a dashboard shows token counts. Raw volume without yield or goodput is accounting, not economics. The article’s framework is only useful if the uncomfortable metrics come with it.

Treating flat-price AI tools as flat costs is the second trap. When teams around you adopt AI SaaS tooling, those subscriptions consume tokens on someone’s meter. Budgeting them as fixed line items repeats the exact mistake the article warns about, one procurement layer up.

Optimizing the cache before understanding the traffic is the last one. A semantic cache with an aggressive threshold saves tokens and quietly serves wrong answers. Tune against logged real queries, never against the token savings number alone. I learned this at 0.85, and the number that’s right for a weather agent is wrong for an agent where two similar-sounding questions need different answers.

Closing

The FinOps article gives this space the vocabulary it needs, and the Citadel Platform gives me somewhere to test that vocabulary against running policy XML. The scorecard: attribution and caching, solid. Goodput, yield, and routing: real gaps with concrete next steps.

The bigger takeaway is architectural. Every improvement on that list lands in the same place, the gateway. APIM started this series as a governance layer. It’s ending it as the FinOps instrumentation layer too, and I don’t think that’s a coincidence. The choke point that can say no to a request is the same choke point that can tell you what the request cost.

If you’re metering your own agent platform, I’d like to hear which of these gaps you closed first, and whether the yield numbers surprised you.

Azure Messaging and Orchestration for Integration Architects

In the Azure PaaS map post, the integration layer got a single paragraph. It named four services: Logic Apps, API Management, Service Bus, and Event Grid, and moved on. This post takes Azure messaging and orchestration apart into the decisions underneath that paragraph.

I won’t tour features here. Two of these services already have their own deep series on this blog, so re-covering them would waste your time. Instead, I’ll stay at the decision layer. When do you reach for which? And why do teams so often reach wrong? Those are the questions that actually cost you in production.

Azure messaging and orchestration: two axes decide almost everything

Four services sound like four choices. In practice, though, only two questions matter, and they cut across the whole layer.

  • First: is this messaging or orchestration? Messaging moves events and data between systems. Orchestration coordinates a multi-step process toward an outcome. The two look similar on a whiteboard, but they fail differently, scale differently, and belong to different services. So separate them before anything else.
  • Second: what does failure cost, and what shape is the work? Once you know whether you’re moving messages or coordinating steps, the follow-up question splits the choice further. For messaging, the cost of a lost message decides it. For orchestration, the complexity of the flow and the team who owns it decide it.

Get those two axes clear, and the service almost picks itself. Skip them, and you end up with Event Grid where you needed guarantees, or a Logic App doing work that belonged in code.

Messaging: Service Bus vs Event Grid

Both move things between systems. That’s where the similarity ends.

  • Service Bus is the durable, ordered, transactional backbone. Reach for it when delivery has to be guaranteed. It gives you sessions for ordered processing, dead-lettering for messages that can’t be handled, and transactional handling across multiple operations. Topics and subscriptions add pub/sub without a separate broker. So Service Bus fits business messages: an order, a payment, a claim, where losing one is an incident.
  • Event Grid is a lightweight, high-volume router. It broadcasts events to whoever cares: resource state changes, custom application events, and telemetry. It’s built for throughput and fire-and-forget delivery, not guaranteed processing. Therefore, it fits notifications and reactive triggers, where a missed event is a shrug rather than a page.

Here’s the rule of thumb I give teams new to Azure messaging. If losing a message would be a business incident, it belongs on Service Bus. If losing it would just mean a missed notification, Event Grid is fine.

And often you use both. A common pattern pairs them: Event Grid fans out a notification, and a subscriber drops a durable message onto Service Bus for guaranteed processing. That way you get Event Grid’s reach and Service Bus’s reliability in one flow, each doing the job it’s good at.

The honest note, though, is that this is where teams get burned. Event Grid looks simpler, so teams default to it. Then, weeks later, they discover the workload actually needed ordering or delivery guarantees. Now they’re bolting reliability onto a service that was never designed for it. So decide on the message-loss cost first, before the “which feels easier” instinct takes over.

Orchestration: Logic Apps vs code

Messaging moves things. Orchestration coordinates them. The decision here isn’t about reliability; it’s about complexity and ownership.

  • Logic Apps is the designer-first route. It shines when you need enterprise connectors SAP, IBM MQ, mainframe hosts, the long tail of line-of-business systems without a modern REST API. Standard Logic Apps also closes the old gaps that made it hard in regulated environments: VNet integration, built-in state, per-workflow scaling. So for a workflow that a less code-heavy team will own and maintain, Logic Apps is often the right call even when a Function would be more elegant.
  • Code is the route once complexity climbs. Designer workflows are fast to build and easy to read at first. Past a certain size, though, they get hard to reason about and harder to code-review. In my experience, the practical ceiling sits around a dozen actions with a couple of branches. Beyond that, do one of two things. Either decompose the workflow into smaller ones, or move the logic into a Function where a proper language and real tests take over.

The deciding questions, then, are simple. Who maintains this: a low-code team or engineers? How complex is the flow really? And can you review it a year from now? For the deeper mechanics of building agentic workflows in Logic Apps, I covered that ground in the Logic Apps Agent Loop series so that I won’t repeat it here.

Where API Management fits

Azure API Management (APIM) isn’t messaging or orchestration. Instead, it’s the control point in front of both.

It sits between consumers and whatever does the real work: a Logic App, a Function, an App Service backend. From there, it enforces rate limits, authentication, transformation, and policy-based routing. So when multiple consumers hit a shared set of backend capabilities, APIM lets you change the implementation behind them without breaking anyone, and lets you enforce policy without touching application code.

That’s all I’ll say here, because APIM earns a series of its own. I went deep on it in the APIM for AI workloads series, including how it behaves as an AI gateway. For this layer, treat it as the front door that governs whatever messaging and orchestration sit behind it.

Where each is the wrong answer

Every service here has a failure mode when you reach for it by reflex. So, to keep this honest:

  • Service Bus is wrong for high-volume telemetry: If you’re routing millions of fire-and-forget events and none of them individually matter, Service Bus is expensive overkill. Use Event Grid.
  • Event Grid is wrong for anything needing order: The moment sequence or guaranteed delivery matters, Event Grid stops fitting. Move to Service Bus before the gap bites.
  • Logic Apps is wrong past its complexity ceiling: A workflow with thirty actions and nested branches is a maintenance liability in the designer. Decompose it, or move it to code.
  • Code is wrong for something a citizen developer should own: Not every integration belongs in a repo. If a low-code team can own and maintain a simple connector-driven flow, hand-writing it in a Function just centralizes work that didn’t need to be centralized.

The shape of it

Azure gives you four integration services, but you don’t choose between four things. You answer two questions. Is this messaging or orchestration? And then what does failure cost, and who owns the work? Answer those, and Service Bus, Event Grid, Logic Apps, or a Function each falls out naturally, with APIM governing the front.

In practice, real platforms use several together: APIM, fronting Logic Apps, and Functions, Event Grid fanning out to Service Bus for reliable processing. So the craft of Azure messaging and orchestration isn’t picking a winner. It’s drawing clean boundaries between them.

Want the layer above this one? The Azure PaaS map puts integration in context against compute, data, and governance, and walks the five-question framework for choosing across all of them.