Messages, Events or Streams: Choosing Azure Messaging by What You Send

I’ve worked in the integration space for a long time, and messaging has been at the center of it for most of that time. Queues, topics, publish/subscribe, event routing, streaming: the patterns stay, while the products and their names keep changing. The first question I ask in a design review hasn’t changed either. What is actually on the wire, and what does the sender expect to happen next? When teams ask me to compare Azure messaging services, I answer with that question before I name a single product.

In July I wrote Azure Messaging and Orchestration for Integration Architects as part of the Azure PaaS for Integration Architects series. That post drew one line through messaging: “If losing a message would be a business incident, it belongs on Service Bus. If losing it would just mean a missed notification, Event Grid is fine.” I still stand by it. But it deliberately covered only two of the Azure messaging services, and there are five options for moving data between systems: Storage Queues, Service Bus queues, Service Bus topics, Event Grid and Event Hubs.

Recently I saw a LinkedIn post with an infographic listing all five, and it’s a good moment for me to paint a clearer picture. Cheat sheets like that one usually sort by product. After years of fixing integrations, I look at the payload first. This post makes that distinction explicit, gives a clear rule for when to use what, and backs it with one order flow that uses all five services. You can deploy it yourself from github.com/steefjan1/messaging-choices-azure.

Messages, events and streams

Microsoft Learn’s comparison of the messaging services puts the distinction into two definitions:

  • An event is “a lightweight notification of a condition or state change. The publisher has no expectation about how the event is handled.”
  • A message is “raw data produced by a service to be consumed or stored elsewhere.” And: “A contract exists between publisher and consumer.”

Events split once more. A discrete event reports a state change you can act on. An event series is part of a time-ordered stream that you analyze. One reading from a delivery van tells you nothing; ten thousand of them tell you which van is overheating.

That gives three kinds of payload, and each maps to a family of Azure messaging services. Throughput, cost and the technology stack all matter, but they come second.

Azure messaging services: when to use what

Here’s the short version I give teams, one rule per option.

UseWhenNot when
Storage QueueYou have background work that must get done eventually, and losing a message costs you a rerun, not a customer. Reports, thumbnails, file jobs.You need dead-lettering, ordering, duplicate detection or transactions. You’d end up building them yourself.
Service Bus queueYou’re sending a command with exactly one owner, and the business depends on it: ShipOrder, IssueRefund, PostInvoice.Nobody would notice the loss, or several teams need the same message.
Service Bus topicA business fact has several parties that each owe it something: OrderPlaced goes to payment, inventory and email, each with its own copy, retries and dead-letter queue.The subscribers owe nothing and just want to know. That’s an event.
Event GridSomething happened, and whoever cares may react: a blob landed, a resource changed, a domain event fired. The publisher doesn’t know or care who listens.The handler must process it or the business breaks. Put a queue behind Event Grid for that.
Event HubsYou ingest a continuous stream (telemetry, clickstream, logs) where the value is in the aggregate and several readers need the same data.You need a retry or an acknowledgment per message. A stream has neither.

Two quick tests settle most debates. Would losing one of these be a business incident? Then it’s Service Bus, as in the July post. Would anyone notice a single missing item? If not, and there are thousands per second, it’s a stream.

Why the rules fall where they do

The table is the short version. The reasons come from how each service behaves when something fails.

Storage Queue or Service Bus

Storage Queue is cheap, it’s HTTP, and it scales to a backlog larger than 80 GB. Learn’s side-by-side comparison is blunt about what you give up: messages up to 64 KB, no ordering guarantee, no duplicate detection, no transactions and no dead-letter queue. To find poison messages, “the application examines the DequeueCount property” and moves them itself. If you use Azure Functions, the host does that for you and fills a -poison queue. That’s a Functions feature, though, not a Storage one.

Service Bus is built for work the business depends on. The broker owns a dead-letter queue per entity and records a reason on each message. You get sessions for FIFO per key, duplicate detection on MessageId, and transactions. Standard tier caps messages at 256 KB; Premium goes to 100 MB and gives you isolated capacity. One tier catch: Basic supports queues only, with no topics, sessions or duplicate detection.

Queue or topic comes down to ownership: a command has one owner, a business fact has several. OrderPlaced sounds like an event, which is why people reach for Event Grid. By Learn’s definition it isn’t one. The publisher very much expects someone to take the money. That expectation is a contract, and a contract belongs on Service Bus.

Event Grid routes, it doesn’t queue

Event Grid pushes to Functions, webhooks, Logic Apps and queues, and namespace topics add pull delivery and an MQTT broker. Its delivery behavior tells you what it’s built for. The default retry policy is 30 attempts within 1,440 minutes, backing off from 10 seconds to 12 hours. “Event Grid doesn’t guarantee order for event delivery.” And “by default, Event Grid doesn’t turn on dead-lettering”: an event that runs out of retries is simply dropped unless you configured a storage container for it.

That’s fine for a notification, not for work. So the pattern that holds up in production puts a queue behind Event Grid: the event says “a blob arrived,” and a queue holds the job of processing it, with its own pace, retries and poison handling.

Event Hubs is a log, not a queue

Event Hubs is the one service in the set where reading a message doesn’t remove it. It’s a partitioned, append-only log. Consumers track their own position with checkpoints, and “checkpointing is the consumer’s responsibility.” Events stay until the retention period expires: up to 7 days on Standard and 90 on Premium and Dedicated.

Three things follow from that. Ordering holds only within a partition, so you pick a partition key (the van id, the device id) to keep related events in order. Consumer groups give independent readers of the same stream, each with its own checkpoint. And there’s no per-message acknowledgment and no dead-letter queue. A consumer that can’t handle event 4,711 has to decide for itself what to do and then move on. That’s why it’s excellent for telemetry and wrong for order processing, however high the volume.

One order flow, all five services

The companion repo puts this into one e-commerce flow. It runs on a single Azure Functions app (.NET 8 isolated, Flex Consumption) and deploys with azd up to Sweden Central. All connections use the app’s managed identity, and local auth is disabled on Service Bus, Event Hubs and both storage accounts.

Orders: messages on Service Bus

POST /api/orders publishes OrderPlaced to a Service Bus topic with four subscriptions. The fraud-review subscription uses a SQL rule, total >= 1000, so it only sees expensive orders. Rules evaluate application properties, not the body, and that’s why the publisher sets them explicitly:

var message = new ServiceBusMessage(BinaryData.FromObjectAsJson(order, Json))
{
MessageId = order.OrderId, // duplicate detection keys on this
Subject = nameof(OrderPlaced),
};
message.ApplicationProperties["total"] = (double)order.Total; // the SQL rule reads this

The payment handler settles its own messages. It dead-letters an order it can never charge with reason InvalidTotal. If it hits a failure that looks transient, it throws and lets the broker retry, and after three deliveries the broker dead-letters the message with MaxDeliveryCountExceeded. On success, it sends a ShipOrder command to a session-enabled queue with SessionId = customerId.

One detail is easy to miss: sending the command and completing the incoming message are two operations, not one transaction. After a crash between them, the order is redelivered and the command sent again. Duplicate detection on shipments drops the second copy, because the command’s MessageId is derived from the order id. At-least-once delivery means idempotent handlers, and broker features like this take some of that work off your hands.

Product images: an event, then work

A blob upload to product-images raises BlobCreated on an Event Grid system topic. The subscription is set to 30 attempts within 24 hours and has dead-lettering turned on, since it’s off by default. OnImageUploaded doesn’t process the image. It drops an ImageJob on a Storage queue and returns. ImageWorker does the work, and anything that fails three dequeues lands in image-jobs-poison. The uploader knows nothing about any of this, which is the point of an event.

Telemetry: a stream on Event Hubs

POST /api/telemetry simulates delivery vans and sends batches to an event hub with four partitions, keyed by van id. Two functions read the same events through two consumer groups: one computes averages, the other flags overheating engines. Take the alerts function down for an hour and it resumes from its checkpoint, and the aggregator never notices.

What the demo run showed

scripts/demo.ps1 drives every scenario, including a duplicate order, a poison order and a corrupt image, and then peeks at where the failures ended up. On my run against Sweden Central, the output looked like this:

WhereMessageWhat the platform recorded
orders/payment dead-letter queuezero-total orderreason InvalidTotal, description “Order total 0 is not payable.”, delivery count 0
orders/payment dead-letter queuepoison customerreason MaxDeliveryCountExceeded, “could not be consumed after 3 delivery attempts”, delivery count 3
shipments dead-letter queuenoneempty: every paid order shipped
image-jobs-poison Storage queuecorrupt-104734.pngthe job body and a dequeue count of 0

That table is the lesson in one picture, and it’s where Azure messaging services differ most. The Service Bus dead-letter queue belongs to the broker, and every entry tells you why it’s there and how often it was tried. The dead-lettered order also never burned a retry, because the handler knew it could never succeed. The Storage poison queue only exists because the Functions host moved the message there, and it arrives with no reason and a dequeue count reset to 0. The history of three failed attempts is gone. If a Storage queue carries work you’ll have to explain in an incident review, you’ll be rebuilding that history yourself.

The publisher also can’t tell a duplicate from a first send. Both POSTs of the same order came back accepted, and the API logged “published” twice. The broker drops the second copy silently, inside its 10-minute duplicate detection window.

Where this framing is the wrong answer

No rule for picking Azure messaging services survives every context. These are the ones I’ve seen bend it.

  • When the volume argument is real. A few thousand orders per second is still a message workload. Before you move it to Event Hubs and rebuild dead-lettering yourself, look at Service Bus Premium with more messaging units.
  • When the devices need identity or commands back. Real vans don’t write straight to Event Hubs. If you need per-device authentication or cloud-to-device messages, look at IoT Hub or the Event Grid MQTT broker in front of the stream.
  • When you don’t need a broker at all. A synchronous call with a retry policy is simpler than a queue with one producer and one consumer in the same deployment. Add a queue when you need to decouple, not by default.
  • When the event crosses an organizational boundary. Service Bus subscriptions are cheap inside one team. Across teams or companies, Event Grid with CloudEvents and filtered subscriptions is usually the looser and better contract.
  • When cost dominates. Event Hubs Standard bills per throughput unit whether you send anything or not, and Service Bus Standard has a base charge. Run azd down --purge after the demo.

Takeaway

Choosing between Azure messaging services doesn’t start with “which Azure service?” It starts with “what am I sending, and what does the sender expect?” A command goes to a Service Bus queue, a business fact with several owners to a topic, and cheap background work to a Storage queue. A notification without expectations goes to Event Grid, with a queue behind it for the actual work. A stream goes to Event Hubs. As I wrote in July, the craft isn’t picking a winner. Most real systems use several of these, and the interesting design work happens at the seams between them.

The repo, diagrams and demo script are at github.com/steefjan1/messaging-choices-azure.

Sources

Event-Driven Services in the Cloud: Azure Event Grid, AWS Event Bridge, and Google EventArc

I mentioned Azure Event Grid in a scenario with D365FO Business Events in a previous blog post. It is a Platform as a Service (PaaS) capability in Azure or eventing platform or event bus (I see various terms describing the service) allowing you to centrally manage events. In addition, it supports direct event filtering based on event type, prefix, or suffix, so your application will only receive events that are relevant to it.

Whether you want to handle built-in Azure events, such as a file being added to storage, or create your own custom events and event handlers, Event Grid supports both options via the same underlying model. Thus, regardless of the service or use case, intelligent routing and filtering capabilities apply to every event scenario and ensure that your apps focus on core business logic rather than worrying about event routing.

In this blog post, I like to dive into Azure Event Grid and competitive offering on the two other big cloud providers, AWS and Google.

Azure Event Grid

In 2017 Microsoft introduced Azure Event Grid as a fully-managed event routing service and the first of its kind (meaning the public cloud claimed it was the first offering the service). Dan Rosanova, previously Principal Program Manager Lead at Microsoft, now Director Program Management at Confluent, said in an InfoQ news item on the introduction:

Azure Event Grid fills a gap in the current cloud messaging space, not just in Azure but also across all cloud providers. We have services for messaging, queuing, and telemetry, but nothing for comprehensive eventing, particularly for cross-service or cross-cloud scenarios.

Within Azure service supporting Event Grid generates events routed to several event handlers. These handlers support event filtering and reliable delivery, ranging from Azure Functions to webhooks. Furthermore, underhood, the service relies on Service Fabric and thus can scale automatically to handle millions of events per second.

Source: https://docs.microsoft.com/en-us/azure/event-grid/overview

The Event Grid concept revolves around events emitted from a source (publisher), an Azure service, or a third-party source that adheres to the event schema (proprietary schema or the CNCF Cloud Events schema). For example, IoT Hub, Storage, and others are all event publishers in Azure. Following that, the events are sent to a topic in Event Grid, and each topic can have one or more subscribers (event handlers). A topic can be set up with the event publisher, or it can be a custom topic for custom events. Finally, event handlers respond to and process the events. Functions, WebHooks, and Event Hubs are examples of event handlers in Azure.

Azure Event Grid generally became available (GA) in February 2018 and Clemens Vasters, Principal Architect Messaging Services at Microsoft, said:

Event Grid is catching everyone’s attention because it unlocks new architectural possibilities for cloud platforms and applications: it’s the glue that enables information flow between services, and Event Grid allows expanding the capabilities of existing services by extension.

And that’s what also triggered or got the attention of AWS as they released EventBridge in July 2019, labeled as a serverless event bus that allows AWS services, Software-as-a-Service (SaaS), and custom applications to communicate with each other using events.

Since the GA, Azure EventGrid received several updates, including advanced filtering, retry policies, and support for CloudEvents. More details and samples are available on the Microsoft documentation and GitHub. Note that there is also an introductory paper available on Azure Event Grid and GitHub from Clemens.

AWS Eventbridge

You can use EventBridge to build and manage event-driven solutions by centrally controlling event ingestion, delivery, security, authorization, and error handling. Furthermore, you do not have to manage any infrastructure or scaling and only pay for the events that their applications consume, similar to Azure Event Grid. Moreover, the concepts are the same too.

Source: https://aws.amazon.com/eventbridge/

However, Amazon Eventbridge surpasses Azure Event Grid with features (as you can see from the diagram above). It has a schema registry allowing you to discover, create, and manage OpenAPI schemas for events on EventBridge. According to the documentation, you can find schemas for existing AWS services, create and upload custom schemas, or generate a schema based on events located on an event bus. Furthermore, EventBridge enables you to generate and download code bindings for all event schemas to help quickly build applications that use those events.

Next to the schema registry, the service integrates easily with third-party services like Zendesk, Pagerduty, and SignalFx. Amazon has set up an extensive partner program for these integrations. Event Grid supports partner events (still preview) yet only has one with Auth0.  

Another capability Amazon EventBridge offers is event replay and archive –  allowing you to archive events so that you can easily replay them later by starting an event replay. Again, a capability that is not available in Azure Event Grid. Although it is something, you can find in Azure Event Hubs. You can configure the archive capability with the actions menu on the EventBridge Console and set the events’ retention period (ranging from zero days to infinite). Subsequently, you can optionally set a pattern matching filter for which events to archive. Later, when events run through the event bus, you can replay the events by selecting the appropriate archive.

Sample Implementation AWS EventBridge

Since the inception of Event Grid, I followed its evolution and wrote and presented on it. Moreover, I followed its competitive solution on AWS and, next to writing about it on InfoQ, built a simple demo around it using .NET in combination with AWS EventBridge. Below you will find a diagram of the demo I created.

From .NET code, I send an event to a custom event bus containing a rule to send the event to a destination, an Amazon Simple Queue. Subsequently, an AWS Lambda function can poll the queue and receive the message – below shows the steps until the SQS queue.

You can find a live demo on YouTube with demoing the above (minute 19). Furthermore, you can look at other samples like in the AWS documentation or on GitHub.

Google Eventarc

With Azure and AWS offering a service to centrally manage events, Google followed in October 2020 with Eventarc to provide customers with a service to connect Cloud Run services with events from various sources, adhering to the CloudEvents standard. It became generally available in January 2021.

Eventarc’s underlying delivery mechanism is Pub/Sub, which includes topics and subscriptions similar to previously discussed Event Grid and EventBridge. Event sources create events and publish them in any format on the Pub/Sub topic. The events are then delivered to the Google Run sinks. For applications running on Cloud Run, you can use Eventarc to use a Cloud Storage event (via Cloud Audit Logs) to trigger a data processing pipeline or an event from custom sources (publishing to Cloud Pub/Sub) to signal between microservices.

Source: https://codelabs.developers.google.com/codelabs/cloud-run-events#1

The diagram above shows what Google hopes to achieve with Eventarc. Currently, you can Cloud Run Service as a destination, and recently Cloud Run for Anthos has been added. Additionally, you can leverage a UI through the Google Cloud console allowing you to view, edit, and delete EventArc triggers. Lastly, you can find more details and samples on GitHub.

CloudEvents Schema

Before I end the blog post with some conclusions, I like to discuss the CloudEvent schema. CloudEvents is an open-source specification for consistently describing event data to make event declaration and delivery easier across services, platforms, and beyond. The Cloud Native Computing Foundation (CNCF) is the driving force behind the specification, which reached the version 1.0 milestone in October 2019.

Clemens Vasters, Principal Architect Messaging Services at Microsoft, stated in an InfoQ news item on CloudEvents:

The goal was to provide an industry definition and open framework for what an “event” is, what its minimal semantic elements are, and how events are encoded for transfer and how they are transferred and do so using the major encodings and application protocols in use today rather than inventing new ones.

Earlier I mentioned that Azure Event Grid has its own proprietary schema and supports CloudEvent schema. The differences are shown below:

Note that Azure Event Grid and Google Eventarc support the CloudEvent schema; however, AWS EventBridge does not, leading to customization.

Conclusion

From this blog post, you can probably conclude that AWS with Eventbridge delivers the most complete event bus or eventing platform in the cloud than Event Grid and Eventarc. If I rank each, AWS comes first, Azure second, and Eventarc third based on features and maturity. The service overlap in concepts, yet implementation, support, and features differ dramatically. Interestingly, they all support changes in their respective storage service. Azure Event Grid brings support for events like when blobs are created, and EventBridge supports S3 notifications and Eventarc triggers for Cloud storage. You can think of various scenarios regarding storage and events, for instance, the pipe and filters pattern implementation discussed in my first blog post.