‹ All study guides

SQS vs SNS vs EventBridge on AWS

Updated 12 October 2026 · 8 min read

Two parts of a system need to talk without being wired directly to each other. An order service should not fail because billing is down for maintenance, and it should not need a code change every time a new team wants to hear about orders. AWS offers three managed services for this, and they look interchangeable at first glance: Amazon SQS, Amazon SNS and Amazon EventBridge. They are not. Each one makes a different promise about who receives a message, when, and how many times.

Exam questions about decoupling are almost always a choice between these three (with Kinesis Data Streams as the occasional fourth option). The scenario tells you how many consumers there are, whether they must work at their own pace, whether order matters and whether events need to be routed by content. Match those clues to the right service and the question is done.

SQS: a queue that holds work

An SQS queue stores messages until a consumer asks for them, processes them and deletes them. Consumers pull; nothing is pushed to them. That is why SQS is the answer whenever a consumer might be slow, offline or scaled to zero: messages simply wait. Retention is 4 days by default and can be set from 1 minute up to 14 days.

Each message is meant to be handled by one consumer. Add more workers and they share the backlog, which makes a queue the standard way to absorb spikes in front of an Auto Scaling group or a Lambda function. It is not a way to give three different systems their own copy of every message.

Visibility timeout

When a consumer receives a message, SQS does not delete it. It hides it for the visibility timeout, 30 seconds by default and up to 12 hours. If the consumer deletes the message in time, the work is done. If it crashes, or simply takes longer than the timeout, the message reappears and another consumer picks it up. That is how SQS avoids losing work, and it is also the classic cause of duplicate processing: a job that used to take 20 seconds now takes 45, the default timeout expires mid-job, and a second worker repeats it. The fix is to set the timeout above the longest processing time, or have the worker extend it with ChangeMessageVisibility while it runs.

Standard and FIFO queues

A standard queue gives very high throughput, at-least-once delivery and best-effort ordering. Occasional duplicates and out-of-order messages are possible, so consumers should be idempotent.

A FIFO queue (its name must end in .fifo) delivers messages in order and removes duplicates sent within a 5-minute deduplication window, using either a deduplication ID you supply or a hash of the body. Ordering is per message group ID: messages in one group are handed out strictly one after another, while different groups are processed in parallel. Use a customer or account ID as the group ID and you get order per customer with plenty of parallelism. Use one constant group ID and you have built a single-threaded queue. Without high throughput mode, a FIFO queue handles 300 API calls per second per action, or 3,000 messages per second with batches of 10; high throughput mode raises that considerably.

Dead-letter queues

Some messages will never succeed, such as a malformed payload that crashes the consumer every time. Without help, that message comes back after every visibility timeout forever. A dead-letter queue (DLQ) fixes this: the source queue's redrive policy sets maxReceiveCount, and once a message has been received that many times without being deleted, SQS moves it to the DLQ. You can inspect it there, alarm on the DLQ's depth, and after fixing the bug use redrive to source to move the messages back. The DLQ must be in the same account and Region and be the same type as its source (FIFO with FIFO). For a standard queue a message keeps its original enqueue time when it moves, so give the DLQ a longer retention period than the source.

SNS: a topic that pushes copies

An SNS topic receives a message and immediately pushes a copy to every subscription: SQS queues, Lambda functions, HTTPS endpoints, email, SMS and mobile push. It stores nothing for later. If a subscriber is an HTTPS endpoint that is down for hours, SNS retries for a while and then the message is gone, unless that subscription has its own DLQ.

SNS is about one-to-many. The publisher sends once and does not know or care how many subscribers exist. Subscription filter policies let each subscriber receive only the messages whose attributes (or body) match.

Fan-out: SNS plus SQS

The most tested pattern in this area combines the two: publish to an SNS topic and subscribe one SQS queue per consuming system. Every system gets its own durable copy, each consumes at its own pace, and a consumer that is offline for maintenance finds its messages waiting when it returns. Adding a consumer means adding a queue and a subscription, with no change to the publisher. The queue's access policy must allow the topic to send to it.

SNS FIFO topics keep order and deduplicate, using the same message group and deduplication IDs. They deliver only to SQS queues (FIFO queues if you need the order preserved all the way through), not to email, SMS or HTTPS. SNS FIFO topic to one SQS FIFO queue per consumer is the answer when every consumer needs every event in order per key.

EventBridge: a router for events

An EventBridge event bus receives JSON events and routes them with rules. Each rule has an event pattern that can match on any field in the event (source, detail type, values deep in the payload, prefixes, numeric ranges) and sends matches to targets such as Lambda, SQS, SNS, Step Functions, Systems Manager Automation, API destinations or another event bus, including one in another account. On the rules-based bus a rule has at most five targets, so wide fan-out is done with more rules or an SNS topic as a target.

Three things set EventBridge apart:

  • AWS services already publish to it. EC2 state changes, CodePipeline failures, AWS Health events, GuardDuty findings and many others arrive on the default bus. Automation that reacts to something happening in your account almost always starts with an EventBridge rule.
  • Content-based routing across teams and accounts. Producers publish to one bus; consumers own their own rules. Onboarding a team is a new rule, not a producer change.
  • Archive and replay. An archive keeps matching events for a period you choose (indefinitely by default), and a replay sends a time range of them back to the same bus. That is how you reprocess events after fixing a consumer bug without asking the producer to resend.

EventBridge is not a queue. Like SNS, it pushes and keeps nothing for a consumer that is not ready, so put an SQS queue between a rule and any consumer that needs buffering. Its ordering is not guaranteed either. For schedules, use EventBridge Scheduler, which AWS now recommends over scheduled rules.

Where Kinesis fits

Kinesis Data Streams is a log, not a queue. Records are kept for 24 hours by default and up to 365 days, ordered within a shard by partition key, and every consumer reads the whole stream at its own position. Reading does not delete anything, so several applications can read the same data in order and any of them can rewind and reprocess. Choose it for high-volume streaming data, ordered per key, with several independent readers or a need to replay. The cost is managing shards (or paying for on-demand capacity).

How to choose

  1. One consumer group sharing work, possibly slow or offline: SQS.
  2. Strict order per key with one consumer group: SQS FIFO with the key as message group ID.
  3. Several independent systems each needing every message, at their own pace: SNS topic with an SQS queue per system.
  4. The same, but order and deduplication matter: SNS FIFO topic with SQS FIFO queues.
  5. Routing on event content, reacting to AWS service events, cross-account distribution, or replaying past events: EventBridge.
  6. A continuous high-volume stream that several apps read in order and may rewind: Kinesis Data Streams.

Common exam traps

  • Several consumers polling one queue. They split the messages; none of them sees all of them. If each system needs every message, the answer is fan-out.
  • SNS straight to an HTTPS endpoint for a consumer that goes offline. SNS does not hold messages for days. Put a queue in between.
  • Shortening the visibility timeout to fix duplicates. It makes duplicates worse. The timeout must be longer than processing time.
  • Raising retention to deal with poison messages. Retention only keeps the bad message circulating longer. A DLQ with maxReceiveCount is the fix.
  • Expecting a standard queue to keep order or a FIFO queue to scale with one message group.
  • Rebuilding archive and replay by hand. If events went through EventBridge and an archive exists, a replay is the low-effort answer.

A worked example

An order service emits an event for each order. Inventory, billing and analytics must each receive every event. Billing is taken offline for several hours of maintenance each month and must catch up afterwards. The order service should publish each event once.

A single queue fails because the three systems would split the events between them. SNS with HTTPS subscriptions fails because billing would miss what was sent while it was down. EventBridge alone has the same problem. The design that meets everything is an SNS topic with three SQS queues subscribed, one per system: each gets its own copy, billing's queue fills up during maintenance and drains afterwards, and a DLQ on each queue catches anything a consumer cannot process. If the requirement added "route only high-value orders to a fraud team" or "replay last week's events for a new team", EventBridge rules feeding the queues would become the better front door.

Practice questions

Further reading

Practice questions