What is a message queue and when do you need one?
A message queue sits between a producer of work and a consumer of it, holding messages until they can be processed. It converts a synchronous call into an asynchronous one.
What it gives you:
Decoupling. The producer does not need the consumer to be available, or even to exist yet. It publishes and moves on. Services can be deployed, restarted and scaled independently.
Load levelling (buffering). This is often the main reason. Traffic arrives in spikes; processing capacity is fixed. A queue absorbs the spike and lets consumers work through it at a sustainable rate, rather than the system collapsing under load. A sudden burst becomes a longer queue, not an outage.
Responsiveness. A web request that triggers slow work — sending email, generating a PDF, processing video, calling a slow third party — can enqueue the work and return immediately.
Retries and failure isolation. A failed message can be retried with backoff, and one persistently failing message can be moved to a dead letter queue for inspection rather than blocking everything behind it.
Independent scaling. Add consumers to process faster without changing the producer.
When you need one: background jobs, work that can tolerate delay, spiky traffic, fan-out to multiple consumers, integration with unreliable external services, and anything where a user should not wait.
When you do not. A queue adds real operational complexity — another component to run, monitor and reason about. If work is fast and must be synchronous, a direct call is simpler and easier to debug.
Things that catch people out:
Delivery semantics. Most systems provide at-least-once delivery, meaning duplicates are possible. Consumers should therefore be idempotent — processing the same message twice must be safe. Exactly-once is much harder than it sounds.
Ordering is not guaranteed by default in many systems.
Queues hide problems — a growing backlog needs alerting, or failure is silent.