What is observability, and how is it different from monitoring?
Monitoring tells you whether the things you anticipated are behaving. Observability is the property of being able to answer questions you did not anticipate, from the data the system already emits.
The distinction concretely. Monitoring is dashboards and alerts on known failure modes: CPU above a threshold, error rate above a threshold, disk filling. You decided in advance what to watch. That works well for failures you have seen before.
Observability addresses the other case — a user reports that checkout is slow for some customers in one region on mobile, and no dashboard exists for that. Can you find out without deploying new code? If yes, the system is observable.
The three commonly cited signals:
Logs — discrete timestamped events. Rich in detail, expensive at volume, and hard to aggregate. Structured logging — emitting key-value fields rather than formatted sentences — is what makes them queryable rather than greppable.
Metrics — numeric measurements aggregated over time. Cheap, efficient, excellent for trends and alerting. Their limitation is cardinality: you cannot attach a user ID to a metric without exploding storage, so metrics tell you that something changed and rarely for whom.
Traces — the path of a single request across services, with timing at each step. In distributed systems this is frequently the only way to locate where latency actually accumulates.
The ideas that matter more than the three-pillar framing:
High cardinality and wide events. The practical route to answering unanticipated questions is emitting events with many attached dimensions — user, version, region, feature flag, device — and being able to slice by any of them arbitrarily.
Correlation IDs threaded through every log and span, so one request can be reconstructed.
Sampling, because full-fidelity tracing at scale is unaffordable, with tail-based sampling keeping the interesting traces.
OpenTelemetry has become the standard instrumentation layer, which matters mainly because it decouples instrumentation from vendor.
Alert on symptoms users feel, not on causes.