What does being on-call actually involve?
Being responsible for responding to production incidents outside working hours — and it is one of the more consequential aspects of software work for people's lives, and one of the least well designed in many organisations.
The arrangement. A rota assigns responsibility for a period — typically a week. During that period, alerts route to the on-call engineer, who must acknowledge within a defined time, assess the situation, and either resolve it or escalate.
What it actually costs. Not simply the incidents. The constraint is the availability requirement: remaining within reach of a working connection, able to respond within minutes, for the whole period. That restricts sleep, travel, social arrangements and childcare regardless of whether anything happens — and the burden is frequently accounted for only in terms of hours worked, which substantially understates it.
Sleep disruption is the specific harm. Being woken repeatedly has documented health effects, and a single night's alert affects performance for days.
What distinguishes a sustainable rota from a damaging one:
Enough people. A rota of three is punishing; six or more is manageable.
Compensation — payment for being on call, not only for incidents, and time off in lieu after disrupted nights.
Actionable alerts only. This is the single biggest factor. An alert that does not require immediate human action should not page anyone. Alerting on causes rather than symptoms produces noise, and alert fatigue is a genuine safety problem — once people stop trusting alerts, real ones are missed.
Runbooks for known scenarios, so a tired person is following a procedure rather than reasoning from scratch.
Authority to act, including to wake someone else. An on-call engineer without permission to escalate is in an impossible position.
Clear escalation paths, with defined secondary cover.
Follow-up. Incidents producing action items that are actually completed, so the same page does not recur indefinitely.
You build it, you run it is the argument that teams operating their own services build more operable software — which is well supported, and only works if the team has the authority to change the system.
Alert volume is a leading indicator of attrition, and organisations that track nothing else should track that.