How do platforms actually roll out changes?
Gradually, to small randomised slices of users, measured against a control group — which is why your feed can differ from a friend's on the same day, and why "they changed the algorithm" is frequently both true and unverifiable.
The core method: A/B testing. Users are randomly assigned to a control group, which sees the existing behaviour, and one or more treatment groups, which see a variant. Metrics are compared. If the variant performs better against the chosen measure, it moves forward.
Why randomisation matters. Without it, you cannot tell whether a difference was caused by the change or by who happened to be in each group. Randomisation is what makes the comparison causal rather than correlational.
The staged rollout:
Internal testing — employees only, sometimes called dogfooding.
A small experiment — frequently 1% or less of users, enough to detect large problems.
Progressive expansion if metrics hold, to larger percentages.
Full rollout, sometimes region by region.
Feature flags make this possible: the code ships to everyone, but behaviour is switched on per user by configuration. It also allows an instant rollback without a new release, which is why platforms can undo a change within minutes.
What this explains:
Why your app differs from someone else's on the same version — you are in different buckets.
Why features appear and disappear. A test ended, or you were rolled back.
Why support cannot help. Front-line staff frequently do not know which experiments a given account is in.
Why creators' reports are unreliable evidence. Any individual's experience is a sample of one within an unknown experimental condition.
What gets measured. Typically engagement, retention and revenue, alongside guardrail metrics intended to catch harms — reports, blocks, sessions abandoned. What is chosen as the success metric determines what the platform becomes, and that choice is rarely disclosed.
The ethical dimension. Users are enrolled without specific consent, under terms of service. Some past experiments — notably emotional contagion research — caused significant controversy, and platforms now run internal review processes, with limited external visibility.
Holdback groups are sometimes kept on old behaviour long-term to measure cumulative effects.