How do you find a performance bottleneck?
By measuring before changing anything, because intuition about where time goes is reliably wrong — and the great majority of wasted optimisation effort comes from skipping that step.
The order that works:
Define the problem numerically. "Slow" is not actionable. Which operation, at which percentile, under what load, and what would be acceptable? Averages hide everything — a p50 of 100ms with a p99 of 8 seconds is a completely different problem from a uniform 400ms.
Reproduce it, ideally in a way you can run repeatedly.
Measure the whole path first. Use tracing or timing to find which stage dominates — network, application, database, external call, serialisation, or the client. Optimising a stage that accounts for 3% of the time is the classic wasted week.
Then profile that stage. A sampling profiler interrupts periodically and records the stack, giving a low-overhead picture of where time is actually spent; a flame graph makes the hot path obvious at a glance.
Form a hypothesis, change one thing, re-measure. Multiple simultaneous changes make attribution impossible.
The usual culprits, roughly in order of frequency:
Database access — missing indexes, N+1 query patterns, unnecessary columns, queries inside loops, and connection pool exhaustion. This is where most application slowness actually lives.
Doing work repeatedly that could be done once or cached.
Serial network calls that could run concurrently — latency adds up brutally.
Serialisation and data transfer volume, including sending far more data than the client uses.
Lock contention and blocking calls in an async context.
Memory pressure, causing excessive garbage collection or swapping.
Algorithmic complexity, which matters enormously when it bites and is less often the cause than programmers expect.
What to avoid: optimising without a target; micro-optimisations in cold code; measuring in a development environment with tiny data; and profiling with a debugger attached, which distorts everything.
Keep the measurement in place afterwards, so a regression is detected rather than rediscovered.