What is the difference between a CPU core, a thread and a GPU?
A core is a physical processing unit; a thread is a stream of instructions a core executes; a GPU is a fundamentally different processor designed for a different shape of problem. Confusing them leads to buying the wrong hardware for the work.
Cores. A modern CPU contains several physical cores, each capable of executing instructions independently. More cores means more genuinely simultaneous work — provided the software can use them, which is the constant caveat. Many tasks are inherently sequential and gain nothing.
Threads and simultaneous multithreading. One physical core can present as two logical processors, interleaving two instruction streams to fill gaps when one stalls waiting for memory. This is not two cores: the typical gain is perhaps 20–30%, not 100%, and it varies by workload.
Performance and efficiency cores. Many modern processors mix fast power-hungry cores with slower efficient ones, with a scheduler assigning work — which is why core counts are no longer comparable between designs.
Clock speed and IPC. Speed alone is meaningless across architectures, since instructions per cycle differ. A slower chip with better IPC outperforms a faster one with worse.
Cache, which frequently matters more than either, because waiting for main memory dominates many workloads.
GPUs, and why they are different in kind. A CPU has a few powerful cores optimised for latency — finishing one complex task quickly, with branch prediction and large caches. A GPU has thousands of simple cores optimised for throughput — performing the same operation on enormous amounts of data simultaneously.
What that suits: graphics, where every pixel undergoes similar maths; and machine learning, which is overwhelmingly matrix multiplication. This is why GPUs became the hardware of AI — not because they were designed for it, but because the shape of the problem matched.
What GPUs are bad at: branching logic, sequential dependencies, and anything where each step depends on the last.
Memory bandwidth is frequently the real constraint on a GPU, not core count.
NPUs are a further specialisation for efficient on-device inference.