What is the difference between on-device and cloud AI?
Where the computation happens — and the trade-off is between privacy, latency and offline capability on one side, and capability, scale and updatability on the other.
Cloud AI. Your input is sent to a remote data centre, processed on specialised hardware, and the result returned.
Advantages: access to the largest models, which cannot fit on a phone; consistent capability regardless of device; updated continuously; and no battery or thermal cost to the device.
Costs: your data leaves the device, which is the fundamental privacy consideration; a network connection is required; latency includes a round trip; the provider bears real per-query cost, which is passed on; and availability depends on their service.
On-device AI. The model runs locally, on the phone's or computer's own processing hardware — increasingly a dedicated neural accelerator.
Advantages: data never leaves the device, which is a categorical rather than a policy guarantee; very low latency with no network round trip; works offline; no per-query cost; and it does not depend on a service continuing to exist.
Costs: models must be small enough to fit in memory and run within a power budget, so capability is lower; battery and heat; storage; and updating requires shipping a new model.
How smaller models are made to fit: quantisation, reducing numerical precision so the model uses less memory with modest quality loss; distillation, training a small model to imitate a large one; and pruning.
Hybrid approaches, which is where most consumer products have landed. Simple and privacy-sensitive tasks run locally; complex requests are sent to the cloud, sometimes with the device deciding per request. Some implementations add attested private cloud processing, where the provider designs the system so it cannot retain or inspect the data.
What runs well locally already: speech recognition, keyboard prediction, photograph categorisation, face and object recognition, translation, noise suppression and summarisation of short text.
The direction of travel is toward more local processing as accelerators improve and small models become substantially more capable — which is driven by cost as much as by privacy.