The question that decides everything else
Where the model runs — on the device or in the cloud — determines your failure modes, your privacy story and your operating cost.
What actually differs
| Dimension | On-device | Cloud |
|---|---|---|
| Network outage | Checkout keeps working | Checkout stops |
| Per-transaction cost | Fixed hardware cost | Grows with volume |
| Model updates | Rolling, per store | Instant, fleet-wide |
| Data leaving the store | Images stay local | Images uploaded |
| Latency | 30–150 ms | 200–800 ms |
Why latency matters more than accuracy here
A cashier processes a transaction every 5–15 seconds. An extra 500 ms of network latency, repeated on every item, is felt as a slow register — even if the recognition itself is perfect.
When cloud is the right answer
- Very large catalogues that change daily
- You already have reliable, low-latency links per store
- You need centralised model iteration more than you need offline resilience
The pragmatic pattern
Run inference on-device for the latency-critical path, and sync metadata, not frames, back to the cloud. Keep a local queue so an outage becomes delayed upload rather than lost data.