The question that decides everything else

Where the model runs — on the device or in the cloud — determines your failure modes, your privacy story and your operating cost.

What actually differs

Dimension On-device Cloud
Network outage Checkout keeps working Checkout stops
Per-transaction cost Fixed hardware cost Grows with volume
Model updates Rolling, per store Instant, fleet-wide
Data leaving the store Images stay local Images uploaded
Latency 30–150 ms 200–800 ms

Why latency matters more than accuracy here

A cashier processes a transaction every 5–15 seconds. An extra 500 ms of network latency, repeated on every item, is felt as a slow register — even if the recognition itself is perfect.

When cloud is the right answer

  • Very large catalogues that change daily
  • You already have reliable, low-latency links per store
  • You need centralised model iteration more than you need offline resilience

The pragmatic pattern

Run inference on-device for the latency-critical path, and sync metadata, not frames, back to the cloud. Keep a local queue so an outage becomes delayed upload rather than lost data.