Industrial data

Sizing an IIoT data budget: sampling, deadband and retention

How tag count, sampling interval, payload size, overhead and retention combine into bandwidth and storage numbers, and how deadband and compression change them.

The model in four lines

Capacity planning for industrial data does not need a clever tool. It needs four lines of arithmetic and honest assumptions:

  1. Samples per second = number of tags ÷ sampling interval (seconds)
  2. Bytes per second = samples per second × bytes per sample × protocol and metadata overhead
  3. Data per day = bytes per second × 86,400
  4. Storage = data per day × retention days × stored copies ÷ compression ratio

Take 1,000 tags sampled every second at 24 bytes each, with a 1.5× overhead, kept for 30 days in two copies at 2:1 compression:

QuantityResult
Samples per second1,000
Bytes per second36,000 (0.288 Mbps)
Data per day3.11 GB
Storage for 30 days, 2 copies, 2:193.3 GB

You can reproduce this in the IIoT data budget calculator, which uses these values as its default example and shows how the result changes when you alter one assumption.

Where the bytes come from

The “bytes per sample” figure hides most of the uncertainty. A raw value takes between 1 and 8 bytes, but a stored sample also carries a timestamp (commonly 8 bytes), a quality code and some way of identifying the tag. A JSON message published over MQTT can be many times larger than the same sample in a binary historian format. The default of 24 bytes is a planning assumption, not a measurement: capture a few minutes of your real traffic, divide the size by the number of samples and use that figure.

Sampling versus change

Fixed-rate sampling records every tag every interval. Report by exception records a sample only when the value moves by more than a deadband, usually with a maximum interval so that a flat signal still produces an occasional heartbeat. Savings depend entirely on how the signals behave:

  • A slowly changing temperature or a level in a large tank can drop to a small fraction of the fixed-rate volume.
  • A noisy flow signal or a vibration measurement barely benefits, because it keeps leaving the deadband.
  • Discrete states (running, stopped, alarm) are naturally change-based and cheap.

Do not guess the saving. Replay a day of real data through your deadband setting and count what survives. Keep the deadband narrow enough that it cannot hide behavior you need for diagnosis, and consider holding raw data for a short window before reducing it.

Compression and retention tiers

Compression ratios depend on the data and on the engine: highly repetitive data compresses well, noisy data does not. Treat the ratio as something to measure on a sample. It is also worth separating retention tiers: raw samples for a short period, then aggregated values (averages, minimums, maximums per interval) for far longer. The articles on time-series compression and historian retention planning go deeper.

How each assumption moves the answer

Using the example above, changing one thing at a time gives:

ChangeStorage
Baseline93.3 GB
Sample twice as often (0.5 s)186.6 GB
Half the tags46.7 GB
Keep data twice as long (60 days)186.6 GB
Compression twice as good (4:1)46.7 GB

Sampling rate, tag count and retention scale storage linearly, so the cheapest saving is usually to question the requirement: does this tag really need one-second resolution for 30 days?

What the model leaves out

  • Indexes, replication, backups and growth reserve. Add a margin, and decide it in advance.
  • Catch-up traffic. After a link outage a store-and-forward gateway replays its buffer, so the network and the database see a burst well above the steady-state rate. Plan for it.
  • Query load. Dashboards and analytics read the data as well as write it.
  • Licensing and infrastructure cost. The calculator estimates volume only, not price.

If you are choosing between a historian and a data lake, the trade-offs are covered in Historian vs data lake.

THE NEXT STEP

From calculation to implementation.

Let’s look at your machine, your data flow or your production goal together. Describe your situation in a few sentences and the ASP Dijital team will reply by email.

Talk to ASP Dijital