Sizing an IIoT data budget: sampling, deadband and retention
How tag count, sampling interval, payload size, overhead and retention combine into bandwidth and storage numbers, and how deadband and compression change them.
ASP Dijital · IT Hub4 min readEnglish
The model in four lines
Capacity planning for industrial data does not need a clever tool. It needs four lines of arithmetic and honest assumptions:
Samples per second = number of tags ÷ sampling interval (seconds)
Bytes per second = samples per second × bytes per sample × protocol and metadata overhead
Data per day = bytes per second × 86,400
Storage = data per day × retention days × stored copies ÷ compression ratio
Take 1,000 tags sampled every second at 24 bytes each, with a 1.5× overhead, kept for 30 days in two copies at 2:1 compression:
Quantity
Result
Samples per second
1,000
Bytes per second
36,000 (0.288 Mbps)
Data per day
3.11 GB
Storage for 30 days, 2 copies, 2:1
93.3 GB
You can reproduce this in the IIoT data budget calculator, which uses these values as its default example and shows how the result changes when you alter one assumption.
Where the bytes come from
The “bytes per sample” figure hides most of the uncertainty. A raw value takes between 1 and 8 bytes, but a stored sample also carries a timestamp (commonly 8 bytes), a quality code and some way of identifying the tag. A JSON message published over MQTT can be many times larger than the same sample in a binary historian format. The default of 24 bytes is a planning assumption, not a measurement: capture a few minutes of your real traffic, divide the size by the number of samples and use that figure.
Sampling versus change
Fixed-rate sampling records every tag every interval. Report by exception records a sample only when the value moves by more than a deadband, usually with a maximum interval so that a flat signal still produces an occasional heartbeat. Savings depend entirely on how the signals behave:
A slowly changing temperature or a level in a large tank can drop to a small fraction of the fixed-rate volume.
A noisy flow signal or a vibration measurement barely benefits, because it keeps leaving the deadband.
Discrete states (running, stopped, alarm) are naturally change-based and cheap.
Do not guess the saving. Replay a day of real data through your deadband setting and count what survives. Keep the deadband narrow enough that it cannot hide behavior you need for diagnosis, and consider holding raw data for a short window before reducing it.
Compression and retention tiers
Compression ratios depend on the data and on the engine: highly repetitive data compresses well, noisy data does not. Treat the ratio as something to measure on a sample. It is also worth separating retention tiers: raw samples for a short period, then aggregated values (averages, minimums, maximums per interval) for far longer. The articles on time-series compression and historian retention planning go deeper.
How each assumption moves the answer
Using the example above, changing one thing at a time gives:
Change
Storage
Baseline
93.3 GB
Sample twice as often (0.5 s)
186.6 GB
Half the tags
46.7 GB
Keep data twice as long (60 days)
186.6 GB
Compression twice as good (4:1)
46.7 GB
Sampling rate, tag count and retention scale storage linearly, so the cheapest saving is usually to question the requirement: does this tag really need one-second resolution for 30 days?
What the model leaves out
Indexes, replication, backups and growth reserve. Add a margin, and decide it in advance.
Catch-up traffic. After a link outage a store-and-forward gateway replays its buffer, so the network and the database see a burst well above the steady-state rate. Plan for it.
Query load. Dashboards and analytics read the data as well as write it.
Licensing and infrastructure cost. The calculator estimates volume only, not price.
If you are choosing between a historian and a data lake, the trade-offs are covered in Historian vs data lake.
Let’s look at your machine, your data flow or your production goal together. Describe your situation in a few sentences and the ASP Dijital team will reply by email.