Engineering /Control Systems

ReferenceWorking6 min read

Stating a determinism budget for a control loop

TL;DR

Before writing loop code, fix four numbers: sample rate, allowed jitter on the sampling instant, worst-case actuation latency, and the deadline-miss policy. Every downstream choice then checks against the budget.

View as Markdown

Most real-time control problems are argued in adjectives — “fast enough”, “low jitter”, “hard real-time” — and adjectives cannot be verified. A determinism budget replaces them with four numbers agreed before implementation starts. After that, every architectural question has a right answer: does option A fit inside the budget or not.

This page is about writing that budget. It is protocol- and silicon-agnostic on purpose; the numbers change per project, the structure does not.

The four numbers

1. Sample rate (f_s)

The rate at which the loop reads its inputs and updates its output. Set it from the plant, not from what the CPU can do:

  • Rule of thumb: f_s between 10× and 20× the closed-loop bandwidth you need. Below 10× the phase lag from sampling eats your phase margin; far above 20× you are burning CPU and amplifying sensor noise through the derivative term for no dynamic benefit.
  • For a mechanical position loop with a 50 Hz bandwidth target, that is roughly 1–2 kHz. For a motor current loop it is typically 8–20 kHz because the electrical time constant is short. For a temperature loop, 1–10 Hz is often plenty and a faster loop just wastes power.
  • Write down the reasoning, not just the number. The next engineer needs to know whether 2 kHz was a plant requirement or a guess.

2. Sampling jitter (Δt_s)

The allowed variation in when the input is actually sampled, relative to the ideal period 1/f_s. This is the number most specs omit and most loops get bitten by.

Jitter matters because a control law assumes a fixed Δt. If the real interval wanders by ±15% and the code still divides by the nominal Δt in the derivative and integral terms, you have injected a disturbance proportional to the jitter and the signal slew rate. Effects:

  • The I term accumulates the wrong area.
  • The D term produces spikes on jittered intervals.
  • At the control-bandwidth frequency, timing jitter aliases into the loop as broadband noise you cannot filter out without also hurting the response.

Budget it as a percentage of the period and as an absolute time. “≤ 2% of period or ≤ 5 µs, whichever is larger” is a typical starting point for a mid-rate loop. For a current loop synchronised to PWM, the sampling instant is usually locked to a timer/PWM trigger and an ADC hardware trigger, and the jitter budget there is tens of nanoseconds — a software-timer-driven sample cannot meet it and the budget is what tells you that up front.

Two mitigations worth stating in the budget itself:

  • Timestamp every sample and feed the actual Δt into the loop maths, so jitter degrades gracefully instead of injecting noise.
  • Trigger sampling in hardware (timer-to-ADC, no CPU in the path) and let the ISR only consume the result.

3. Worst-case actuation latency (L_wc)

The time from “the sampling instant” to “the new actuator command is in effect” — end to end, worst case, not typical. Its components:

L_wc = t_acq        ADC conversion + settling
     + t_dispatch   interrupt latency + scheduler to the loop task
     + t_compute    the control law, worst-case path
     + t_output     DAC/PWM update, or the bus transaction to a remote drive
     + t_actuator   the actuator's own transport lag (often the biggest term)

Rules for filling it in:

  • Use worst case for every term. Interrupt latency under maximum interrupt load, t_compute with the branch that runs the anti-windup and the fault checks, bus latency including retransmission if the link allows it.
  • L_wc shows up in the loop as pure dead time. Dead time destroys phase margin fast: as a guide, keep L_wc under about 1/10 of the loop period, and treat anything above 1/4 of the period as a redesign trigger.
  • If the actuator is across a fieldbus, the bus cycle time and its determinism are now inside your control budget. A 1 kHz loop commanding a drive over a 1 ms-cycle bus has spent its entire latency budget on transport before any control happens. This is the calculation that decides whether the loop runs local to the drive or on a central controller.

4. Deadline-miss policy

Define what a missed deadline is and what the system does about it. “It should not happen” is not a policy.

  • What counts as a miss: loop iteration N has not completed before iteration N+1’s release. Instrument it — a GPIO toggled at loop start/end on a scope, or a software counter of overruns exported to diagnostics.
  • Allowed miss rate: for a soft loop, “≤ 1 in 10⁶ iterations and never two consecutive” might be fine. For a loop tied to a safety function, the answer is usually zero within the safety analysis and the response is a defined safe state, not a retry.
  • The response: hold last output, ramp to a safe value, trip a fault, or extrapolate one step. Each has failure modes — holding last output during a fast transient can be worse than a brief zero. State the choice and why.
  • Escalation: what happens on the 2nd, 10th, 100th consecutive miss.

What the budget lets you decide without arguing

Once the four numbers exist, these stop being opinions:

Question Decided by
Bare-metal, RTOS, or Linux with PREEMPT_RT? Can its worst-case interrupt-to-task latency + scheduler jitter fit inside Δt_s and the t_dispatch share of L_wc, measured, under load?
Loop in an ISR or a task? If t_compute is far below the period and the jitter budget is tight, use the ISR. If t_compute is a meaningful fraction of the period, use a task so it can be preempted by faster ISRs.
Control local to the actuator or central? Does the bus transport lag fit inside L_wc?
Which fieldbus? Its cycle time and cycle-time jitter versus Δt_s and L_wc.
Is this core allowed to run anything else? Only if the other work’s worst-case interference still leaves the budget intact — usually meaning a shielded/isolated core or a dedicated MCU.

Measuring against it

A budget you cannot measure is a wish. The standard instrumentation:

  • GPIO + scope/logic analyser: toggle a pin at ISR entry, at loop-math start, at output write. Persistence mode on the scope shows the jitter envelope directly. This is the ground truth; trust it over any software timestamp.
  • Cycle counter (DWT->CYCCNT on Cortex-M, PMCCNTR / perf on Cortex-A) around the compute path, logging min/max/histogram, not just mean.
  • Overrun counter in the loop, exported to the same diagnostics channel as everything else so field units report it.
  • Soak under worst-case load: every interrupt source firing, the comms stack saturated, the file system busy, cache cold. The typical case is not the number in the budget.

Hold the measurement for long enough to see the tail. Latency distributions in real systems have long tails driven by rare cache/TLB/bus-contention events, and the 99.99th percentile is often 3–10× the median. The budget is about the tail, so the measurement has to reach it.

The honest trade-off

Writing this budget costs a day or two of analysis and a negotiation with whoever owns the plant requirements, before any loop code exists. On a schedule under pressure that day is tempting to skip, and the loop will usually “work” on the bench without it.

What you lose by skipping it is the ability to say why it works, and the ability to catch — at design time rather than during integration — the cases where the chosen bus or OS cannot meet the timing. Those cases are expensive exactly in proportion to how late they are found. The budget moves that discovery to the cheapest possible point.

A tighter budget is not free either. Driving jitter to nanoseconds means hardware triggering, a shielded core, and no shared bus — real BOM and integration cost. The budget should be as loose as the plant genuinely allows, and no looser.

Talk to an engineer

Ask about this directly — the person who wrote it answers, not a sales desk.