{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Exubits Engineering",
  "home_page_url": "https://exubits.com/engineering",
  "feed_url": "https://exubits.com/feed.json",
  "description": "Reference pages, deep dives, teardowns and field notes on embedded product engineering.",
  "language": "en-IN",
  "items": [
    {
      "id": "https://exubits.com/engineering/can-canfd-ethernet-industrial-links",
      "url": "https://exubits.com/engineering/can-canfd-ethernet-industrial-links",
      "title": "Choosing between CAN, CAN FD and Ethernet industrial links",
      "summary": "CAN suits cheap, rugged, low-rate event traffic; CAN FD when payloads outgrew 8 bytes but a bus still fits; industrial Ethernet (PROFINET, EtherCAT, TSN) for megabits or tight cycle times. Expect to bridge, not replace.",
      "content_html": "<p>This decision gets made badly in two directions. Teams keep classic CAN because\n“it has always worked” and then spend a year fighting bus-load and 8-byte frames.\nOr they jump to industrial Ethernet everywhere and discover it costs more per\nnode, needs managed switches, and still does not give them determinism unless\nthey configured it for it.</p>\n<p>Here is how the three actually differ, and what each is bad at.</p>\n<h2 id=\"classic-can-iso-11898-1-2\">Classic CAN (ISO 11898-1/-2)</h2>\n<p><strong>What it is</strong>: a multi-drop differential bus, up to 1 Mbit/s, 8 data bytes per\nframe, non-destructive bitwise arbitration by identifier, strong error detection\n(15-bit CRC, bit monitoring, form/stuff checks) with automatic retransmission.</p>\n<p><strong>What it is genuinely good at</strong>:</p>\n<ul>\n<li><strong>Cost and ruggedness.</strong> A CAN transceiver is cents. The bus is two wires, no\nswitch, no hub, tolerant of ground shift and EMI, and it degrades gracefully —\na node dropping off does not take the segment down.</li>\n<li><strong>Event-driven traffic with priority.</strong> Arbitration means the lowest-numbered\nID always wins the bus with no collisions and no configuration. For “report\nthis alarm now” traffic that is exactly right.</li>\n<li><strong>Determinism at low load.</strong> Below roughly 30–40% bus load, worst-case latency\nfor a given priority is calculable and small.</li>\n<li><strong>Ecosystem.</strong> CANopen, J1939, and decades of tooling, diagnostics, and\nengineers who know it.</li>\n</ul>\n<p><strong>What it is bad at</strong> — state these before choosing it:</p>\n<ul>\n<li><strong>8 bytes.</strong> A 12-axis sensor packet or any structured record needs\nfragmentation (ISO-TP), which adds latency, state, and failure modes.</li>\n<li><strong>1 Mbit/s ceiling</strong>, and you only reach it on a short bus. At 1 Mbit/s the\npractical bus length is on the order of 40 m; longer buses force lower rates\nbecause arbitration needs the signal to propagate to the far end within a bit.</li>\n<li><strong>Bus load cliff.</strong> Above ~50–60% load, low-priority frames see large and\nhard-to-bound latency, and a chatty node can starve them entirely.</li>\n<li><strong>No native security.</strong> No authentication, no encryption; any node can send any\nID. Mitigations exist (bus guardians, MACsec-style add-ons, segmentation) but\nthey are bolt-ons.</li>\n<li><strong>A single retransmitting faulty node</strong> can dominate the bus until error\nconfinement kicks in.</li>\n</ul>\n<h2 id=\"can-fd-iso-11898-12015\">CAN FD (ISO 11898-1:2015)</h2>\n<p><strong>What it changes</strong>: up to <strong>64 data bytes</strong> per frame, and a <strong>second, faster\nbit rate</strong> for the data phase (commonly 2–5 Mbit/s, 8 Mbit/s achievable with\ncare) while the arbitration phase stays at the classic rate. Better CRC (17- or\n21-bit) for the larger payload.</p>\n<p><strong>Why it is often the right upgrade</strong>:</p>\n<ul>\n<li><strong>The 8-byte problem disappears</strong> without fragmentation. One 64-byte frame\ninstead of eight classic frames means less protocol overhead and less software\nstate.</li>\n<li><strong>Effective throughput rises several-fold</strong> for payload-heavy traffic, because\nboth the payload is bigger and the data phase is faster.</li>\n<li><strong>You keep the topology, the transceivers’ ruggedness, the arbitration model,\nand most of the tooling.</strong> It is an incremental change, not a re-architecture.</li>\n<li>CANopen FD and J1939 variants exist to carry the ecosystem forward.</li>\n</ul>\n<p><strong>What it is bad at / what to watch</strong>:</p>\n<ul>\n<li><strong>The arbitration phase is still limited to the classic bit rate.</strong> CAN FD does\nnot raise the number of <em>frames per second</em> you can arbitrate; it raises the\nbytes per frame. A system that is frame-rate-bound, not payload-bound, gets\nlittle from it.</li>\n<li><strong>Signal integrity gets harder.</strong> The fast data phase shrinks bit times to\nhundreds of nanoseconds; ringing, stub length, connector quality, and\ntransceiver loop delay now matter. A messy bus that tolerated 500 kbit/s\nclassic will not tolerate 5 Mbit/s FD without topology work and possibly ring\nor point-to-point rather than long stubs.</li>\n<li><strong>Mixed classic/FD segments need care.</strong> A classic-only node on an FD bus sees\nFD frames as errors and floods the bus unless partial networking / FD-tolerant\ntransceivers are used. Usually the whole segment must be FD.</li>\n<li><strong>Controller support</strong>: older MCUs have classic-only CAN peripherals. Check the\nsilicon, not the datasheet marketing.</li>\n</ul>\n<h2 id=\"switched-industrial-ethernet--profinet-ethercat-and-tsn\">Switched industrial Ethernet — PROFINET, EtherCAT, and TSN</h2>\n<p>Not one thing. “Industrial Ethernet” covers protocols with very different\ndeterminism stories, all on 100 Mbit/s or 1 Gbit/s PHYs.</p>\n<ul>\n<li><strong>PROFINET RT</strong> — standard switched Ethernet, prioritised frames, cycle times\ndown to ~1 ms; <strong>PROFINET IRT</strong> adds hardware time-slotting for sub-millisecond,\nlow-jitter cycles but needs IRT-capable switches.</li>\n<li><strong>EtherCAT</strong> — a single frame passes through every node, each reading and\nwriting its slice on the fly (“processing on the fly”). Extremely efficient,\ncycle times to tens of microseconds, but it is a specific master/slave topology\nwith EtherCAT slave controllers in each node, not general IP networking.</li>\n<li><strong>TSN</strong> (IEEE 802.1 set — 802.1AS, Qbv, Qci, CB, …) — brings scheduled,\nbounded-latency traffic to <em>standard</em> Ethernet, letting control traffic and IT\ntraffic share the same infrastructure. It is the direction the industry is\nmoving; it also requires TSN-capable switches and endpoints and non-trivial\nnetwork engineering (a schedule computed and pushed to every bridge).</li>\n</ul>\n<p><strong>What Ethernet buys you</strong>:</p>\n<ul>\n<li><strong>Bandwidth</strong>: 100–1000 Mbit/s, so video, firmware images, and high-channel-\ncount data are no longer a problem.</li>\n<li><strong>Routable, IT-integratable</strong>: TCP/IP, TLS, MQTT/OPC UA to the cloud on the\nsame wire (with segmentation and a firewall).</li>\n<li><strong>Determinism <em>if configured for it</em></strong>: IRT, EtherCAT, or TSN give bounded\nlatency and low jitter — but plain “Ethernet” does not; a standard switch under\nload has queueing delay and no guarantees.</li>\n</ul>\n<p><strong>What it is bad at</strong>:</p>\n<ul>\n<li><strong>Cost and complexity per node</strong>: a PHY, magnetics or connector, often a\nmanaged or special switch, more powerful MCU/SoC, and a stack. Multiples of a\nCAN node.</li>\n<li><strong>Topology and cabling</strong>: star or line with switches, 100 m segment limit per\nhop, more connectors, more to go wrong mechanically in a vibrating machine.</li>\n<li><strong>Determinism is opt-in and fragile to misconfiguration.</strong> A well-meaning IT\nchange, a non-TSN device plugged into a TSN segment, or an unmanaged switch\ndropped in “temporarily” can quietly break the timing guarantees.</li>\n<li><strong>Security surface</strong>: now you are on an IP network, with everything that\nimplies. IEC 62443 becomes your problem.</li>\n</ul>\n<h2 id=\"a-decision-order-that-works\">A decision order that works</h2>\n<ol>\n<li><strong>What is the largest single payload, and how often?</strong> ≤ 8 bytes, event-rate\n→ classic CAN is still fine. 8–64 bytes → CAN FD. Kilobytes, or any streaming\n→ Ethernet.</li>\n<li><strong>What is the tightest cycle time with bounded jitter you must hit?</strong>\nAbove 5–10 ms and modest, CAN/CAN FD or PROFINET RT. Around 1 ms, PROFINET IRT\nor CAN FD on a lightly loaded bus. Under 250 µs with many axes, EtherCAT. Mixed\ncritical and best-effort on one wire, TSN.</li>\n<li><strong>Node count, cost target, and environment.</strong> Dozens of cheap rugged nodes on\na harness → bus wins. A cell with a handful of high-value devices and a video\nfeed → Ethernet.</li>\n<li><strong>Does it need to talk to IT / cloud directly?</strong> Yes and at rate → Ethernet.\nOccasionally → keep the fieldbus and put the IP stack in the gateway.</li>\n<li><strong>Who maintains it for ten years?</strong> A plant electrician can fault-find a CAN\nbus with a multimeter. A TSN schedule needs someone who understands 802.1Qbv.</li>\n</ol>\n<h2 id=\"the-realistic-answer-is-usually-bridge\">The realistic answer is usually “bridge”</h2>\n<p>Greenfield-everything-Ethernet is rare. The common shape is CAN or CAN FD at the\nmachine edge for sensors and drives, aggregated by a <strong>gateway</strong> that speaks the\nfieldbus on one side and OPC UA / MQTT over Ethernet on the other, doing protocol\ntranslation, buffering across link outages, and store-and-forward. That gateway is\nalso the right place to put the security boundary, the data-model mapping, and the\nfirmware-update path — rather than pushing all of that cost into every edge node.</p>\n<p>Designing that seam well — which data is published at what rate, what is buffered,\nhow backpressure and reconnection behave — is usually more of the engineering\neffort than picking the physical layer.</p>\n",
      "image": "https://exubits.com/engineering/og/can-canfd-ethernet-industrial-links.png",
      "date_published": "2026-09-03T00:00:00.000Z",
      "date_modified": "2026-09-03T00:00:00.000Z",
      "authors": [
        {
          "name": "Exubits Engineering",
          "url": "https://exubits.com/authors/exubits-engineering"
        }
      ],
      "tags": [
        "Reference",
        "can-fd",
        "ethercat",
        "tsn",
        "profinet"
      ]
    },
    {
      "id": "https://exubits.com/engineering/production-handover-pack",
      "url": "https://exubits.com/engineering/production-handover-pack",
      "title": "What a production hand-over pack has to contain",
      "summary": "It lets someone else build, test, certify and change the product without the original team: design data, a reproducible firmware build, factory tests with limits, the compliance file, an SBOM, and known issues.",
      "content_html": "<p>A hand-over pack has one job: let a competent engineer who was not on the project\nbuild the product, prove it works, keep it compliant, support it in the field, and\nchange it — without phoning the original team. Anything less is a dependency, not\na delivery.</p>\n<p>This is the contents list we work to. It spans hardware, firmware, test, and\ncompliance, because a product that splits those across separate incomplete\nhand-overs has no owner for the seams.</p>\n<h2 id=\"1-released-versioned-design-data\">1. Released, versioned design data</h2>\n<p>Not “the latest files” — a <strong>release</strong>, tagged, with a revision and a date.</p>\n<p><strong>Hardware</strong></p>\n<ul>\n<li>Schematics (PDF) and source project, at the released revision.</li>\n<li>PCB layout source, Gerbers, drill, netlist, fab drawing with stack-up and\nimpedance spec, assembly drawings.</li>\n<li><strong>BOM</strong> with manufacturer part numbers, approved alternates, and DNP items\nmarked. Lifecycle status noted for anything NRND or single-source.</li>\n<li>Mechanical: enclosure models, gaskets, thermal interface parts, fasteners,\nlabels and their artwork.</li>\n<li>The <strong>ECO/ECR history</strong> — every change since the first build, why it was made,\nand which build it entered. This is how the factory knows a v3 board is not a\nv2 board with a sticker.</li>\n</ul>\n<p><strong>Firmware and software</strong></p>\n<ul>\n<li>Source, in a repo the client controls, tagged for the release.</li>\n<li>A <strong>reproducible build</strong>: pinned toolchain (container or documented exact\nversions), pinned dependencies, one command. The test is that a fresh machine\nreproduces the released binary — or the released image manifest — bit for bit.\n(For a Yocto-based product, this is its own checklist; see the Embedded Linux\nhub’s BSP hand-over reference.)</li>\n<li>The released binaries themselves, with checksums and signatures, and the public\nkeys to verify them.</li>\n<li>Build and flashing instructions from empty machine to programmed unit.</li>\n</ul>\n<h2 id=\"2-requirements-and-traceability\">2. Requirements and traceability</h2>\n<ul>\n<li>The <strong>requirements</strong> the product was built to, at their released version.</li>\n<li>A <strong>traceability matrix</strong>: requirement → design element → verification test →\nresult. Even a lightweight one. This is what a customer audit, a safety\nassessor, or a future change-impact analysis needs, and it cannot be\nreconstructed later without re-deriving intent.</li>\n<li>Interface control documents for every external interface (connector pinouts,\nprotocols, message definitions, electrical limits).</li>\n</ul>\n<h2 id=\"3-verification-and-validation-records\">3. Verification and validation records</h2>\n<ul>\n<li><strong>Test plans and procedures</strong>, with pass/fail criteria that are numbers, not\nadjectives.</li>\n<li><strong>Results</strong> for the released revision: what was run, on how many units, with\nwhat outcome, signed and dated. Raw data retained, not just a summary.</li>\n<li><strong>Coverage statement</strong>: what was tested, what was explicitly not, and why. A\nhand-over that implies everything was verified when it was not is worse than\none that is honest about the gaps.</li>\n<li>Regression suite and how to run it, so the client can re-verify after a change.</li>\n</ul>\n<h2 id=\"4-compliance--certification-file\">4. Compliance / certification file</h2>\n<p>Whatever regimes apply (EMC, safety, radio, environmental, sector-specific):</p>\n<ul>\n<li><strong>Test reports</strong> from the accredited lab, in full, not just the certificate.</li>\n<li>The <strong>Declaration of Conformity</strong> and the technical file / construction file\nbehind it.</li>\n<li>The <strong>exact configuration tested</strong>: hardware revision, firmware version,\ncables, orientation, peripherals, any test-mode firmware used. A certificate\nthat does not pin the configuration does not transfer to production cleanly.</li>\n<li><strong>Applied standards and their editions.</strong> Standard editions change; the file\nmust say which one the product was assessed against.</li>\n<li><strong>Conditions and limitations</strong> of the approval, and the class/limits the\nproduct passed against (with margin, ideally).</li>\n<li>What invalidates the certification — which changes force re-test. This is the\nsingle most useful sentence for the team that inherits it.</li>\n<li>Radio module certs (modular approval IDs) and the conditions attached to using\nthem under that approval.</li>\n</ul>\n<h2 id=\"5-software-bill-of-materials-and-update-story\">5. Software bill of materials and update story</h2>\n<ul>\n<li><strong>SBOM</strong> (SPDX or CycloneDX) for the shipped image: every component, version,\nlicence, and source. Increasingly a legal requirement, and the basis for\nanswering “are we affected by CVE-XXXX” quickly.</li>\n<li><strong>Licence compliance</strong>: the licence manifest, written offers where copyleft\nrequires them, and the corresponding source archive.</li>\n<li><strong>Secure/verified boot</strong>: key hierarchy, what signs what, where keys are stored\nor fused, key-rotation procedure, and the recovery path for a unit that fails a\nsignature check.</li>\n<li><strong>Field update mechanism</strong>: how an update is packaged, signed, delivered,\napplied, and rolled back; A/B or recovery-slot behaviour; what happens on power\nloss mid-update; version-compatibility rules.</li>\n<li><strong>Provisioning</strong>: per-unit secrets, certificates, serial numbers — how they are\ngenerated, injected at the factory, and stored/backed up.</li>\n</ul>\n<h2 id=\"6-manufacturing-pack\">6. Manufacturing pack</h2>\n<ul>\n<li><strong>Factory test procedure</strong> with a defined sequence, measurement points, and\n<strong>limits</strong> for every measured parameter. This is what makes a contract\nmanufacturer able to accept or reject a unit without judgement calls.</li>\n<li><strong>Test fixtures</strong>: design, firmware, calibration procedure and interval, and a\nspare or the data to build one.</li>\n<li><strong>Programming</strong>: what gets programmed, in what order, with what tool, and how\nthe line verifies it took.</li>\n<li><strong>Calibration</strong>: what is calibrated, against what reference, to what tolerance,\nand where the calibration constants are stored.</li>\n<li><strong>Serialisation and labelling</strong>: numbering scheme, label content and placement,\nwhat is recorded per unit and where that record lives.</li>\n<li><strong>Packaging and ESD/transport</strong> requirements.</li>\n<li><strong>First-article inspection</strong> report and the golden-unit definition.</li>\n<li><strong>Yield and known failure modes</strong> from the builds so far, with the diagnosis\nfor each — so the line does not re-learn them.</li>\n</ul>\n<h2 id=\"7-field-support-material\">7. Field support material</h2>\n<ul>\n<li><strong>Service manual</strong>: diagnosis flow, safe-to-replace parts, what needs a\ndepot, recovery procedures (including “how to un-brick”).</li>\n<li><strong>Diagnostics</strong>: what the product logs, how to extract it, and how to read it.</li>\n<li><strong>RMA criteria</strong> and the data to capture on a return.</li>\n<li><strong>Spares strategy</strong>: which parts, expected failure rates if known, last-time-buy\nexposure on any component.</li>\n</ul>\n<h2 id=\"8-the-known-issues-and-deferred-work-list\">8. The known-issues and deferred-work list</h2>\n<p>Every project has one. A hand-over without it means the client discovers the list\none incident at a time.</p>\n<ul>\n<li>Open bugs with severity, reproduction, and any workaround.</li>\n<li>Deferred features and the reason they were cut.</li>\n<li>Design compromises made under schedule pressure and what a proper fix looks\nlike.</li>\n<li>Component or supplier risks not yet resolved.</li>\n<li>Anything that “works but we do not fully understand why”.</li>\n</ul>\n<h2 id=\"9-contacts-licences-and-accounts\">9. Contacts, licences, and accounts</h2>\n<ul>\n<li>Which third-party licences, subscriptions, or accounts the product depends on\n(cloud tenant, certificate authority, code-signing service, paid libraries),\nwho holds them, and renewal dates.</li>\n<li>Escrow arrangements if any.</li>\n<li>A named transition contact and a defined support window, so the hand-over is a\nramp, not a cliff.</li>\n</ul>\n<h2 id=\"how-to-know-the-pack-is-complete\">How to know the pack is complete</h2>\n<p>Run the test that matters: give the pack to an engineer who was not on the\nproject and have them, using only what is in the box, (a) build the firmware and\nget a matching binary, (b) build and pass one unit through the factory test, and\n(c) answer “what change would invalidate the CE marking”. If any of the three\nneeds a phone call, the pack has a hole, and the hole is cheaper to fill now.</p>\n<h2 id=\"the-trade-off\">The trade-off</h2>\n<p>Assembling this properly is weeks of work, concentrated at the end of a project\nwhen budget and patience are thinnest, and most of it produces no visible feature.\nIt is genuinely tempting to ship the design files and a README and call the rest\n“support”.</p>\n<p>The cost of that shortcut is entirely borne later and by someone else: the first\nCM build that fails because the test limits were in someone’s head, the CVE\nresponse that takes three weeks because there is no SBOM, the field failure with\nno service procedure, the change that quietly invalidates a certification because\nnobody wrote down what the certification depended on. The pack is insurance, and\nlike insurance its value is invisible right up until the moment it is the only\nthing that helps.</p>\n",
      "image": "https://exubits.com/engineering/og/production-handover-pack.png",
      "date_published": "2026-09-03T00:00:00.000Z",
      "date_modified": "2026-09-03T00:00:00.000Z",
      "authors": [
        {
          "name": "Exubits Engineering",
          "url": "https://exubits.com/authors/exubits-engineering"
        }
      ],
      "tags": [
        "Reference",
        "v-and-v",
        "traceability",
        "certification"
      ]
    },
    {
      "id": "https://exubits.com/engineering/schematic-review-before-layout",
      "url": "https://exubits.com/engineering/schematic-review-before-layout",
      "title": "What a schematic review should catch before layout starts",
      "summary": "A schematic review catches the errors that get far more expensive after layout: power sequencing and budgets, reset and strapping, every net's return path, connector pinout and ESD, test access, and BOM risk.",
      "content_html": "<p>The schematic review is the last cheap checkpoint. A missed pull-up found here is\na two-minute edit. Found after layout it is a re-route; found after fab it is a\nbodge wire and a respin; found in the field it is a recall. The review’s purpose\nis to spend an hour now to not spend a month later.</p>\n<p>This is the checklist we run schematics against before releasing them to layout.\nIt assumes the design is functionally “done” — the point is to find what “done”\nmissed.</p>\n<h2 id=\"power-the-section-that-causes-the-most-respins\">Power: the section that causes the most respins</h2>\n<h3 id=\"rail-inventory-and-budget\">Rail inventory and budget</h3>\n<ul>\n<li><strong>List every rail</strong>: voltage, tolerance, estimated and worst-case current,\nsource (which regulator), and every load on it. A spreadsheet, not a mental\nmodel.</li>\n<li><strong>Check each regulator against its worst-case load</strong> including inrush and\ntransient, not typical. Add margin for the load estimate being wrong — 30% is\ncommon at this stage.</li>\n<li><strong>Thermal</strong>: <code>(Vin − Vout) × I</code> for every linear regulator. An LDO dropping\n3.3 V at 500 mA is dissipating 1.65 W and needs a heatsinking copper plan that\nlayout has to know about <em>now</em>.</li>\n<li><strong>Efficiency and input current</strong>: does the upstream supply / connector / fuse\nactually deliver the sum of input currents at low line?</li>\n</ul>\n<h3 id=\"sequencing\">Sequencing</h3>\n<ul>\n<li><strong>Does the SoC/FPGA have a required power-up and power-down order?</strong> Most do\n(core before I/O, or a specific ramp relationship). Violating it can cause\nlatch-up or excess current through internal diodes. Check the datasheet’s\nsequencing section and confirm the design enforces it — enable-pin daisy\nchains, a sequencer IC, or a supervisor.</li>\n<li><strong>Power-down order matters too</strong>, and is more often missed. A rail collapsing\nin the wrong order can back-drive an I/O bank.</li>\n<li><strong>What happens on a brown-out</strong> — does the sequence restart cleanly, or can it\nhang half-powered?</li>\n</ul>\n<h3 id=\"decoupling-as-schematic-intent\">Decoupling, as schematic intent</h3>\n<ul>\n<li>Every power pin has decoupling specified with values and count per the device\ndatasheet. Bulk capacitance per rail sized for the load step.</li>\n<li>It is fine that placement is layout’s job — but the schematic must show the\n<em>intent</em> so the reviewer can check nothing is missing and layout knows the\ntarget.</li>\n</ul>\n<h2 id=\"reset-clocks-and-strapping\">Reset, clocks, and strapping</h2>\n<ul>\n<li><strong>Every reset</strong>: source, polarity, pull resistor, RC or supervisor timing,\nand which devices it reaches. Open-drain resets wired together need one pull-up,\nnot none and not five.</li>\n<li><strong>Reset supervisor threshold</strong> matched to the SoC’s minimum operating voltage,\nwith hysteresis, so a sagging rail does not leave the part running out of spec.</li>\n<li><strong>Watchdog</strong>: present, wired to actually reset the system, and not defeatable\nby the failure it is meant to catch.</li>\n<li><strong>Boot strapping / mode pins</strong>: this is a classic post-layout disaster. Every\nstrap pin identified, its required level confirmed against the boot-mode table,\nand the resistor value chosen so it wins against the pin’s other function\n(often the same pin is a functional I/O with its own loading). Note which straps\nneed to be <em>changeable</em> for bring-up and give them a header or a 0 Ω option.</li>\n<li><strong>Clocks</strong>: every oscillator/crystal has the load caps per the crystal spec and\nthe oscillator’s drive level checked. PLL supplies filtered. Spread-spectrum\nchoice made deliberately (it helps EMC, it can hurt a camera or ADC).</li>\n</ul>\n<h2 id=\"every-net-reference-and-return-path\">Every net: reference and return path</h2>\n<p>Signal integrity is decided in the schematic by what you connect to what, before\na single trace exists.</p>\n<ul>\n<li><strong>For each interface, name the return-current path.</strong> A differential pair\nreferenced to a plane that has a split under it will radiate and fail EMC — and\nthe fix is a stitching-cap or plane change that is far easier to plan now.</li>\n<li><strong>Series termination / source termination</strong> resistors placed in the schematic\nfor fast single-ended nets (RGMII, parallel memory, fast GPIO). Value TBD by\nlayout, but the footprint must exist.</li>\n<li><strong>Differential pairs</strong> (USB, Ethernet, MIPI, PCIe, LVDS): correct AC coupling\nwhere the standard requires it, correct common-mode termination, correct\npolarity, and pairs that the connector pinout does not force to cross.</li>\n<li><strong>Unused inputs tied off</strong>, not floating. Unused outputs left open. Unused\nop-amp sections wired as followers to a mid-rail, not left open-loop.</li>\n<li><strong>Level shifting</strong> wherever two voltage domains meet — confirm direction,\nspeed, and that the translator’s supported data rate exceeds the bus.</li>\n</ul>\n<h2 id=\"connectors-and-the-outside-world\">Connectors and the outside world</h2>\n<p>Connectors are where field failures enter.</p>\n<ul>\n<li><strong>Pinout sanity</strong>: power and ground pins adjacent enough to carry the current;\nhot-plug order (ground first, then power, then signal) if the connector is\never mated live.</li>\n<li><strong>ESD/EOS protection on every externally accessible pin</strong>: USB, Ethernet\n(plus the magnetics and Bob Smith termination), buttons, connector I/O.\nTVS diodes with the right stand-off and clamping voltage, placed at the\nconnector.</li>\n<li><strong>Reverse-polarity and overvoltage</strong> protection on the power input, sized for\nthe real worst case (a field tech with a 24 V supply on a 12 V input).</li>\n<li><strong>Miswire survival</strong>: what happens if the harness is plugged in shifted by one\npin? On rugged products this is worth a deliberate answer.</li>\n<li><strong>Test points</strong> on every rail, reset, boot strap, and key signal — see DFT.</li>\n</ul>\n<h2 id=\"design-for-test-and-manufacture\">Design for test and manufacture</h2>\n<ul>\n<li><strong>DFT</strong>: bed-of-nails or flying-probe access to rails and critical nets. A\nJTAG/SWD header (even if depopulated in production). A UART console broken out.\nA way to hold the board in reset and in each boot mode. Programming access for\nevery programmable device, and the programming order/method noted.</li>\n<li><strong>DFM</strong>: no parts on both sides that force two reflow passes unless intended;\npackage choices the assembler can place; no 0201s where 0402 would do; fiducials\npresent; polarity markings that survive assembly.</li>\n<li><strong>First-article bring-up plan</strong> implied by the schematic: can you bring rails up\none at a time? Is there a “populate R1, leave R2 off” staged power-on option?</li>\n</ul>\n<h2 id=\"bom-and-component-risk\">BOM and component risk</h2>\n<ul>\n<li><strong>Lifecycle</strong>: any part NRND or single-source? Flag it now; a second source or\na footprint that accepts alternates is a schematic decision.</li>\n<li><strong>Ratings with margin</strong>: capacitor voltage derating (ceramics lose capacitance\nwith DC bias — a 6.3 V X5R on a 5 V rail may be at half its rated value),\nresistor power, inductor saturation current at the real peak, MOSFET SOA.</li>\n<li><strong>Tolerance stack-up</strong> on anything that sets a threshold, a timing, or a\nfeedback divider.</li>\n<li><strong>Passives count</strong>: every DNP is intentional and labelled; every “we’ll tune\nthis in bring-up” has a real footprint and a starting value.</li>\n</ul>\n<h2 id=\"how-to-actually-run-it\">How to actually run it</h2>\n<ul>\n<li><strong>Two reviewers minimum</strong>, at least one who did not draw the schematic.</li>\n<li><strong>Page by page, net by net on the critical interfaces</strong>, out loud, against the\ndevice datasheets open on the table — not a glance-through.</li>\n<li><strong>Every datasheet’s “layout guidelines” and “design checklist” section read</strong>,\nbecause vendors put the expensive mistakes there.</li>\n<li><strong>Track findings as a list with owners and states.</strong> The review is not done\nwhen the meeting ends; it is done when the list is closed and re-checked.</li>\n<li><strong>Re-review after the fixes.</strong> Fixes introduce errors.</li>\n</ul>\n<h2 id=\"the-trade-off\">The trade-off</h2>\n<p>A proper schematic review on a medium-complexity board is the better part of a\nday for two engineers, plus the fix cycle — call it two to three engineer-days\nbefore layout even starts. It feels like a delay when the schematic “looks\nfinished”.</p>\n<p>What it does not do: it will not catch problems that only exist in the physical\nlayout — coupling, plane resonance, actual trace impedance, thermal reality,\nmechanical fit. Those need a layout review and, ultimately, measurement on\nhardware. The schematic review is necessary and cheap; it is not sufficient, and\ntreating a passed schematic review as “the hard part is over” is its own failure\nmode.</p>\n",
      "image": "https://exubits.com/engineering/og/schematic-review-before-layout.png",
      "date_published": "2026-09-03T00:00:00.000Z",
      "date_modified": "2026-09-03T00:00:00.000Z",
      "authors": [
        {
          "name": "Exubits Engineering",
          "url": "https://exubits.com/authors/exubits-engineering"
        }
      ],
      "tags": [
        "Reference",
        "signal-integrity",
        "power-design",
        "dfm"
      ]
    },
    {
      "id": "https://exubits.com/engineering/sensor-to-display-latency",
      "url": "https://exubits.com/engineering/sensor-to-display-latency",
      "title": "Where the milliseconds go between sensor and display",
      "summary": "Glass-to-glass latency is exposure + sensor readout + CSI-2 + ISP + buffer and queue waits + composition + one or two display refreshes. Most of it is exposure and buffering, not code; shrink queues and sync to vsync.",
      "content_html": "<p>“The camera feels laggy” is a latency-budget problem, and it is almost always\nlost in places people do not instrument: exposure time, the number of buffers in\neach queue, and the wait for the next display refresh. Optimising the application\ncode is usually the smallest available win.</p>\n<p>This page walks the whole path from photons to photons and gives the order of\nmagnitude of each stage, so you know which term to attack.</p>\n<h2 id=\"the-pipeline-stage-by-stage\">The pipeline, stage by stage</h2>\n<h3 id=\"1-exposure-integration-time\">1. Exposure (integration time)</h3>\n<p>The sensor integrates light for the exposure time. The <em>photons that form a\nframe</em> are spread across that whole window, so the effective latency contribution\nis roughly <strong>half the exposure time</strong> for motion (the centroid of the exposure),\nbut the frame is not readable until the full exposure ends, so for end-to-end\ntiming you count the <strong>full exposure time</strong>.</p>\n<ul>\n<li>Bright scene, short exposure (1–2 ms): negligible.</li>\n<li>Indoor / auto-exposure (8–33 ms): this is often the single largest term, and it\nis invisible in any software profiler.</li>\n<li>Low light with a 1/30 s exposure: 33 ms before anything else has happened.</li>\n</ul>\n<p>Lever: cap the maximum auto-exposure time and accept more gain (noise) if latency\nmatters more than SNR. This is an ISP/3A tuning decision, not a code change.</p>\n<h3 id=\"2-sensor-readout\">2. Sensor readout</h3>\n<p>The pixel array is read out row by row and streamed. Readout takes on the order\nof <strong>one frame period</strong> at the sensor’s current mode — a rolling shutter sensor\nat 60 fps takes ~16 ms to clock the whole frame out, and the bottom of the frame\nis genuinely ~16 ms “younger” than the top.</p>\n<ul>\n<li>Higher frame-rate modes read out faster (and usually crop or bin).</li>\n<li>Global shutter removes the top-to-bottom skew but not the readout transport\ntime.</li>\n<li>Lever: run the sensor faster than the display rate if the SoC can take it —\nreading a 60 fps stream for a 30 fps display halves this term and the exposure\ncap becomes easier to hit.</li>\n</ul>\n<h3 id=\"3-mipi-csi-2-transport\">3. MIPI CSI-2 transport</h3>\n<p>Serialised over the D-PHY/C-PHY lanes into the SoC’s CSI receiver. At any sane\nlane rate this is <strong>sub-millisecond per frame</strong> for typical resolutions — it is\nalmost never the problem. It shows up only if lanes are marginal and you are\ngetting retimed/corrupted lines that force retry or drop.</p>\n<h3 id=\"4-receiver-dma-to-memory\">4. Receiver DMA to memory</h3>\n<p>The CSI receiver writes frames into a ring of buffers in DRAM. Two things here:</p>\n<ul>\n<li>The write itself is fast (bounded by memory bandwidth, sub-ms to low-ms).</li>\n<li><strong>The number of buffers in this ring is a latency knob.</strong> More buffers =\nsmoother under jitter, but a frame can sit in the ring for up to\n<code>(buffers − 1) × frame_period</code> before anything consumes it. Four buffers at\n30 fps is up to 100 ms of potential sit time. This is the classic hidden lag.</li>\n</ul>\n<h3 id=\"5-isp\">5. ISP</h3>\n<p>Debayer, black level, lens shading, denoise, sharpening, tone mapping, colour\nconversion, scaling. On a hardware ISP this is <strong>a few milliseconds</strong> and often\npipelined with readout. In software (libcamera soft ISP, OpenCV) it can be tens\nof milliseconds and steals CPU from everything else.</p>\n<ul>\n<li>3A (auto-exposure/white-balance/focus) runs here and feeds <em>the next</em> frame —\nit adds a frame of loop latency to convergence, not to the display path, but a\nslow-converging AE means longer exposures for longer, which loops back to\nstage 1.</li>\n<li>Lever: use the hardware ISP path (V4L2 media controller / libcamera with the\nvendor pipeline handler) rather than a software fallback.</li>\n</ul>\n<h3 id=\"6-application--buffer-handoff\">6. Application / buffer handoff</h3>\n<p>The frame becomes a <code>dmabuf</code> handed to whatever draws it — a Qt <code>QVideoSink</code>, a\nGStreamer <code>appsink</code>/<code>glimagesink</code>, an LVGL canvas, a Weston/Wayland client, a\ndirect DRM/KMS plane.</p>\n<ul>\n<li><strong>Every queue between elements is <code>depth × frame_period</code> of potential latency.</strong>\nGStreamer’s default queue sizes, a <code>v4l2src</code> with many buffers, a compositor\nthat triple-buffers — each is a place frames wait.</li>\n<li><strong>Zero-copy or not.</strong> If the frame is memcpy’d (or worse, uploaded to a GL\ntexture) at each stage, that is memory-bandwidth time and CPU stalls, a few ms\neach and more under load. <code>dmabuf</code> import all the way to the display plane\navoids it.</li>\n<li>Lever: shortest possible pipeline. A camera frame on its own DRM/KMS overlay\nplane, composited by the display controller, skips GPU composition entirely.</li>\n</ul>\n<h3 id=\"7-composition\">7. Composition</h3>\n<p>If the UI draws the video into a scene (overlays, HUD, controls), the GPU\ncomposites. One frame period at the render rate, plus the GPU’s own latency.\nUsing a hardware overlay plane for the video and only compositing the UI chrome\navoids putting the video through this stage.</p>\n<h3 id=\"8-display-refresh-and-panel-response\">8. Display refresh and panel response</h3>\n<ul>\n<li>The scanout waits for the next <strong>vsync</strong> to start sending the new frame:\n0 to one refresh period of wait (average half). At 60 Hz that is up to 16.7 ms.</li>\n<li>Double/triple buffering adds <strong>one or two more refresh periods</strong> depending on\nwhether a frame misses its flip deadline.</li>\n<li>The panel’s own response: LCD pixel response and its internal frame buffer /\noverdrive processing, commonly <strong>one frame</strong> for an LCD, sometimes more for a\npanel with a scaler or “image enhancement” it will not let you disable.</li>\n<li>Lever: <code>PAGE_FLIP</code> synced to vsync with exactly the buffers you need (often\ndouble, not triple), and pick a panel/timing controller with a documented,\nlow, fixed latency. Disable panel-side processing.</li>\n</ul>\n<h2 id=\"adding-it-up\">Adding it up</h2>\n<p>A representative indoor 30 fps preview, no tuning:</p>\n<table>\n<thead>\n<tr>\n<th>Stage</th>\n<th>Typical</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Exposure</td>\n<td>15–33 ms</td>\n</tr>\n<tr>\n<td>Readout</td>\n<td>~16–33 ms</td>\n</tr>\n<tr>\n<td>CSI-2 + DMA</td>\n<td>1–3 ms</td>\n</tr>\n<tr>\n<td>Buffer ring wait (4 deep)</td>\n<td>0–100 ms</td>\n</tr>\n<tr>\n<td>ISP (hardware)</td>\n<td>2–6 ms</td>\n</tr>\n<tr>\n<td>Pipeline queues (GStreamer defaults)</td>\n<td>30–100 ms</td>\n</tr>\n<tr>\n<td>GPU composition</td>\n<td>~16–33 ms</td>\n</tr>\n<tr>\n<td>Vsync wait + buffering</td>\n<td>16–50 ms</td>\n</tr>\n<tr>\n<td>Panel</td>\n<td>8–33 ms</td>\n</tr>\n</tbody>\n</table>\n<p>That easily reaches <strong>150–250 ms</strong> — the “laggy” complaint — and only ~5 ms of it\nis anything a code profiler would show you.</p>\n<p>The same pipeline, tuned:</p>\n<ul>\n<li>Sensor at 60 fps, AE capped at 8 ms.</li>\n<li>Two CSI buffers, not four.</li>\n<li>Hardware ISP path.</li>\n<li>No GStreamer queues, or <code>leaky=downstream max-size-buffers=1</code>.</li>\n<li>Video on a dedicated KMS overlay plane, UI composited separately.</li>\n<li>Double-buffered page flip synced to vsync, panel processing off.</li>\n</ul>\n<p>lands in the <strong>40–70 ms</strong> range, and most of what remains is exposure, readout,\nand one refresh period — the irreducible physics.</p>\n<h2 id=\"how-to-measure-it\">How to measure it</h2>\n<p>Do not trust per-stage estimates; measure glass-to-glass:</p>\n<ul>\n<li><strong>Photodiode + LED + scope.</strong> Flash an LED in frame, detect it with a\nphotodiode taped to the display, measure LED-on to display-brightens on the\nscope. This is the only number that counts.</li>\n<li><strong>On-screen millisecond counter filmed with a 120–240 fps camera</strong> alongside a\nphysical stopwatch/timer in the same shot; count frames of offset.</li>\n<li><strong>Per-stage</strong>: V4L2 buffer timestamps (<code>v4l2_buffer.timestamp</code>),\n<code>GST_DEBUG</code> with <code>GST_TRACERS=latency</code>, DRM/KMS <code>PAGE_FLIP</code> event timestamps,\nand a GPIO toggle at each app-level handoff. Line them up against the\nglass-to-glass number to find the fat stage.</li>\n</ul>\n<h2 id=\"the-trade-offs-you-are-actually-making\">The trade-offs you are actually making</h2>\n<ul>\n<li><strong>Fewer buffers = lower latency, less tolerance to scheduling jitter.</strong> Drop a\ndeadline with two buffers and you get a visible stutter; with four you get lag\nbut no stutter. Pick per product — a welding HUD wants low latency, a\nsecurity-review monitor wants no dropped frames.</li>\n<li><strong>Shorter exposure = lower latency, more noise.</strong> You are trading SNR for\nresponsiveness, frame by frame, in the AE tuning.</li>\n<li><strong>Overlay plane = low latency, less compositing flexibility.</strong> Overlay planes\nhave format, scaling, and count limits set by the display controller; a complex\nUI that must blend with the video may not fit and has to go through the GPU.</li>\n<li><strong>Running the sensor faster = lower latency and easier AE, more MIPI and memory\nbandwidth, more power and heat.</strong> On a thermally constrained enclosure that can\nbe the deciding constraint.</li>\n</ul>\n",
      "image": "https://exubits.com/engineering/og/sensor-to-display-latency.png",
      "date_published": "2026-09-03T00:00:00.000Z",
      "date_modified": "2026-09-03T00:00:00.000Z",
      "authors": [
        {
          "name": "Exubits Engineering",
          "url": "https://exubits.com/authors/exubits-engineering"
        }
      ],
      "tags": [
        "Reference",
        "latency-budget",
        "mipi-csi-2",
        "isp-tuning"
      ]
    },
    {
      "id": "https://exubits.com/engineering/stating-a-determinism-budget",
      "url": "https://exubits.com/engineering/stating-a-determinism-budget",
      "title": "Stating a determinism budget for a control loop",
      "summary": "Before writing loop code, fix four numbers: sample rate, allowed jitter on the sampling instant, worst-case actuation latency, and the deadline-miss policy. Every downstream choice then checks against the budget.",
      "content_html": "<p>Most real-time control problems are argued in adjectives — “fast enough”,\n“low jitter”, “hard real-time” — and adjectives cannot be verified. A determinism\nbudget replaces them with four numbers agreed before implementation starts. After\nthat, every architectural question has a right answer: does option A fit inside\nthe budget or not.</p>\n<p>This page is about writing that budget. It is protocol- and silicon-agnostic on\npurpose; the numbers change per project, the structure does not.</p>\n<h2 id=\"the-four-numbers\">The four numbers</h2>\n<h3 id=\"1-sample-rate-f_s\">1. Sample rate (<code>f_s</code>)</h3>\n<p>The rate at which the loop reads its inputs and updates its output. Set it from\nthe plant, not from what the CPU can do:</p>\n<ul>\n<li><strong>Rule of thumb</strong>: <code>f_s</code> between 10× and 20× the closed-loop bandwidth you\nneed. Below 10× the phase lag from sampling eats your phase margin; far above\n20× you are burning CPU and amplifying sensor noise through the derivative term\nfor no dynamic benefit.</li>\n<li>For a mechanical position loop with a 50 Hz bandwidth target, that is roughly\n1–2 kHz. For a motor current loop it is typically 8–20 kHz because the\nelectrical time constant is short. For a temperature loop, 1–10 Hz is often\nplenty and a faster loop just wastes power.</li>\n<li>Write down the reasoning, not just the number. The next engineer needs to know\nwhether 2 kHz was a plant requirement or a guess.</li>\n</ul>\n<h3 id=\"2-sampling-jitter-δt_s\">2. Sampling jitter (<code>Δt_s</code>)</h3>\n<p>The allowed variation in <em>when</em> the input is actually sampled, relative to the\nideal period <code>1/f_s</code>. This is the number most specs omit and most loops get bitten\nby.</p>\n<p>Jitter matters because a control law assumes a fixed <code>Δt</code>. If the real interval\nwanders by ±15% and the code still divides by the nominal <code>Δt</code> in the derivative\nand integral terms, you have injected a disturbance proportional to the jitter and\nthe signal slew rate. Effects:</p>\n<ul>\n<li>The <code>I</code> term accumulates the wrong area.</li>\n<li>The <code>D</code> term produces spikes on jittered intervals.</li>\n<li>At the control-bandwidth frequency, timing jitter aliases into the loop as\nbroadband noise you cannot filter out without also hurting the response.</li>\n</ul>\n<p>Budget it as a percentage of the period and as an absolute time. “≤ 2% of period\nor ≤ 5 µs, whichever is larger” is a typical starting point for a mid-rate loop.\nFor a current loop synchronised to PWM, the sampling instant is usually locked to\na timer/PWM trigger and an ADC hardware trigger, and the jitter budget there is\ntens of nanoseconds — a software-timer-driven sample cannot meet it and the\nbudget is what tells you that up front.</p>\n<p>Two mitigations worth stating in the budget itself:</p>\n<ul>\n<li><strong>Timestamp every sample</strong> and feed the <em>actual</em> <code>Δt</code> into the loop maths, so\njitter degrades gracefully instead of injecting noise.</li>\n<li><strong>Trigger sampling in hardware</strong> (timer-to-ADC, no CPU in the path) and let the\nISR only <em>consume</em> the result.</li>\n</ul>\n<h3 id=\"3-worst-case-actuation-latency-l_wc\">3. Worst-case actuation latency (<code>L_wc</code>)</h3>\n<p>The time from “the sampling instant” to “the new actuator command is in effect” —\nend to end, worst case, not typical. Its components:</p>\n<pre class=\"astro-code github-dark\" style=\"background-color:#24292e;color:#e1e4e8;overflow-x:auto\" tabindex=\"0\" data-language=\"plaintext\"><code><span class=\"line\"><span>L_wc = t_acq        ADC conversion + settling</span></span>\n<span class=\"line\"><span>     + t_dispatch   interrupt latency + scheduler to the loop task</span></span>\n<span class=\"line\"><span>     + t_compute    the control law, worst-case path</span></span>\n<span class=\"line\"><span>     + t_output     DAC/PWM update, or the bus transaction to a remote drive</span></span>\n<span class=\"line\"><span>     + t_actuator   the actuator&#39;s own transport lag (often the biggest term)</span></span></code></pre>\n<p>Rules for filling it in:</p>\n<ul>\n<li>Use <strong>worst case</strong> for every term. Interrupt latency under maximum interrupt\nload, <code>t_compute</code> with the branch that runs the anti-windup and the fault\nchecks, bus latency including retransmission if the link allows it.</li>\n<li><code>L_wc</code> shows up in the loop as pure dead time. Dead time destroys phase margin\nfast: as a guide, keep <code>L_wc</code> under about 1/10 of the loop period, and treat\nanything above 1/4 of the period as a redesign trigger.</li>\n<li>If the actuator is across a fieldbus, the bus cycle time and its determinism\nare now inside your control budget. A 1 kHz loop commanding a drive over a\n1 ms-cycle bus has spent its entire latency budget on transport before any\ncontrol happens. This is the calculation that decides whether the loop runs\nlocal to the drive or on a central controller.</li>\n</ul>\n<h3 id=\"4-deadline-miss-policy\">4. Deadline-miss policy</h3>\n<p>Define what a missed deadline <em>is</em> and what the system does about it. “It should\nnot happen” is not a policy.</p>\n<ul>\n<li><strong>What counts as a miss</strong>: loop iteration N has not completed before iteration\nN+1’s release. Instrument it — a GPIO toggled at loop start/end on a scope, or\na software counter of overruns exported to diagnostics.</li>\n<li><strong>Allowed miss rate</strong>: for a soft loop, “≤ 1 in 10⁶ iterations and never two\nconsecutive” might be fine. For a loop tied to a safety function, the answer is\nusually zero within the safety analysis and the response is a defined safe\nstate, not a retry.</li>\n<li><strong>The response</strong>: hold last output, ramp to a safe value, trip a fault, or\nextrapolate one step. Each has failure modes — holding last output during a\nfast transient can be worse than a brief zero. State the choice and why.</li>\n<li><strong>Escalation</strong>: what happens on the 2nd, 10th, 100th consecutive miss.</li>\n</ul>\n<h2 id=\"what-the-budget-lets-you-decide-without-arguing\">What the budget lets you decide without arguing</h2>\n<p>Once the four numbers exist, these stop being opinions:</p>\n<table>\n<thead>\n<tr>\n<th>Question</th>\n<th>Decided by</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Bare-metal, RTOS, or Linux with <code>PREEMPT_RT</code>?</td>\n<td>Can its worst-case interrupt-to-task latency + scheduler jitter fit inside <code>Δt_s</code> and the <code>t_dispatch</code> share of <code>L_wc</code>, measured, under load?</td>\n</tr>\n<tr>\n<td>Loop in an ISR or a task?</td>\n<td>If <code>t_compute</code> is far below the period and the jitter budget is tight, use the ISR. If <code>t_compute</code> is a meaningful fraction of the period, use a task so it can be preempted by faster ISRs.</td>\n</tr>\n<tr>\n<td>Control local to the actuator or central?</td>\n<td>Does the bus transport lag fit inside <code>L_wc</code>?</td>\n</tr>\n<tr>\n<td>Which fieldbus?</td>\n<td>Its cycle time and cycle-time jitter versus <code>Δt_s</code> and <code>L_wc</code>.</td>\n</tr>\n<tr>\n<td>Is this core allowed to run anything else?</td>\n<td>Only if the other work’s worst-case interference still leaves the budget intact — usually meaning a shielded/isolated core or a dedicated MCU.</td>\n</tr>\n</tbody>\n</table>\n<h2 id=\"measuring-against-it\">Measuring against it</h2>\n<p>A budget you cannot measure is a wish. The standard instrumentation:</p>\n<ul>\n<li><strong>GPIO + scope/logic analyser</strong>: toggle a pin at ISR entry, at loop-math start,\nat output write. Persistence mode on the scope shows the jitter envelope\ndirectly. This is the ground truth; trust it over any software timestamp.</li>\n<li><strong>Cycle counter</strong> (<code>DWT-&gt;CYCCNT</code> on Cortex-M, <code>PMCCNTR</code> / <code>perf</code> on\nCortex-A) around the compute path, logging min/max/histogram, not just mean.</li>\n<li><strong>Overrun counter</strong> in the loop, exported to the same diagnostics channel as\neverything else so field units report it.</li>\n<li><strong>Soak under worst-case load</strong>: every interrupt source firing, the comms stack\nsaturated, the file system busy, cache cold. The typical case is not the\nnumber in the budget.</li>\n</ul>\n<p>Hold the measurement for long enough to see the tail. Latency distributions in\nreal systems have long tails driven by rare cache/TLB/bus-contention events, and\nthe 99.99th percentile is often 3–10× the median. The budget is about the tail,\nso the measurement has to reach it.</p>\n<h2 id=\"the-honest-trade-off\">The honest trade-off</h2>\n<p>Writing this budget costs a day or two of analysis and a negotiation with whoever\nowns the plant requirements, before any loop code exists. On a schedule under\npressure that day is tempting to skip, and the loop will usually “work” on the\nbench without it.</p>\n<p>What you lose by skipping it is the ability to say <em>why</em> it works, and the ability\nto catch — at design time rather than during integration — the cases where the\nchosen bus or OS cannot meet the timing. Those cases are expensive exactly in\nproportion to how late they are found. The budget moves that discovery to the\ncheapest possible point.</p>\n<p>A tighter budget is not free either. Driving jitter to nanoseconds means hardware\ntriggering, a shielded core, and no shared bus — real BOM and integration cost.\nThe budget should be as loose as the plant genuinely allows, and no looser.</p>\n",
      "image": "https://exubits.com/engineering/og/stating-a-determinism-budget.png",
      "date_published": "2026-09-03T00:00:00.000Z",
      "date_modified": "2026-09-03T00:00:00.000Z",
      "authors": [
        {
          "name": "Exubits Engineering",
          "url": "https://exubits.com/authors/exubits-engineering"
        }
      ],
      "tags": [
        "Reference",
        "determinism",
        "real-time",
        "closed-loop"
      ]
    },
    {
      "id": "https://exubits.com/engineering/yocto-bsp-handover",
      "url": "https://exubits.com/engineering/yocto-bsp-handover",
      "title": "What a Yocto BSP hand-over should actually contain",
      "summary": "A usable Yocto BSP hand-over: a pinned manifest, vendored layers, a documented DISTRO/MACHINE, a reproducible build, a signed image with its SPDX SBOM, and the rationale for every kernel and U-Boot config choice.",
      "content_html": "<p>A Yocto BSP hand-over fails in a predictable way. Six months after the last\ninvoice, someone runs the build on a fresh machine and it breaks: a layer moved,\n<code>meta-openembedded</code> advanced a branch, an SRC_URI 404s, the host GCC is now too\nnew for a fetched tarball. The image that shipped is fine. The ability to\n<em>rebuild</em> it is gone. That is the thing a hand-over has to protect, and most\nhand-overs do not.</p>\n<p>This page is the checklist we hold our own BSP deliveries to. It is deliberately\nabout artefacts and their properties, not about a particular board.</p>\n<h2 id=\"the-one-property-that-matters-bit-for-bit-rebuildability\">The one property that matters: bit-for-bit rebuildability</h2>\n<p>Everything below is in service of a single test. Take the delivery, put it on a\nmachine that has never seen the project, follow the written steps, and get an\nimage whose package manifest matches the one that shipped. If that test passes,\nthe BSP is maintainable. If it does not, you have bought a binary with source\nattached, not a BSP.</p>\n<p>Yocto gives you the machinery for this — <code>BB_HASHSERVE</code>, hash-equivalence,\n<code>buildhistory</code>, reproducible-builds class — but none of it is on by a default\nthat survives a vendor change. It has to be configured, and the configuration is\npart of the deliverable.</p>\n<h2 id=\"1-a-pinned-self-contained-source-manifest\">1. A pinned, self-contained source manifest</h2>\n<ul>\n<li><strong>A <code>repo</code> manifest or a kas file</strong> that pins every layer to a <strong>commit SHA</strong>,\nnot a branch. <code>honister</code>, <code>kirkstone</code>, <code>scarthgap</code> — a branch name is a moving\ntarget. <code>kirkstone</code> today is not <code>kirkstone</code> from the release date.</li>\n<li><strong>The BitBake and OE-Core revision</strong> pinned in the same file.</li>\n<li><strong>No layer fetched from a URL the client does not control.</strong> If a layer lives\nonly on a vendor’s GitLab, it is a single point of failure with someone else’s\nuptime. Vendored into the client’s own Git, with the upstream URL and SHA\nrecorded in the commit message so the provenance is not lost.</li>\n<li>The <code>conf/bblayers.conf</code> <code>BBLAYERS</code> order documented, because layer priority\nchanges which <code>.bbappend</code> wins.</li>\n</ul>\n<p>The test: <code>kas checkout</code> (or <code>repo sync</code>) on an air-gapped mirror reproduces the\nexact tree. No “then update meta-freescale to the tip”.</p>\n<h2 id=\"2-your-own-layer-and-only-your-changes-in-it\">2. Your own layer, and only your changes in it</h2>\n<p>There should be exactly one layer that contains the project’s work —\n<code>meta-&lt;project&gt;</code> — and it should be the only layer with local commits. Every\nchange to a BSP or upstream recipe lives there as a <code>.bbappend</code> or a versioned\nrecipe copy, never as an edit to <code>meta-ti</code>, <code>meta-freescale</code>, or <code>poky</code>.</p>\n<p>Why this is a hand-over issue and not a style preference: when the client later\nmoves from <code>kirkstone</code> to the next LTS, the migration work is <em>reviewing one\nlayer</em>. If changes are smeared across five vendor layers, the migration is\narchaeology, and the estimate for it triples.</p>\n<p>What belongs in <code>meta-&lt;project&gt;</code>:</p>\n<ul>\n<li>The machine <code>.conf</code> (or a <code>.bbappend</code> to the vendor’s) with every deviation\ncommented — why <code>PREFERRED_VERSION_linux-*</code> is pinned, why a <code>MACHINE_FEATURE</code>\nwas removed.</li>\n<li>The image recipe(s). One production image, one development image, and the\ndelta between them stated explicitly (dev adds <code>debug-tweaks</code>, <code>openssh-sftp</code>,\n<code>gdbserver</code> — and production must be verified <em>not</em> to).</li>\n<li>The distro config if the project defines its own <code>DISTRO</code>. It usually should:\ninheriting <code>poky</code> and overriding twelve variables in <code>local.conf</code> means the\nconfig only exists on the machine that has that <code>local.conf</code>.</li>\n</ul>\n<h2 id=\"3-localconf-is-not-part-of-the-delivery--the-distro-config-is\">3. <code>local.conf</code> is not part of the delivery — the distro config is</h2>\n<p><code>local.conf</code> is per-developer scratch. Anything load-bearing that lives there\nwill be lost. The hand-over must move every meaningful setting into version\ncontrol:</p>\n<ul>\n<li><code>DISTRO_FEATURES</code> / <code>DISTRO_FEATURES_remove</code> — particularly <code>systemd</code> vs\n<code>sysvinit</code>, <code>wayland</code>/<code>x11</code>, <code>pam</code>, <code>usrmerge</code>.</li>\n<li><code>IMAGE_FSTYPES</code> and how they map to what the factory actually flashes\n(<code>wic.gz</code>, <code>wic.bmap</code>, a <code>.swu</code> for SWUpdate).</li>\n<li><code>EXTRA_IMAGE_FEATURES</code>, <code>IMAGE_INSTALL:append</code>.</li>\n<li>The <code>PACKAGE_CLASSES</code> choice (<code>package_rpm</code>/<code>ipk</code>/<code>deb</code>) — this affects the\non-target update mechanism and cannot be changed casually later.</li>\n<li>Any <code>PREMIRRORS</code> / <code>SSTATE_MIRRORS</code> pointing at internal infrastructure.</li>\n</ul>\n<h2 id=\"4-the-kernel-a-defconfig-fragment-and-a-reason-for-every-line\">4. The kernel: a defconfig fragment and a reason for every line</h2>\n<p>A kernel handed over as a 6,000-line <code>.config</code> is not maintainable, because\nnobody can tell an essential setting from an accident of <code>make oldconfig</code>.</p>\n<ul>\n<li><strong><code>defconfig</code> plus fragments</strong>, wired through <code>KERNEL_CONFIG_FRAGMENTS</code> or a\n<code>linux-*.bbappend</code>. The fragment is small and every line is intentional.</li>\n<li><strong>A written rationale</strong> for the non-obvious ones: which <code>CONFIG_</code> enables the\nEthernet PHY, which sets the RT behaviour, which was needed for a USB gadget\nmode, which disables an unused subsystem to cut attack surface and boot time.</li>\n<li><strong>The kernel provenance</strong>: mainline version, the vendor SoC tree it is based\non, the patch stack applied on top — as a quilt series or Git history, not a\nsingle squashed diff. When a CVE lands, someone needs to know whether the tree\nalready carries the fix.</li>\n<li><strong>Out-of-tree modules</strong> identified, with their licence and their source. An\nout-of-tree Wi-Fi driver with no upstream is a maintenance liability that\nshould be named in the hand-over, not discovered later.</li>\n</ul>\n<h2 id=\"5-device-tree-the-boards-hardware-description-reviewed\">5. Device tree: the board’s hardware description, reviewed</h2>\n<p>The device tree is where “the BSP works” and “the BSP is correct” diverge. A\nnode can be missing and the board still boots.</p>\n<ul>\n<li>The board <code>.dts</code> in <code>meta-&lt;project&gt;</code>, including via <code>.dtsi</code>, not patched into\nthe vendor’s file in place.</li>\n<li>Every pinmux group traceable to the schematic net name. A comment linking\n<code>MX8MM_IOMUXC_SD2_CD_B_GPIO2_IO12</code> to the actual card-detect net saves the next\nengineer an afternoon with a multimeter.</li>\n<li>Regulators modelled properly — <code>regulator-always-on</code>, <code>regulator-boot-on</code> and\nthe supply chain to each peripheral. Half of “the peripheral randomly does not\nenumerate” bugs are a missing or lazy regulator description.</li>\n<li>Overlays, if used, with the base-plus-overlay combination that the running\nsystem actually uses documented. Applied by U-Boot or by the kernel — state\nwhich.</li>\n</ul>\n<h2 id=\"6-u-boot-the-config-the-environment-and-the-boot-flow\">6. U-Boot: the config, the environment, and the boot flow</h2>\n<ul>\n<li>The U-Boot defconfig and any board patches, same discipline as the kernel.</li>\n<li><strong>The boot environment as a file</strong>, not as whatever is currently in the SPI\nflash of the one golden board. The <code>boot.cmd</code>/<code>boot.scr</code> or the extlinux\nconfig, in version control.</li>\n<li>The <strong>boot flow written out</strong>: ROM → SPL → U-Boot proper → kernel → init,\nwith where each stage lives (eMMC boot partition, offset, GPT) and what the\nfallback path is if the primary kernel does not come up.</li>\n<li>If secure boot is in play: which keys sign what, where the public keys are\nfused or stored, and — critically — a documented recovery path for a board\nwhose signature check fails. A secure-boot BSP with no recovery story is a\nbrick generator.</li>\n</ul>\n<h2 id=\"7-the-signed-image-and-its-bill-of-materials\">7. The signed image and its bill of materials</h2>\n<ul>\n<li>The exact image that was released, plus its <strong>signature</strong> and the public key\nto verify it.</li>\n<li><strong><code>buildhistory</code> output</strong> committed for that build: the package list with\nversions, image size, dependency graph, and the diff against the previous\nrelease. This is what makes “what changed between v1.2 and v1.3” answerable in\nminutes.</li>\n<li>An <strong>SBOM</strong> — Yocto emits SPDX 2.2 JSON via <code>create-spdx</code> — covering every\npackage in the image with its version and licence. Increasingly this is a\ncontractual and regulatory requirement (the EU Cyber Resilience Act among\nthem), and it is far cheaper to generate at build time than to reconstruct.</li>\n<li>The <strong>licence manifest</strong> (<code>license.manifest</code>, <code>deploy/licenses/</code>) and the\nsources for anything under a copyleft licence, or a working <code>bitbake -c archiver</code> configuration that regenerates them.</li>\n</ul>\n<h2 id=\"8-the-build-environment-itself\">8. The build environment itself</h2>\n<p>“Works on my machine” is not a hand-over. Pin the host too:</p>\n<ul>\n<li>A <strong>container</strong> (the <a href=\"https://github.com/crops/poky-container\">CROPS</a> base or\na project Dockerfile) that fixes the host distro, the essential host packages,\nPython, and locale. Yocto is sensitive to host GCC and host Python; a 2024\nbuild host and a 2027 build host are not equivalent.</li>\n<li>The exact <code>bitbake</code> invocation, target names, and any <code>BB_ENV_PASSTHROUGH</code>\nadditions.</li>\n<li>Expected build resources: disk for <code>TMPDIR</code> and <code>SSTATE_DIR</code>, RAM, rough wall\ntime on a stated machine — so the client can size CI.</li>\n<li>A <strong>populated sstate and downloads mirror</strong>, or instructions to build one.\nWithout it the first rebuild pulls hundreds of source archives from the\ninternet, and some of those URLs will be dead. An offline <code>DL_DIR</code> archive is\nthe single most valuable non-obvious artefact in the box.</li>\n</ul>\n<h2 id=\"9-documentation-that-is-task-shaped\">9. Documentation that is task-shaped</h2>\n<p>Not a wiki dump. Five documents, each answering a question someone will actually\nask:</p>\n<ol>\n<li><strong>Build it</strong> — from empty machine to flashable image, copy-paste commands.</li>\n<li><strong>Flash it</strong> — factory path and field path, including the recovery/USB-download\nprocedure for a bricked board.</li>\n<li><strong>Change the kernel config / add a package / bump a recipe</strong> — the routine\nmaintenance loop.</li>\n<li><strong>Cut a release</strong> — versioning, signing, what gets tagged, what gets archived.</li>\n<li><strong>Known issues and deferred work</strong> — the honest list. Every BSP has one; a\nhand-over without it just means the client finds them the hard way.</li>\n</ol>\n<h2 id=\"what-this-costs\">What this costs</h2>\n<p>Producing this is real work — on the order of one to two engineer-weeks on top of\nthe BSP itself, most of it in items 1, 7, and 8. It is worth being explicit about\nthat in a statement of work rather than discovering it at the end. The payoff is\nentirely deferred: it shows up the first time the client rebuilds without the\noriginal team, and it is the difference between an afternoon and a re-engagement.</p>\n<p>The trade-off we make deliberately: we do <em>not</em> hand over a build that tracks\nupstream branches, even though that would look more “up to date” on delivery day.\nA pinned build is frozen and will drift out of security currency until someone\ndeliberately bumps it — that is the cost. It is the right cost, because a build\nthat silently changes under the client is not a build they own.</p>\n",
      "image": "https://exubits.com/engineering/og/yocto-bsp-handover.png",
      "date_published": "2026-09-03T00:00:00.000Z",
      "date_modified": "2026-09-03T00:00:00.000Z",
      "authors": [
        {
          "name": "Exubits Engineering",
          "url": "https://exubits.com/authors/exubits-engineering"
        }
      ],
      "tags": [
        "Reference",
        "yocto",
        "bitbake",
        "kernel",
        "device-tree"
      ]
    }
  ]
}