<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"
     xmlns:content="http://purl.org/rss/1.0/modules/content/"
     xmlns:dc="http://purl.org/dc/elements/1.1/"
     xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Exubits Engineering</title>
    <link>https://exubits.com/engineering</link>
    <atom:link href="https://exubits.com/rss.xml" rel="self" type="application/rss+xml" />
    <description>Reference pages, deep dives, teardowns and field notes on embedded product engineering.</description>
    <language>en-IN</language>
    <lastBuildDate>Thu, 03 Sep 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>Choosing between CAN, CAN FD and Ethernet industrial links</title>
      <link>https://exubits.com/engineering/can-canfd-ethernet-industrial-links</link>
      <guid isPermaLink="true">https://exubits.com/engineering/can-canfd-ethernet-industrial-links</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Exubits Engineering</dc:creator>
      <category>Reference</category>
      <description>CAN suits cheap, rugged, low-rate event traffic; CAN FD when payloads outgrew 8 bytes but a bus still fits; industrial Ethernet (PROFINET, EtherCAT, TSN) for megabits or tight cycle times. Expect to bridge, not replace.</description>
      <content:encoded><![CDATA[<p>This decision gets made badly in two directions. Teams keep classic CAN because
“it has always worked” and then spend a year fighting bus-load and 8-byte frames.
Or they jump to industrial Ethernet everywhere and discover it costs more per
node, needs managed switches, and still does not give them determinism unless
they configured it for it.</p>
<p>Here is how the three actually differ, and what each is bad at.</p>
<h2 id="classic-can-iso-11898-1-2">Classic CAN (ISO 11898-1/-2)</h2>
<p><strong>What it is</strong>: a multi-drop differential bus, up to 1 Mbit/s, 8 data bytes per
frame, non-destructive bitwise arbitration by identifier, strong error detection
(15-bit CRC, bit monitoring, form/stuff checks) with automatic retransmission.</p>
<p><strong>What it is genuinely good at</strong>:</p>
<ul>
<li><strong>Cost and ruggedness.</strong> A CAN transceiver is cents. The bus is two wires, no
switch, no hub, tolerant of ground shift and EMI, and it degrades gracefully —
a node dropping off does not take the segment down.</li>
<li><strong>Event-driven traffic with priority.</strong> Arbitration means the lowest-numbered
ID always wins the bus with no collisions and no configuration. For “report
this alarm now” traffic that is exactly right.</li>
<li><strong>Determinism at low load.</strong> Below roughly 30–40% bus load, worst-case latency
for a given priority is calculable and small.</li>
<li><strong>Ecosystem.</strong> CANopen, J1939, and decades of tooling, diagnostics, and
engineers who know it.</li>
</ul>
<p><strong>What it is bad at</strong> — state these before choosing it:</p>
<ul>
<li><strong>8 bytes.</strong> A 12-axis sensor packet or any structured record needs
fragmentation (ISO-TP), which adds latency, state, and failure modes.</li>
<li><strong>1 Mbit/s ceiling</strong>, and you only reach it on a short bus. At 1 Mbit/s the
practical bus length is on the order of 40 m; longer buses force lower rates
because arbitration needs the signal to propagate to the far end within a bit.</li>
<li><strong>Bus load cliff.</strong> Above ~50–60% load, low-priority frames see large and
hard-to-bound latency, and a chatty node can starve them entirely.</li>
<li><strong>No native security.</strong> No authentication, no encryption; any node can send any
ID. Mitigations exist (bus guardians, MACsec-style add-ons, segmentation) but
they are bolt-ons.</li>
<li><strong>A single retransmitting faulty node</strong> can dominate the bus until error
confinement kicks in.</li>
</ul>
<h2 id="can-fd-iso-11898-12015">CAN FD (ISO 11898-1:2015)</h2>
<p><strong>What it changes</strong>: up to <strong>64 data bytes</strong> per frame, and a <strong>second, faster
bit rate</strong> for the data phase (commonly 2–5 Mbit/s, 8 Mbit/s achievable with
care) while the arbitration phase stays at the classic rate. Better CRC (17- or
21-bit) for the larger payload.</p>
<p><strong>Why it is often the right upgrade</strong>:</p>
<ul>
<li><strong>The 8-byte problem disappears</strong> without fragmentation. One 64-byte frame
instead of eight classic frames means less protocol overhead and less software
state.</li>
<li><strong>Effective throughput rises several-fold</strong> for payload-heavy traffic, because
both the payload is bigger and the data phase is faster.</li>
<li><strong>You keep the topology, the transceivers’ ruggedness, the arbitration model,
and most of the tooling.</strong> It is an incremental change, not a re-architecture.</li>
<li>CANopen FD and J1939 variants exist to carry the ecosystem forward.</li>
</ul>
<p><strong>What it is bad at / what to watch</strong>:</p>
<ul>
<li><strong>The arbitration phase is still limited to the classic bit rate.</strong> CAN FD does
not raise the number of <em>frames per second</em> you can arbitrate; it raises the
bytes per frame. A system that is frame-rate-bound, not payload-bound, gets
little from it.</li>
<li><strong>Signal integrity gets harder.</strong> The fast data phase shrinks bit times to
hundreds of nanoseconds; ringing, stub length, connector quality, and
transceiver loop delay now matter. A messy bus that tolerated 500 kbit/s
classic will not tolerate 5 Mbit/s FD without topology work and possibly ring
or point-to-point rather than long stubs.</li>
<li><strong>Mixed classic/FD segments need care.</strong> A classic-only node on an FD bus sees
FD frames as errors and floods the bus unless partial networking / FD-tolerant
transceivers are used. Usually the whole segment must be FD.</li>
<li><strong>Controller support</strong>: older MCUs have classic-only CAN peripherals. Check the
silicon, not the datasheet marketing.</li>
</ul>
<h2 id="switched-industrial-ethernet--profinet-ethercat-and-tsn">Switched industrial Ethernet — PROFINET, EtherCAT, and TSN</h2>
<p>Not one thing. “Industrial Ethernet” covers protocols with very different
determinism stories, all on 100 Mbit/s or 1 Gbit/s PHYs.</p>
<ul>
<li><strong>PROFINET RT</strong> — standard switched Ethernet, prioritised frames, cycle times
down to ~1 ms; <strong>PROFINET IRT</strong> adds hardware time-slotting for sub-millisecond,
low-jitter cycles but needs IRT-capable switches.</li>
<li><strong>EtherCAT</strong> — a single frame passes through every node, each reading and
writing its slice on the fly (“processing on the fly”). Extremely efficient,
cycle times to tens of microseconds, but it is a specific master/slave topology
with EtherCAT slave controllers in each node, not general IP networking.</li>
<li><strong>TSN</strong> (IEEE 802.1 set — 802.1AS, Qbv, Qci, CB, …) — brings scheduled,
bounded-latency traffic to <em>standard</em> Ethernet, letting control traffic and IT
traffic share the same infrastructure. It is the direction the industry is
moving; it also requires TSN-capable switches and endpoints and non-trivial
network engineering (a schedule computed and pushed to every bridge).</li>
</ul>
<p><strong>What Ethernet buys you</strong>:</p>
<ul>
<li><strong>Bandwidth</strong>: 100–1000 Mbit/s, so video, firmware images, and high-channel-
count data are no longer a problem.</li>
<li><strong>Routable, IT-integratable</strong>: TCP/IP, TLS, MQTT/OPC UA to the cloud on the
same wire (with segmentation and a firewall).</li>
<li><strong>Determinism <em>if configured for it</em></strong>: IRT, EtherCAT, or TSN give bounded
latency and low jitter — but plain “Ethernet” does not; a standard switch under
load has queueing delay and no guarantees.</li>
</ul>
<p><strong>What it is bad at</strong>:</p>
<ul>
<li><strong>Cost and complexity per node</strong>: a PHY, magnetics or connector, often a
managed or special switch, more powerful MCU/SoC, and a stack. Multiples of a
CAN node.</li>
<li><strong>Topology and cabling</strong>: star or line with switches, 100 m segment limit per
hop, more connectors, more to go wrong mechanically in a vibrating machine.</li>
<li><strong>Determinism is opt-in and fragile to misconfiguration.</strong> A well-meaning IT
change, a non-TSN device plugged into a TSN segment, or an unmanaged switch
dropped in “temporarily” can quietly break the timing guarantees.</li>
<li><strong>Security surface</strong>: now you are on an IP network, with everything that
implies. IEC 62443 becomes your problem.</li>
</ul>
<h2 id="a-decision-order-that-works">A decision order that works</h2>
<ol>
<li><strong>What is the largest single payload, and how often?</strong> ≤ 8 bytes, event-rate
→ classic CAN is still fine. 8–64 bytes → CAN FD. Kilobytes, or any streaming
→ Ethernet.</li>
<li><strong>What is the tightest cycle time with bounded jitter you must hit?</strong>
Above 5–10 ms and modest, CAN/CAN FD or PROFINET RT. Around 1 ms, PROFINET IRT
or CAN FD on a lightly loaded bus. Under 250 µs with many axes, EtherCAT. Mixed
critical and best-effort on one wire, TSN.</li>
<li><strong>Node count, cost target, and environment.</strong> Dozens of cheap rugged nodes on
a harness → bus wins. A cell with a handful of high-value devices and a video
feed → Ethernet.</li>
<li><strong>Does it need to talk to IT / cloud directly?</strong> Yes and at rate → Ethernet.
Occasionally → keep the fieldbus and put the IP stack in the gateway.</li>
<li><strong>Who maintains it for ten years?</strong> A plant electrician can fault-find a CAN
bus with a multimeter. A TSN schedule needs someone who understands 802.1Qbv.</li>
</ol>
<h2 id="the-realistic-answer-is-usually-bridge">The realistic answer is usually “bridge”</h2>
<p>Greenfield-everything-Ethernet is rare. The common shape is CAN or CAN FD at the
machine edge for sensors and drives, aggregated by a <strong>gateway</strong> that speaks the
fieldbus on one side and OPC UA / MQTT over Ethernet on the other, doing protocol
translation, buffering across link outages, and store-and-forward. That gateway is
also the right place to put the security boundary, the data-model mapping, and the
firmware-update path — rather than pushing all of that cost into every edge node.</p>
<p>Designing that seam well — which data is published at what rate, what is buffered,
how backpressure and reconnection behave — is usually more of the engineering
effort than picking the physical layer.</p>
]]></content:encoded>
    </item>
    <item>
      <title>What a production hand-over pack has to contain</title>
      <link>https://exubits.com/engineering/production-handover-pack</link>
      <guid isPermaLink="true">https://exubits.com/engineering/production-handover-pack</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Exubits Engineering</dc:creator>
      <category>Reference</category>
      <description>It lets someone else build, test, certify and change the product without the original team: design data, a reproducible firmware build, factory tests with limits, the compliance file, an SBOM, and known issues.</description>
      <content:encoded><![CDATA[<p>A hand-over pack has one job: let a competent engineer who was not on the project
build the product, prove it works, keep it compliant, support it in the field, and
change it — without phoning the original team. Anything less is a dependency, not
a delivery.</p>
<p>This is the contents list we work to. It spans hardware, firmware, test, and
compliance, because a product that splits those across separate incomplete
hand-overs has no owner for the seams.</p>
<h2 id="1-released-versioned-design-data">1. Released, versioned design data</h2>
<p>Not “the latest files” — a <strong>release</strong>, tagged, with a revision and a date.</p>
<p><strong>Hardware</strong></p>
<ul>
<li>Schematics (PDF) and source project, at the released revision.</li>
<li>PCB layout source, Gerbers, drill, netlist, fab drawing with stack-up and
impedance spec, assembly drawings.</li>
<li><strong>BOM</strong> with manufacturer part numbers, approved alternates, and DNP items
marked. Lifecycle status noted for anything NRND or single-source.</li>
<li>Mechanical: enclosure models, gaskets, thermal interface parts, fasteners,
labels and their artwork.</li>
<li>The <strong>ECO/ECR history</strong> — every change since the first build, why it was made,
and which build it entered. This is how the factory knows a v3 board is not a
v2 board with a sticker.</li>
</ul>
<p><strong>Firmware and software</strong></p>
<ul>
<li>Source, in a repo the client controls, tagged for the release.</li>
<li>A <strong>reproducible build</strong>: pinned toolchain (container or documented exact
versions), pinned dependencies, one command. The test is that a fresh machine
reproduces the released binary — or the released image manifest — bit for bit.
(For a Yocto-based product, this is its own checklist; see the Embedded Linux
hub’s BSP hand-over reference.)</li>
<li>The released binaries themselves, with checksums and signatures, and the public
keys to verify them.</li>
<li>Build and flashing instructions from empty machine to programmed unit.</li>
</ul>
<h2 id="2-requirements-and-traceability">2. Requirements and traceability</h2>
<ul>
<li>The <strong>requirements</strong> the product was built to, at their released version.</li>
<li>A <strong>traceability matrix</strong>: requirement → design element → verification test →
result. Even a lightweight one. This is what a customer audit, a safety
assessor, or a future change-impact analysis needs, and it cannot be
reconstructed later without re-deriving intent.</li>
<li>Interface control documents for every external interface (connector pinouts,
protocols, message definitions, electrical limits).</li>
</ul>
<h2 id="3-verification-and-validation-records">3. Verification and validation records</h2>
<ul>
<li><strong>Test plans and procedures</strong>, with pass/fail criteria that are numbers, not
adjectives.</li>
<li><strong>Results</strong> for the released revision: what was run, on how many units, with
what outcome, signed and dated. Raw data retained, not just a summary.</li>
<li><strong>Coverage statement</strong>: what was tested, what was explicitly not, and why. A
hand-over that implies everything was verified when it was not is worse than
one that is honest about the gaps.</li>
<li>Regression suite and how to run it, so the client can re-verify after a change.</li>
</ul>
<h2 id="4-compliance--certification-file">4. Compliance / certification file</h2>
<p>Whatever regimes apply (EMC, safety, radio, environmental, sector-specific):</p>
<ul>
<li><strong>Test reports</strong> from the accredited lab, in full, not just the certificate.</li>
<li>The <strong>Declaration of Conformity</strong> and the technical file / construction file
behind it.</li>
<li>The <strong>exact configuration tested</strong>: hardware revision, firmware version,
cables, orientation, peripherals, any test-mode firmware used. A certificate
that does not pin the configuration does not transfer to production cleanly.</li>
<li><strong>Applied standards and their editions.</strong> Standard editions change; the file
must say which one the product was assessed against.</li>
<li><strong>Conditions and limitations</strong> of the approval, and the class/limits the
product passed against (with margin, ideally).</li>
<li>What invalidates the certification — which changes force re-test. This is the
single most useful sentence for the team that inherits it.</li>
<li>Radio module certs (modular approval IDs) and the conditions attached to using
them under that approval.</li>
</ul>
<h2 id="5-software-bill-of-materials-and-update-story">5. Software bill of materials and update story</h2>
<ul>
<li><strong>SBOM</strong> (SPDX or CycloneDX) for the shipped image: every component, version,
licence, and source. Increasingly a legal requirement, and the basis for
answering “are we affected by CVE-XXXX” quickly.</li>
<li><strong>Licence compliance</strong>: the licence manifest, written offers where copyleft
requires them, and the corresponding source archive.</li>
<li><strong>Secure/verified boot</strong>: key hierarchy, what signs what, where keys are stored
or fused, key-rotation procedure, and the recovery path for a unit that fails a
signature check.</li>
<li><strong>Field update mechanism</strong>: how an update is packaged, signed, delivered,
applied, and rolled back; A/B or recovery-slot behaviour; what happens on power
loss mid-update; version-compatibility rules.</li>
<li><strong>Provisioning</strong>: per-unit secrets, certificates, serial numbers — how they are
generated, injected at the factory, and stored/backed up.</li>
</ul>
<h2 id="6-manufacturing-pack">6. Manufacturing pack</h2>
<ul>
<li><strong>Factory test procedure</strong> with a defined sequence, measurement points, and
<strong>limits</strong> for every measured parameter. This is what makes a contract
manufacturer able to accept or reject a unit without judgement calls.</li>
<li><strong>Test fixtures</strong>: design, firmware, calibration procedure and interval, and a
spare or the data to build one.</li>
<li><strong>Programming</strong>: what gets programmed, in what order, with what tool, and how
the line verifies it took.</li>
<li><strong>Calibration</strong>: what is calibrated, against what reference, to what tolerance,
and where the calibration constants are stored.</li>
<li><strong>Serialisation and labelling</strong>: numbering scheme, label content and placement,
what is recorded per unit and where that record lives.</li>
<li><strong>Packaging and ESD/transport</strong> requirements.</li>
<li><strong>First-article inspection</strong> report and the golden-unit definition.</li>
<li><strong>Yield and known failure modes</strong> from the builds so far, with the diagnosis
for each — so the line does not re-learn them.</li>
</ul>
<h2 id="7-field-support-material">7. Field support material</h2>
<ul>
<li><strong>Service manual</strong>: diagnosis flow, safe-to-replace parts, what needs a
depot, recovery procedures (including “how to un-brick”).</li>
<li><strong>Diagnostics</strong>: what the product logs, how to extract it, and how to read it.</li>
<li><strong>RMA criteria</strong> and the data to capture on a return.</li>
<li><strong>Spares strategy</strong>: which parts, expected failure rates if known, last-time-buy
exposure on any component.</li>
</ul>
<h2 id="8-the-known-issues-and-deferred-work-list">8. The known-issues and deferred-work list</h2>
<p>Every project has one. A hand-over without it means the client discovers the list
one incident at a time.</p>
<ul>
<li>Open bugs with severity, reproduction, and any workaround.</li>
<li>Deferred features and the reason they were cut.</li>
<li>Design compromises made under schedule pressure and what a proper fix looks
like.</li>
<li>Component or supplier risks not yet resolved.</li>
<li>Anything that “works but we do not fully understand why”.</li>
</ul>
<h2 id="9-contacts-licences-and-accounts">9. Contacts, licences, and accounts</h2>
<ul>
<li>Which third-party licences, subscriptions, or accounts the product depends on
(cloud tenant, certificate authority, code-signing service, paid libraries),
who holds them, and renewal dates.</li>
<li>Escrow arrangements if any.</li>
<li>A named transition contact and a defined support window, so the hand-over is a
ramp, not a cliff.</li>
</ul>
<h2 id="how-to-know-the-pack-is-complete">How to know the pack is complete</h2>
<p>Run the test that matters: give the pack to an engineer who was not on the
project and have them, using only what is in the box, (a) build the firmware and
get a matching binary, (b) build and pass one unit through the factory test, and
(c) answer “what change would invalidate the CE marking”. If any of the three
needs a phone call, the pack has a hole, and the hole is cheaper to fill now.</p>
<h2 id="the-trade-off">The trade-off</h2>
<p>Assembling this properly is weeks of work, concentrated at the end of a project
when budget and patience are thinnest, and most of it produces no visible feature.
It is genuinely tempting to ship the design files and a README and call the rest
“support”.</p>
<p>The cost of that shortcut is entirely borne later and by someone else: the first
CM build that fails because the test limits were in someone’s head, the CVE
response that takes three weeks because there is no SBOM, the field failure with
no service procedure, the change that quietly invalidates a certification because
nobody wrote down what the certification depended on. The pack is insurance, and
like insurance its value is invisible right up until the moment it is the only
thing that helps.</p>
]]></content:encoded>
    </item>
    <item>
      <title>What a schematic review should catch before layout starts</title>
      <link>https://exubits.com/engineering/schematic-review-before-layout</link>
      <guid isPermaLink="true">https://exubits.com/engineering/schematic-review-before-layout</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Exubits Engineering</dc:creator>
      <category>Reference</category>
      <description>A schematic review catches the errors that get far more expensive after layout: power sequencing and budgets, reset and strapping, every net&apos;s return path, connector pinout and ESD, test access, and BOM risk.</description>
      <content:encoded><![CDATA[<p>The schematic review is the last cheap checkpoint. A missed pull-up found here is
a two-minute edit. Found after layout it is a re-route; found after fab it is a
bodge wire and a respin; found in the field it is a recall. The review’s purpose
is to spend an hour now to not spend a month later.</p>
<p>This is the checklist we run schematics against before releasing them to layout.
It assumes the design is functionally “done” — the point is to find what “done”
missed.</p>
<h2 id="power-the-section-that-causes-the-most-respins">Power: the section that causes the most respins</h2>
<h3 id="rail-inventory-and-budget">Rail inventory and budget</h3>
<ul>
<li><strong>List every rail</strong>: voltage, tolerance, estimated and worst-case current,
source (which regulator), and every load on it. A spreadsheet, not a mental
model.</li>
<li><strong>Check each regulator against its worst-case load</strong> including inrush and
transient, not typical. Add margin for the load estimate being wrong — 30% is
common at this stage.</li>
<li><strong>Thermal</strong>: <code>(Vin − Vout) × I</code> for every linear regulator. An LDO dropping
3.3 V at 500 mA is dissipating 1.65 W and needs a heatsinking copper plan that
layout has to know about <em>now</em>.</li>
<li><strong>Efficiency and input current</strong>: does the upstream supply / connector / fuse
actually deliver the sum of input currents at low line?</li>
</ul>
<h3 id="sequencing">Sequencing</h3>
<ul>
<li><strong>Does the SoC/FPGA have a required power-up and power-down order?</strong> Most do
(core before I/O, or a specific ramp relationship). Violating it can cause
latch-up or excess current through internal diodes. Check the datasheet’s
sequencing section and confirm the design enforces it — enable-pin daisy
chains, a sequencer IC, or a supervisor.</li>
<li><strong>Power-down order matters too</strong>, and is more often missed. A rail collapsing
in the wrong order can back-drive an I/O bank.</li>
<li><strong>What happens on a brown-out</strong> — does the sequence restart cleanly, or can it
hang half-powered?</li>
</ul>
<h3 id="decoupling-as-schematic-intent">Decoupling, as schematic intent</h3>
<ul>
<li>Every power pin has decoupling specified with values and count per the device
datasheet. Bulk capacitance per rail sized for the load step.</li>
<li>It is fine that placement is layout’s job — but the schematic must show the
<em>intent</em> so the reviewer can check nothing is missing and layout knows the
target.</li>
</ul>
<h2 id="reset-clocks-and-strapping">Reset, clocks, and strapping</h2>
<ul>
<li><strong>Every reset</strong>: source, polarity, pull resistor, RC or supervisor timing,
and which devices it reaches. Open-drain resets wired together need one pull-up,
not none and not five.</li>
<li><strong>Reset supervisor threshold</strong> matched to the SoC’s minimum operating voltage,
with hysteresis, so a sagging rail does not leave the part running out of spec.</li>
<li><strong>Watchdog</strong>: present, wired to actually reset the system, and not defeatable
by the failure it is meant to catch.</li>
<li><strong>Boot strapping / mode pins</strong>: this is a classic post-layout disaster. Every
strap pin identified, its required level confirmed against the boot-mode table,
and the resistor value chosen so it wins against the pin’s other function
(often the same pin is a functional I/O with its own loading). Note which straps
need to be <em>changeable</em> for bring-up and give them a header or a 0 Ω option.</li>
<li><strong>Clocks</strong>: every oscillator/crystal has the load caps per the crystal spec and
the oscillator’s drive level checked. PLL supplies filtered. Spread-spectrum
choice made deliberately (it helps EMC, it can hurt a camera or ADC).</li>
</ul>
<h2 id="every-net-reference-and-return-path">Every net: reference and return path</h2>
<p>Signal integrity is decided in the schematic by what you connect to what, before
a single trace exists.</p>
<ul>
<li><strong>For each interface, name the return-current path.</strong> A differential pair
referenced to a plane that has a split under it will radiate and fail EMC — and
the fix is a stitching-cap or plane change that is far easier to plan now.</li>
<li><strong>Series termination / source termination</strong> resistors placed in the schematic
for fast single-ended nets (RGMII, parallel memory, fast GPIO). Value TBD by
layout, but the footprint must exist.</li>
<li><strong>Differential pairs</strong> (USB, Ethernet, MIPI, PCIe, LVDS): correct AC coupling
where the standard requires it, correct common-mode termination, correct
polarity, and pairs that the connector pinout does not force to cross.</li>
<li><strong>Unused inputs tied off</strong>, not floating. Unused outputs left open. Unused
op-amp sections wired as followers to a mid-rail, not left open-loop.</li>
<li><strong>Level shifting</strong> wherever two voltage domains meet — confirm direction,
speed, and that the translator’s supported data rate exceeds the bus.</li>
</ul>
<h2 id="connectors-and-the-outside-world">Connectors and the outside world</h2>
<p>Connectors are where field failures enter.</p>
<ul>
<li><strong>Pinout sanity</strong>: power and ground pins adjacent enough to carry the current;
hot-plug order (ground first, then power, then signal) if the connector is
ever mated live.</li>
<li><strong>ESD/EOS protection on every externally accessible pin</strong>: USB, Ethernet
(plus the magnetics and Bob Smith termination), buttons, connector I/O.
TVS diodes with the right stand-off and clamping voltage, placed at the
connector.</li>
<li><strong>Reverse-polarity and overvoltage</strong> protection on the power input, sized for
the real worst case (a field tech with a 24 V supply on a 12 V input).</li>
<li><strong>Miswire survival</strong>: what happens if the harness is plugged in shifted by one
pin? On rugged products this is worth a deliberate answer.</li>
<li><strong>Test points</strong> on every rail, reset, boot strap, and key signal — see DFT.</li>
</ul>
<h2 id="design-for-test-and-manufacture">Design for test and manufacture</h2>
<ul>
<li><strong>DFT</strong>: bed-of-nails or flying-probe access to rails and critical nets. A
JTAG/SWD header (even if depopulated in production). A UART console broken out.
A way to hold the board in reset and in each boot mode. Programming access for
every programmable device, and the programming order/method noted.</li>
<li><strong>DFM</strong>: no parts on both sides that force two reflow passes unless intended;
package choices the assembler can place; no 0201s where 0402 would do; fiducials
present; polarity markings that survive assembly.</li>
<li><strong>First-article bring-up plan</strong> implied by the schematic: can you bring rails up
one at a time? Is there a “populate R1, leave R2 off” staged power-on option?</li>
</ul>
<h2 id="bom-and-component-risk">BOM and component risk</h2>
<ul>
<li><strong>Lifecycle</strong>: any part NRND or single-source? Flag it now; a second source or
a footprint that accepts alternates is a schematic decision.</li>
<li><strong>Ratings with margin</strong>: capacitor voltage derating (ceramics lose capacitance
with DC bias — a 6.3 V X5R on a 5 V rail may be at half its rated value),
resistor power, inductor saturation current at the real peak, MOSFET SOA.</li>
<li><strong>Tolerance stack-up</strong> on anything that sets a threshold, a timing, or a
feedback divider.</li>
<li><strong>Passives count</strong>: every DNP is intentional and labelled; every “we’ll tune
this in bring-up” has a real footprint and a starting value.</li>
</ul>
<h2 id="how-to-actually-run-it">How to actually run it</h2>
<ul>
<li><strong>Two reviewers minimum</strong>, at least one who did not draw the schematic.</li>
<li><strong>Page by page, net by net on the critical interfaces</strong>, out loud, against the
device datasheets open on the table — not a glance-through.</li>
<li><strong>Every datasheet’s “layout guidelines” and “design checklist” section read</strong>,
because vendors put the expensive mistakes there.</li>
<li><strong>Track findings as a list with owners and states.</strong> The review is not done
when the meeting ends; it is done when the list is closed and re-checked.</li>
<li><strong>Re-review after the fixes.</strong> Fixes introduce errors.</li>
</ul>
<h2 id="the-trade-off">The trade-off</h2>
<p>A proper schematic review on a medium-complexity board is the better part of a
day for two engineers, plus the fix cycle — call it two to three engineer-days
before layout even starts. It feels like a delay when the schematic “looks
finished”.</p>
<p>What it does not do: it will not catch problems that only exist in the physical
layout — coupling, plane resonance, actual trace impedance, thermal reality,
mechanical fit. Those need a layout review and, ultimately, measurement on
hardware. The schematic review is necessary and cheap; it is not sufficient, and
treating a passed schematic review as “the hard part is over” is its own failure
mode.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Where the milliseconds go between sensor and display</title>
      <link>https://exubits.com/engineering/sensor-to-display-latency</link>
      <guid isPermaLink="true">https://exubits.com/engineering/sensor-to-display-latency</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Exubits Engineering</dc:creator>
      <category>Reference</category>
      <description>Glass-to-glass latency is exposure + sensor readout + CSI-2 + ISP + buffer and queue waits + composition + one or two display refreshes. Most of it is exposure and buffering, not code; shrink queues and sync to vsync.</description>
      <content:encoded><![CDATA[<p>“The camera feels laggy” is a latency-budget problem, and it is almost always
lost in places people do not instrument: exposure time, the number of buffers in
each queue, and the wait for the next display refresh. Optimising the application
code is usually the smallest available win.</p>
<p>This page walks the whole path from photons to photons and gives the order of
magnitude of each stage, so you know which term to attack.</p>
<h2 id="the-pipeline-stage-by-stage">The pipeline, stage by stage</h2>
<h3 id="1-exposure-integration-time">1. Exposure (integration time)</h3>
<p>The sensor integrates light for the exposure time. The <em>photons that form a
frame</em> are spread across that whole window, so the effective latency contribution
is roughly <strong>half the exposure time</strong> for motion (the centroid of the exposure),
but the frame is not readable until the full exposure ends, so for end-to-end
timing you count the <strong>full exposure time</strong>.</p>
<ul>
<li>Bright scene, short exposure (1–2 ms): negligible.</li>
<li>Indoor / auto-exposure (8–33 ms): this is often the single largest term, and it
is invisible in any software profiler.</li>
<li>Low light with a 1/30 s exposure: 33 ms before anything else has happened.</li>
</ul>
<p>Lever: cap the maximum auto-exposure time and accept more gain (noise) if latency
matters more than SNR. This is an ISP/3A tuning decision, not a code change.</p>
<h3 id="2-sensor-readout">2. Sensor readout</h3>
<p>The pixel array is read out row by row and streamed. Readout takes on the order
of <strong>one frame period</strong> at the sensor’s current mode — a rolling shutter sensor
at 60 fps takes ~16 ms to clock the whole frame out, and the bottom of the frame
is genuinely ~16 ms “younger” than the top.</p>
<ul>
<li>Higher frame-rate modes read out faster (and usually crop or bin).</li>
<li>Global shutter removes the top-to-bottom skew but not the readout transport
time.</li>
<li>Lever: run the sensor faster than the display rate if the SoC can take it —
reading a 60 fps stream for a 30 fps display halves this term and the exposure
cap becomes easier to hit.</li>
</ul>
<h3 id="3-mipi-csi-2-transport">3. MIPI CSI-2 transport</h3>
<p>Serialised over the D-PHY/C-PHY lanes into the SoC’s CSI receiver. At any sane
lane rate this is <strong>sub-millisecond per frame</strong> for typical resolutions — it is
almost never the problem. It shows up only if lanes are marginal and you are
getting retimed/corrupted lines that force retry or drop.</p>
<h3 id="4-receiver-dma-to-memory">4. Receiver DMA to memory</h3>
<p>The CSI receiver writes frames into a ring of buffers in DRAM. Two things here:</p>
<ul>
<li>The write itself is fast (bounded by memory bandwidth, sub-ms to low-ms).</li>
<li><strong>The number of buffers in this ring is a latency knob.</strong> More buffers =
smoother under jitter, but a frame can sit in the ring for up to
<code>(buffers − 1) × frame_period</code> before anything consumes it. Four buffers at
30 fps is up to 100 ms of potential sit time. This is the classic hidden lag.</li>
</ul>
<h3 id="5-isp">5. ISP</h3>
<p>Debayer, black level, lens shading, denoise, sharpening, tone mapping, colour
conversion, scaling. On a hardware ISP this is <strong>a few milliseconds</strong> and often
pipelined with readout. In software (libcamera soft ISP, OpenCV) it can be tens
of milliseconds and steals CPU from everything else.</p>
<ul>
<li>3A (auto-exposure/white-balance/focus) runs here and feeds <em>the next</em> frame —
it adds a frame of loop latency to convergence, not to the display path, but a
slow-converging AE means longer exposures for longer, which loops back to
stage 1.</li>
<li>Lever: use the hardware ISP path (V4L2 media controller / libcamera with the
vendor pipeline handler) rather than a software fallback.</li>
</ul>
<h3 id="6-application--buffer-handoff">6. Application / buffer handoff</h3>
<p>The frame becomes a <code>dmabuf</code> handed to whatever draws it — a Qt <code>QVideoSink</code>, a
GStreamer <code>appsink</code>/<code>glimagesink</code>, an LVGL canvas, a Weston/Wayland client, a
direct DRM/KMS plane.</p>
<ul>
<li><strong>Every queue between elements is <code>depth × frame_period</code> of potential latency.</strong>
GStreamer’s default queue sizes, a <code>v4l2src</code> with many buffers, a compositor
that triple-buffers — each is a place frames wait.</li>
<li><strong>Zero-copy or not.</strong> If the frame is memcpy’d (or worse, uploaded to a GL
texture) at each stage, that is memory-bandwidth time and CPU stalls, a few ms
each and more under load. <code>dmabuf</code> import all the way to the display plane
avoids it.</li>
<li>Lever: shortest possible pipeline. A camera frame on its own DRM/KMS overlay
plane, composited by the display controller, skips GPU composition entirely.</li>
</ul>
<h3 id="7-composition">7. Composition</h3>
<p>If the UI draws the video into a scene (overlays, HUD, controls), the GPU
composites. One frame period at the render rate, plus the GPU’s own latency.
Using a hardware overlay plane for the video and only compositing the UI chrome
avoids putting the video through this stage.</p>
<h3 id="8-display-refresh-and-panel-response">8. Display refresh and panel response</h3>
<ul>
<li>The scanout waits for the next <strong>vsync</strong> to start sending the new frame:
0 to one refresh period of wait (average half). At 60 Hz that is up to 16.7 ms.</li>
<li>Double/triple buffering adds <strong>one or two more refresh periods</strong> depending on
whether a frame misses its flip deadline.</li>
<li>The panel’s own response: LCD pixel response and its internal frame buffer /
overdrive processing, commonly <strong>one frame</strong> for an LCD, sometimes more for a
panel with a scaler or “image enhancement” it will not let you disable.</li>
<li>Lever: <code>PAGE_FLIP</code> synced to vsync with exactly the buffers you need (often
double, not triple), and pick a panel/timing controller with a documented,
low, fixed latency. Disable panel-side processing.</li>
</ul>
<h2 id="adding-it-up">Adding it up</h2>
<p>A representative indoor 30 fps preview, no tuning:</p>
<table>
<thead>
<tr>
<th>Stage</th>
<th>Typical</th>
</tr>
</thead>
<tbody>
<tr>
<td>Exposure</td>
<td>15–33 ms</td>
</tr>
<tr>
<td>Readout</td>
<td>~16–33 ms</td>
</tr>
<tr>
<td>CSI-2 + DMA</td>
<td>1–3 ms</td>
</tr>
<tr>
<td>Buffer ring wait (4 deep)</td>
<td>0–100 ms</td>
</tr>
<tr>
<td>ISP (hardware)</td>
<td>2–6 ms</td>
</tr>
<tr>
<td>Pipeline queues (GStreamer defaults)</td>
<td>30–100 ms</td>
</tr>
<tr>
<td>GPU composition</td>
<td>~16–33 ms</td>
</tr>
<tr>
<td>Vsync wait + buffering</td>
<td>16–50 ms</td>
</tr>
<tr>
<td>Panel</td>
<td>8–33 ms</td>
</tr>
</tbody>
</table>
<p>That easily reaches <strong>150–250 ms</strong> — the “laggy” complaint — and only ~5 ms of it
is anything a code profiler would show you.</p>
<p>The same pipeline, tuned:</p>
<ul>
<li>Sensor at 60 fps, AE capped at 8 ms.</li>
<li>Two CSI buffers, not four.</li>
<li>Hardware ISP path.</li>
<li>No GStreamer queues, or <code>leaky=downstream max-size-buffers=1</code>.</li>
<li>Video on a dedicated KMS overlay plane, UI composited separately.</li>
<li>Double-buffered page flip synced to vsync, panel processing off.</li>
</ul>
<p>lands in the <strong>40–70 ms</strong> range, and most of what remains is exposure, readout,
and one refresh period — the irreducible physics.</p>
<h2 id="how-to-measure-it">How to measure it</h2>
<p>Do not trust per-stage estimates; measure glass-to-glass:</p>
<ul>
<li><strong>Photodiode + LED + scope.</strong> Flash an LED in frame, detect it with a
photodiode taped to the display, measure LED-on to display-brightens on the
scope. This is the only number that counts.</li>
<li><strong>On-screen millisecond counter filmed with a 120–240 fps camera</strong> alongside a
physical stopwatch/timer in the same shot; count frames of offset.</li>
<li><strong>Per-stage</strong>: V4L2 buffer timestamps (<code>v4l2_buffer.timestamp</code>),
<code>GST_DEBUG</code> with <code>GST_TRACERS=latency</code>, DRM/KMS <code>PAGE_FLIP</code> event timestamps,
and a GPIO toggle at each app-level handoff. Line them up against the
glass-to-glass number to find the fat stage.</li>
</ul>
<h2 id="the-trade-offs-you-are-actually-making">The trade-offs you are actually making</h2>
<ul>
<li><strong>Fewer buffers = lower latency, less tolerance to scheduling jitter.</strong> Drop a
deadline with two buffers and you get a visible stutter; with four you get lag
but no stutter. Pick per product — a welding HUD wants low latency, a
security-review monitor wants no dropped frames.</li>
<li><strong>Shorter exposure = lower latency, more noise.</strong> You are trading SNR for
responsiveness, frame by frame, in the AE tuning.</li>
<li><strong>Overlay plane = low latency, less compositing flexibility.</strong> Overlay planes
have format, scaling, and count limits set by the display controller; a complex
UI that must blend with the video may not fit and has to go through the GPU.</li>
<li><strong>Running the sensor faster = lower latency and easier AE, more MIPI and memory
bandwidth, more power and heat.</strong> On a thermally constrained enclosure that can
be the deciding constraint.</li>
</ul>
]]></content:encoded>
    </item>
    <item>
      <title>Stating a determinism budget for a control loop</title>
      <link>https://exubits.com/engineering/stating-a-determinism-budget</link>
      <guid isPermaLink="true">https://exubits.com/engineering/stating-a-determinism-budget</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Exubits Engineering</dc:creator>
      <category>Reference</category>
      <description>Before writing loop code, fix four numbers: sample rate, allowed jitter on the sampling instant, worst-case actuation latency, and the deadline-miss policy. Every downstream choice then checks against the budget.</description>
      <content:encoded><![CDATA[<p>Most real-time control problems are argued in adjectives — “fast enough”,
“low jitter”, “hard real-time” — and adjectives cannot be verified. A determinism
budget replaces them with four numbers agreed before implementation starts. After
that, every architectural question has a right answer: does option A fit inside
the budget or not.</p>
<p>This page is about writing that budget. It is protocol- and silicon-agnostic on
purpose; the numbers change per project, the structure does not.</p>
<h2 id="the-four-numbers">The four numbers</h2>
<h3 id="1-sample-rate-f_s">1. Sample rate (<code>f_s</code>)</h3>
<p>The rate at which the loop reads its inputs and updates its output. Set it from
the plant, not from what the CPU can do:</p>
<ul>
<li><strong>Rule of thumb</strong>: <code>f_s</code> between 10× and 20× the closed-loop bandwidth you
need. Below 10× the phase lag from sampling eats your phase margin; far above
20× you are burning CPU and amplifying sensor noise through the derivative term
for no dynamic benefit.</li>
<li>For a mechanical position loop with a 50 Hz bandwidth target, that is roughly
1–2 kHz. For a motor current loop it is typically 8–20 kHz because the
electrical time constant is short. For a temperature loop, 1–10 Hz is often
plenty and a faster loop just wastes power.</li>
<li>Write down the reasoning, not just the number. The next engineer needs to know
whether 2 kHz was a plant requirement or a guess.</li>
</ul>
<h3 id="2-sampling-jitter-δt_s">2. Sampling jitter (<code>Δt_s</code>)</h3>
<p>The allowed variation in <em>when</em> the input is actually sampled, relative to the
ideal period <code>1/f_s</code>. This is the number most specs omit and most loops get bitten
by.</p>
<p>Jitter matters because a control law assumes a fixed <code>Δt</code>. If the real interval
wanders by ±15% and the code still divides by the nominal <code>Δt</code> in the derivative
and integral terms, you have injected a disturbance proportional to the jitter and
the signal slew rate. Effects:</p>
<ul>
<li>The <code>I</code> term accumulates the wrong area.</li>
<li>The <code>D</code> term produces spikes on jittered intervals.</li>
<li>At the control-bandwidth frequency, timing jitter aliases into the loop as
broadband noise you cannot filter out without also hurting the response.</li>
</ul>
<p>Budget it as a percentage of the period and as an absolute time. “≤ 2% of period
or ≤ 5 µs, whichever is larger” is a typical starting point for a mid-rate loop.
For a current loop synchronised to PWM, the sampling instant is usually locked to
a timer/PWM trigger and an ADC hardware trigger, and the jitter budget there is
tens of nanoseconds — a software-timer-driven sample cannot meet it and the
budget is what tells you that up front.</p>
<p>Two mitigations worth stating in the budget itself:</p>
<ul>
<li><strong>Timestamp every sample</strong> and feed the <em>actual</em> <code>Δt</code> into the loop maths, so
jitter degrades gracefully instead of injecting noise.</li>
<li><strong>Trigger sampling in hardware</strong> (timer-to-ADC, no CPU in the path) and let the
ISR only <em>consume</em> the result.</li>
</ul>
<h3 id="3-worst-case-actuation-latency-l_wc">3. Worst-case actuation latency (<code>L_wc</code>)</h3>
<p>The time from “the sampling instant” to “the new actuator command is in effect” —
end to end, worst case, not typical. Its components:</p>
<pre class="astro-code github-dark" style="background-color:#24292e;color:#e1e4e8;overflow-x:auto" tabindex="0" data-language="plaintext"><code><span class="line"><span>L_wc = t_acq        ADC conversion + settling</span></span>
<span class="line"><span>     + t_dispatch   interrupt latency + scheduler to the loop task</span></span>
<span class="line"><span>     + t_compute    the control law, worst-case path</span></span>
<span class="line"><span>     + t_output     DAC/PWM update, or the bus transaction to a remote drive</span></span>
<span class="line"><span>     + t_actuator   the actuator&#39;s own transport lag (often the biggest term)</span></span></code></pre>
<p>Rules for filling it in:</p>
<ul>
<li>Use <strong>worst case</strong> for every term. Interrupt latency under maximum interrupt
load, <code>t_compute</code> with the branch that runs the anti-windup and the fault
checks, bus latency including retransmission if the link allows it.</li>
<li><code>L_wc</code> shows up in the loop as pure dead time. Dead time destroys phase margin
fast: as a guide, keep <code>L_wc</code> under about 1/10 of the loop period, and treat
anything above 1/4 of the period as a redesign trigger.</li>
<li>If the actuator is across a fieldbus, the bus cycle time and its determinism
are now inside your control budget. A 1 kHz loop commanding a drive over a
1 ms-cycle bus has spent its entire latency budget on transport before any
control happens. This is the calculation that decides whether the loop runs
local to the drive or on a central controller.</li>
</ul>
<h3 id="4-deadline-miss-policy">4. Deadline-miss policy</h3>
<p>Define what a missed deadline <em>is</em> and what the system does about it. “It should
not happen” is not a policy.</p>
<ul>
<li><strong>What counts as a miss</strong>: loop iteration N has not completed before iteration
N+1’s release. Instrument it — a GPIO toggled at loop start/end on a scope, or
a software counter of overruns exported to diagnostics.</li>
<li><strong>Allowed miss rate</strong>: for a soft loop, “≤ 1 in 10⁶ iterations and never two
consecutive” might be fine. For a loop tied to a safety function, the answer is
usually zero within the safety analysis and the response is a defined safe
state, not a retry.</li>
<li><strong>The response</strong>: hold last output, ramp to a safe value, trip a fault, or
extrapolate one step. Each has failure modes — holding last output during a
fast transient can be worse than a brief zero. State the choice and why.</li>
<li><strong>Escalation</strong>: what happens on the 2nd, 10th, 100th consecutive miss.</li>
</ul>
<h2 id="what-the-budget-lets-you-decide-without-arguing">What the budget lets you decide without arguing</h2>
<p>Once the four numbers exist, these stop being opinions:</p>
<table>
<thead>
<tr>
<th>Question</th>
<th>Decided by</th>
</tr>
</thead>
<tbody>
<tr>
<td>Bare-metal, RTOS, or Linux with <code>PREEMPT_RT</code>?</td>
<td>Can its worst-case interrupt-to-task latency + scheduler jitter fit inside <code>Δt_s</code> and the <code>t_dispatch</code> share of <code>L_wc</code>, measured, under load?</td>
</tr>
<tr>
<td>Loop in an ISR or a task?</td>
<td>If <code>t_compute</code> is far below the period and the jitter budget is tight, use the ISR. If <code>t_compute</code> is a meaningful fraction of the period, use a task so it can be preempted by faster ISRs.</td>
</tr>
<tr>
<td>Control local to the actuator or central?</td>
<td>Does the bus transport lag fit inside <code>L_wc</code>?</td>
</tr>
<tr>
<td>Which fieldbus?</td>
<td>Its cycle time and cycle-time jitter versus <code>Δt_s</code> and <code>L_wc</code>.</td>
</tr>
<tr>
<td>Is this core allowed to run anything else?</td>
<td>Only if the other work’s worst-case interference still leaves the budget intact — usually meaning a shielded/isolated core or a dedicated MCU.</td>
</tr>
</tbody>
</table>
<h2 id="measuring-against-it">Measuring against it</h2>
<p>A budget you cannot measure is a wish. The standard instrumentation:</p>
<ul>
<li><strong>GPIO + scope/logic analyser</strong>: toggle a pin at ISR entry, at loop-math start,
at output write. Persistence mode on the scope shows the jitter envelope
directly. This is the ground truth; trust it over any software timestamp.</li>
<li><strong>Cycle counter</strong> (<code>DWT-&gt;CYCCNT</code> on Cortex-M, <code>PMCCNTR</code> / <code>perf</code> on
Cortex-A) around the compute path, logging min/max/histogram, not just mean.</li>
<li><strong>Overrun counter</strong> in the loop, exported to the same diagnostics channel as
everything else so field units report it.</li>
<li><strong>Soak under worst-case load</strong>: every interrupt source firing, the comms stack
saturated, the file system busy, cache cold. The typical case is not the
number in the budget.</li>
</ul>
<p>Hold the measurement for long enough to see the tail. Latency distributions in
real systems have long tails driven by rare cache/TLB/bus-contention events, and
the 99.99th percentile is often 3–10× the median. The budget is about the tail,
so the measurement has to reach it.</p>
<h2 id="the-honest-trade-off">The honest trade-off</h2>
<p>Writing this budget costs a day or two of analysis and a negotiation with whoever
owns the plant requirements, before any loop code exists. On a schedule under
pressure that day is tempting to skip, and the loop will usually “work” on the
bench without it.</p>
<p>What you lose by skipping it is the ability to say <em>why</em> it works, and the ability
to catch — at design time rather than during integration — the cases where the
chosen bus or OS cannot meet the timing. Those cases are expensive exactly in
proportion to how late they are found. The budget moves that discovery to the
cheapest possible point.</p>
<p>A tighter budget is not free either. Driving jitter to nanoseconds means hardware
triggering, a shielded core, and no shared bus — real BOM and integration cost.
The budget should be as loose as the plant genuinely allows, and no looser.</p>
]]></content:encoded>
    </item>
    <item>
      <title>What a Yocto BSP hand-over should actually contain</title>
      <link>https://exubits.com/engineering/yocto-bsp-handover</link>
      <guid isPermaLink="true">https://exubits.com/engineering/yocto-bsp-handover</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Exubits Engineering</dc:creator>
      <category>Reference</category>
      <description>A usable Yocto BSP hand-over: a pinned manifest, vendored layers, a documented DISTRO/MACHINE, a reproducible build, a signed image with its SPDX SBOM, and the rationale for every kernel and U-Boot config choice.</description>
      <content:encoded><![CDATA[<p>A Yocto BSP hand-over fails in a predictable way. Six months after the last
invoice, someone runs the build on a fresh machine and it breaks: a layer moved,
<code>meta-openembedded</code> advanced a branch, an SRC_URI 404s, the host GCC is now too
new for a fetched tarball. The image that shipped is fine. The ability to
<em>rebuild</em> it is gone. That is the thing a hand-over has to protect, and most
hand-overs do not.</p>
<p>This page is the checklist we hold our own BSP deliveries to. It is deliberately
about artefacts and their properties, not about a particular board.</p>
<h2 id="the-one-property-that-matters-bit-for-bit-rebuildability">The one property that matters: bit-for-bit rebuildability</h2>
<p>Everything below is in service of a single test. Take the delivery, put it on a
machine that has never seen the project, follow the written steps, and get an
image whose package manifest matches the one that shipped. If that test passes,
the BSP is maintainable. If it does not, you have bought a binary with source
attached, not a BSP.</p>
<p>Yocto gives you the machinery for this — <code>BB_HASHSERVE</code>, hash-equivalence,
<code>buildhistory</code>, reproducible-builds class — but none of it is on by a default
that survives a vendor change. It has to be configured, and the configuration is
part of the deliverable.</p>
<h2 id="1-a-pinned-self-contained-source-manifest">1. A pinned, self-contained source manifest</h2>
<ul>
<li><strong>A <code>repo</code> manifest or a kas file</strong> that pins every layer to a <strong>commit SHA</strong>,
not a branch. <code>honister</code>, <code>kirkstone</code>, <code>scarthgap</code> — a branch name is a moving
target. <code>kirkstone</code> today is not <code>kirkstone</code> from the release date.</li>
<li><strong>The BitBake and OE-Core revision</strong> pinned in the same file.</li>
<li><strong>No layer fetched from a URL the client does not control.</strong> If a layer lives
only on a vendor’s GitLab, it is a single point of failure with someone else’s
uptime. Vendored into the client’s own Git, with the upstream URL and SHA
recorded in the commit message so the provenance is not lost.</li>
<li>The <code>conf/bblayers.conf</code> <code>BBLAYERS</code> order documented, because layer priority
changes which <code>.bbappend</code> wins.</li>
</ul>
<p>The test: <code>kas checkout</code> (or <code>repo sync</code>) on an air-gapped mirror reproduces the
exact tree. No “then update meta-freescale to the tip”.</p>
<h2 id="2-your-own-layer-and-only-your-changes-in-it">2. Your own layer, and only your changes in it</h2>
<p>There should be exactly one layer that contains the project’s work —
<code>meta-&lt;project&gt;</code> — and it should be the only layer with local commits. Every
change to a BSP or upstream recipe lives there as a <code>.bbappend</code> or a versioned
recipe copy, never as an edit to <code>meta-ti</code>, <code>meta-freescale</code>, or <code>poky</code>.</p>
<p>Why this is a hand-over issue and not a style preference: when the client later
moves from <code>kirkstone</code> to the next LTS, the migration work is <em>reviewing one
layer</em>. If changes are smeared across five vendor layers, the migration is
archaeology, and the estimate for it triples.</p>
<p>What belongs in <code>meta-&lt;project&gt;</code>:</p>
<ul>
<li>The machine <code>.conf</code> (or a <code>.bbappend</code> to the vendor’s) with every deviation
commented — why <code>PREFERRED_VERSION_linux-*</code> is pinned, why a <code>MACHINE_FEATURE</code>
was removed.</li>
<li>The image recipe(s). One production image, one development image, and the
delta between them stated explicitly (dev adds <code>debug-tweaks</code>, <code>openssh-sftp</code>,
<code>gdbserver</code> — and production must be verified <em>not</em> to).</li>
<li>The distro config if the project defines its own <code>DISTRO</code>. It usually should:
inheriting <code>poky</code> and overriding twelve variables in <code>local.conf</code> means the
config only exists on the machine that has that <code>local.conf</code>.</li>
</ul>
<h2 id="3-localconf-is-not-part-of-the-delivery--the-distro-config-is">3. <code>local.conf</code> is not part of the delivery — the distro config is</h2>
<p><code>local.conf</code> is per-developer scratch. Anything load-bearing that lives there
will be lost. The hand-over must move every meaningful setting into version
control:</p>
<ul>
<li><code>DISTRO_FEATURES</code> / <code>DISTRO_FEATURES_remove</code> — particularly <code>systemd</code> vs
<code>sysvinit</code>, <code>wayland</code>/<code>x11</code>, <code>pam</code>, <code>usrmerge</code>.</li>
<li><code>IMAGE_FSTYPES</code> and how they map to what the factory actually flashes
(<code>wic.gz</code>, <code>wic.bmap</code>, a <code>.swu</code> for SWUpdate).</li>
<li><code>EXTRA_IMAGE_FEATURES</code>, <code>IMAGE_INSTALL:append</code>.</li>
<li>The <code>PACKAGE_CLASSES</code> choice (<code>package_rpm</code>/<code>ipk</code>/<code>deb</code>) — this affects the
on-target update mechanism and cannot be changed casually later.</li>
<li>Any <code>PREMIRRORS</code> / <code>SSTATE_MIRRORS</code> pointing at internal infrastructure.</li>
</ul>
<h2 id="4-the-kernel-a-defconfig-fragment-and-a-reason-for-every-line">4. The kernel: a defconfig fragment and a reason for every line</h2>
<p>A kernel handed over as a 6,000-line <code>.config</code> is not maintainable, because
nobody can tell an essential setting from an accident of <code>make oldconfig</code>.</p>
<ul>
<li><strong><code>defconfig</code> plus fragments</strong>, wired through <code>KERNEL_CONFIG_FRAGMENTS</code> or a
<code>linux-*.bbappend</code>. The fragment is small and every line is intentional.</li>
<li><strong>A written rationale</strong> for the non-obvious ones: which <code>CONFIG_</code> enables the
Ethernet PHY, which sets the RT behaviour, which was needed for a USB gadget
mode, which disables an unused subsystem to cut attack surface and boot time.</li>
<li><strong>The kernel provenance</strong>: mainline version, the vendor SoC tree it is based
on, the patch stack applied on top — as a quilt series or Git history, not a
single squashed diff. When a CVE lands, someone needs to know whether the tree
already carries the fix.</li>
<li><strong>Out-of-tree modules</strong> identified, with their licence and their source. An
out-of-tree Wi-Fi driver with no upstream is a maintenance liability that
should be named in the hand-over, not discovered later.</li>
</ul>
<h2 id="5-device-tree-the-boards-hardware-description-reviewed">5. Device tree: the board’s hardware description, reviewed</h2>
<p>The device tree is where “the BSP works” and “the BSP is correct” diverge. A
node can be missing and the board still boots.</p>
<ul>
<li>The board <code>.dts</code> in <code>meta-&lt;project&gt;</code>, including via <code>.dtsi</code>, not patched into
the vendor’s file in place.</li>
<li>Every pinmux group traceable to the schematic net name. A comment linking
<code>MX8MM_IOMUXC_SD2_CD_B_GPIO2_IO12</code> to the actual card-detect net saves the next
engineer an afternoon with a multimeter.</li>
<li>Regulators modelled properly — <code>regulator-always-on</code>, <code>regulator-boot-on</code> and
the supply chain to each peripheral. Half of “the peripheral randomly does not
enumerate” bugs are a missing or lazy regulator description.</li>
<li>Overlays, if used, with the base-plus-overlay combination that the running
system actually uses documented. Applied by U-Boot or by the kernel — state
which.</li>
</ul>
<h2 id="6-u-boot-the-config-the-environment-and-the-boot-flow">6. U-Boot: the config, the environment, and the boot flow</h2>
<ul>
<li>The U-Boot defconfig and any board patches, same discipline as the kernel.</li>
<li><strong>The boot environment as a file</strong>, not as whatever is currently in the SPI
flash of the one golden board. The <code>boot.cmd</code>/<code>boot.scr</code> or the extlinux
config, in version control.</li>
<li>The <strong>boot flow written out</strong>: ROM → SPL → U-Boot proper → kernel → init,
with where each stage lives (eMMC boot partition, offset, GPT) and what the
fallback path is if the primary kernel does not come up.</li>
<li>If secure boot is in play: which keys sign what, where the public keys are
fused or stored, and — critically — a documented recovery path for a board
whose signature check fails. A secure-boot BSP with no recovery story is a
brick generator.</li>
</ul>
<h2 id="7-the-signed-image-and-its-bill-of-materials">7. The signed image and its bill of materials</h2>
<ul>
<li>The exact image that was released, plus its <strong>signature</strong> and the public key
to verify it.</li>
<li><strong><code>buildhistory</code> output</strong> committed for that build: the package list with
versions, image size, dependency graph, and the diff against the previous
release. This is what makes “what changed between v1.2 and v1.3” answerable in
minutes.</li>
<li>An <strong>SBOM</strong> — Yocto emits SPDX 2.2 JSON via <code>create-spdx</code> — covering every
package in the image with its version and licence. Increasingly this is a
contractual and regulatory requirement (the EU Cyber Resilience Act among
them), and it is far cheaper to generate at build time than to reconstruct.</li>
<li>The <strong>licence manifest</strong> (<code>license.manifest</code>, <code>deploy/licenses/</code>) and the
sources for anything under a copyleft licence, or a working <code>bitbake -c archiver</code> configuration that regenerates them.</li>
</ul>
<h2 id="8-the-build-environment-itself">8. The build environment itself</h2>
<p>“Works on my machine” is not a hand-over. Pin the host too:</p>
<ul>
<li>A <strong>container</strong> (the <a href="https://github.com/crops/poky-container">CROPS</a> base or
a project Dockerfile) that fixes the host distro, the essential host packages,
Python, and locale. Yocto is sensitive to host GCC and host Python; a 2024
build host and a 2027 build host are not equivalent.</li>
<li>The exact <code>bitbake</code> invocation, target names, and any <code>BB_ENV_PASSTHROUGH</code>
additions.</li>
<li>Expected build resources: disk for <code>TMPDIR</code> and <code>SSTATE_DIR</code>, RAM, rough wall
time on a stated machine — so the client can size CI.</li>
<li>A <strong>populated sstate and downloads mirror</strong>, or instructions to build one.
Without it the first rebuild pulls hundreds of source archives from the
internet, and some of those URLs will be dead. An offline <code>DL_DIR</code> archive is
the single most valuable non-obvious artefact in the box.</li>
</ul>
<h2 id="9-documentation-that-is-task-shaped">9. Documentation that is task-shaped</h2>
<p>Not a wiki dump. Five documents, each answering a question someone will actually
ask:</p>
<ol>
<li><strong>Build it</strong> — from empty machine to flashable image, copy-paste commands.</li>
<li><strong>Flash it</strong> — factory path and field path, including the recovery/USB-download
procedure for a bricked board.</li>
<li><strong>Change the kernel config / add a package / bump a recipe</strong> — the routine
maintenance loop.</li>
<li><strong>Cut a release</strong> — versioning, signing, what gets tagged, what gets archived.</li>
<li><strong>Known issues and deferred work</strong> — the honest list. Every BSP has one; a
hand-over without it just means the client finds them the hard way.</li>
</ol>
<h2 id="what-this-costs">What this costs</h2>
<p>Producing this is real work — on the order of one to two engineer-weeks on top of
the BSP itself, most of it in items 1, 7, and 8. It is worth being explicit about
that in a statement of work rather than discovering it at the end. The payoff is
entirely deferred: it shows up the first time the client rebuilds without the
original team, and it is the difference between an afternoon and a re-engagement.</p>
<p>The trade-off we make deliberately: we do <em>not</em> hand over a build that tracks
upstream branches, even though that would look more “up to date” on delivery day.
A pinned build is frozen and will drift out of security currency until someone
deliberately bumps it — that is the cost. It is the right cost, because a build
that silently changes under the client is not a build they own.</p>
]]></content:encoded>
    </item>
  </channel>
</rss>
