Engineering /HMI & Vision

ReferenceWorking6 min read

Where the milliseconds go between sensor and display

TL;DR

Glass-to-glass latency is exposure + sensor readout + CSI-2 + ISP + buffer and queue waits + composition + one or two display refreshes. Most of it is exposure and buffering, not code; shrink queues and sync to vsync.

View as Markdown

“The camera feels laggy” is a latency-budget problem, and it is almost always lost in places people do not instrument: exposure time, the number of buffers in each queue, and the wait for the next display refresh. Optimising the application code is usually the smallest available win.

This page walks the whole path from photons to photons and gives the order of magnitude of each stage, so you know which term to attack.

The pipeline, stage by stage

1. Exposure (integration time)

The sensor integrates light for the exposure time. The photons that form a frame are spread across that whole window, so the effective latency contribution is roughly half the exposure time for motion (the centroid of the exposure), but the frame is not readable until the full exposure ends, so for end-to-end timing you count the full exposure time.

  • Bright scene, short exposure (1–2 ms): negligible.
  • Indoor / auto-exposure (8–33 ms): this is often the single largest term, and it is invisible in any software profiler.
  • Low light with a 1/30 s exposure: 33 ms before anything else has happened.

Lever: cap the maximum auto-exposure time and accept more gain (noise) if latency matters more than SNR. This is an ISP/3A tuning decision, not a code change.

2. Sensor readout

The pixel array is read out row by row and streamed. Readout takes on the order of one frame period at the sensor’s current mode — a rolling shutter sensor at 60 fps takes ~16 ms to clock the whole frame out, and the bottom of the frame is genuinely ~16 ms “younger” than the top.

  • Higher frame-rate modes read out faster (and usually crop or bin).
  • Global shutter removes the top-to-bottom skew but not the readout transport time.
  • Lever: run the sensor faster than the display rate if the SoC can take it — reading a 60 fps stream for a 30 fps display halves this term and the exposure cap becomes easier to hit.

3. MIPI CSI-2 transport

Serialised over the D-PHY/C-PHY lanes into the SoC’s CSI receiver. At any sane lane rate this is sub-millisecond per frame for typical resolutions — it is almost never the problem. It shows up only if lanes are marginal and you are getting retimed/corrupted lines that force retry or drop.

4. Receiver DMA to memory

The CSI receiver writes frames into a ring of buffers in DRAM. Two things here:

  • The write itself is fast (bounded by memory bandwidth, sub-ms to low-ms).
  • The number of buffers in this ring is a latency knob. More buffers = smoother under jitter, but a frame can sit in the ring for up to (buffers − 1) × frame_period before anything consumes it. Four buffers at 30 fps is up to 100 ms of potential sit time. This is the classic hidden lag.

5. ISP

Debayer, black level, lens shading, denoise, sharpening, tone mapping, colour conversion, scaling. On a hardware ISP this is a few milliseconds and often pipelined with readout. In software (libcamera soft ISP, OpenCV) it can be tens of milliseconds and steals CPU from everything else.

  • 3A (auto-exposure/white-balance/focus) runs here and feeds the next frame — it adds a frame of loop latency to convergence, not to the display path, but a slow-converging AE means longer exposures for longer, which loops back to stage 1.
  • Lever: use the hardware ISP path (V4L2 media controller / libcamera with the vendor pipeline handler) rather than a software fallback.

6. Application / buffer handoff

The frame becomes a dmabuf handed to whatever draws it — a Qt QVideoSink, a GStreamer appsink/glimagesink, an LVGL canvas, a Weston/Wayland client, a direct DRM/KMS plane.

  • Every queue between elements is depth × frame_period of potential latency. GStreamer’s default queue sizes, a v4l2src with many buffers, a compositor that triple-buffers — each is a place frames wait.
  • Zero-copy or not. If the frame is memcpy’d (or worse, uploaded to a GL texture) at each stage, that is memory-bandwidth time and CPU stalls, a few ms each and more under load. dmabuf import all the way to the display plane avoids it.
  • Lever: shortest possible pipeline. A camera frame on its own DRM/KMS overlay plane, composited by the display controller, skips GPU composition entirely.

7. Composition

If the UI draws the video into a scene (overlays, HUD, controls), the GPU composites. One frame period at the render rate, plus the GPU’s own latency. Using a hardware overlay plane for the video and only compositing the UI chrome avoids putting the video through this stage.

8. Display refresh and panel response

  • The scanout waits for the next vsync to start sending the new frame: 0 to one refresh period of wait (average half). At 60 Hz that is up to 16.7 ms.
  • Double/triple buffering adds one or two more refresh periods depending on whether a frame misses its flip deadline.
  • The panel’s own response: LCD pixel response and its internal frame buffer / overdrive processing, commonly one frame for an LCD, sometimes more for a panel with a scaler or “image enhancement” it will not let you disable.
  • Lever: PAGE_FLIP synced to vsync with exactly the buffers you need (often double, not triple), and pick a panel/timing controller with a documented, low, fixed latency. Disable panel-side processing.

Adding it up

A representative indoor 30 fps preview, no tuning:

Stage Typical
Exposure 15–33 ms
Readout ~16–33 ms
CSI-2 + DMA 1–3 ms
Buffer ring wait (4 deep) 0–100 ms
ISP (hardware) 2–6 ms
Pipeline queues (GStreamer defaults) 30–100 ms
GPU composition ~16–33 ms
Vsync wait + buffering 16–50 ms
Panel 8–33 ms

That easily reaches 150–250 ms — the “laggy” complaint — and only ~5 ms of it is anything a code profiler would show you.

The same pipeline, tuned:

  • Sensor at 60 fps, AE capped at 8 ms.
  • Two CSI buffers, not four.
  • Hardware ISP path.
  • No GStreamer queues, or leaky=downstream max-size-buffers=1.
  • Video on a dedicated KMS overlay plane, UI composited separately.
  • Double-buffered page flip synced to vsync, panel processing off.

lands in the 40–70 ms range, and most of what remains is exposure, readout, and one refresh period — the irreducible physics.

How to measure it

Do not trust per-stage estimates; measure glass-to-glass:

  • Photodiode + LED + scope. Flash an LED in frame, detect it with a photodiode taped to the display, measure LED-on to display-brightens on the scope. This is the only number that counts.
  • On-screen millisecond counter filmed with a 120–240 fps camera alongside a physical stopwatch/timer in the same shot; count frames of offset.
  • Per-stage: V4L2 buffer timestamps (v4l2_buffer.timestamp), GST_DEBUG with GST_TRACERS=latency, DRM/KMS PAGE_FLIP event timestamps, and a GPIO toggle at each app-level handoff. Line them up against the glass-to-glass number to find the fat stage.

The trade-offs you are actually making

  • Fewer buffers = lower latency, less tolerance to scheduling jitter. Drop a deadline with two buffers and you get a visible stutter; with four you get lag but no stutter. Pick per product — a welding HUD wants low latency, a security-review monitor wants no dropped frames.
  • Shorter exposure = lower latency, more noise. You are trading SNR for responsiveness, frame by frame, in the AE tuning.
  • Overlay plane = low latency, less compositing flexibility. Overlay planes have format, scaling, and count limits set by the display controller; a complex UI that must blend with the video may not fit and has to go through the GPU.
  • Running the sensor faster = lower latency and easier AE, more MIPI and memory bandwidth, more power and heat. On a thermally constrained enclosure that can be the deciding constraint.

Talk to an engineer

Ask about this directly — the person who wrote it answers, not a sales desk.