---
title: "Where the milliseconds go between sensor and display"
tldr: "Glass-to-glass latency is exposure + sensor readout + CSI-2 + ISP + buffer and queue waits + composition + one or two display refreshes. Most of it is exposure and buffering, not code; shrink queues and sync to vsync."
type: "reference"
hub: "HMI & Vision"
published: "2026-09-03"
updated: "2026-09-03"
canonical: "https://exubits.com/engineering/sensor-to-display-latency"
author: "Exubits Engineering"
---

# Where the milliseconds go between sensor and display

> Glass-to-glass latency is exposure + sensor readout + CSI-2 + ISP + buffer and queue waits + composition + one or two display refreshes. Most of it is exposure and buffering, not code; shrink queues and sync to vsync.

"The camera feels laggy" is a latency-budget problem, and it is almost always
lost in places people do not instrument: exposure time, the number of buffers in
each queue, and the wait for the next display refresh. Optimising the application
code is usually the smallest available win.

This page walks the whole path from photons to photons and gives the order of
magnitude of each stage, so you know which term to attack.

## The pipeline, stage by stage

### 1. Exposure (integration time)

The sensor integrates light for the exposure time. The *photons that form a
frame* are spread across that whole window, so the effective latency contribution
is roughly **half the exposure time** for motion (the centroid of the exposure),
but the frame is not readable until the full exposure ends, so for end-to-end
timing you count the **full exposure time**.

- Bright scene, short exposure (1–2 ms): negligible.
- Indoor / auto-exposure (8–33 ms): this is often the single largest term, and it
  is invisible in any software profiler.
- Low light with a 1/30 s exposure: 33 ms before anything else has happened.

Lever: cap the maximum auto-exposure time and accept more gain (noise) if latency
matters more than SNR. This is an ISP/3A tuning decision, not a code change.

### 2. Sensor readout

The pixel array is read out row by row and streamed. Readout takes on the order
of **one frame period** at the sensor's current mode — a rolling shutter sensor
at 60 fps takes ~16 ms to clock the whole frame out, and the bottom of the frame
is genuinely ~16 ms "younger" than the top.

- Higher frame-rate modes read out faster (and usually crop or bin).
- Global shutter removes the top-to-bottom skew but not the readout transport
  time.
- Lever: run the sensor faster than the display rate if the SoC can take it —
  reading a 60 fps stream for a 30 fps display halves this term and the exposure
  cap becomes easier to hit.

### 3. MIPI CSI-2 transport

Serialised over the D-PHY/C-PHY lanes into the SoC's CSI receiver. At any sane
lane rate this is **sub-millisecond per frame** for typical resolutions — it is
almost never the problem. It shows up only if lanes are marginal and you are
getting retimed/corrupted lines that force retry or drop.

### 4. Receiver DMA to memory

The CSI receiver writes frames into a ring of buffers in DRAM. Two things here:

- The write itself is fast (bounded by memory bandwidth, sub-ms to low-ms).
- **The number of buffers in this ring is a latency knob.** More buffers =
  smoother under jitter, but a frame can sit in the ring for up to
  `(buffers − 1) × frame_period` before anything consumes it. Four buffers at
  30 fps is up to 100 ms of potential sit time. This is the classic hidden lag.

### 5. ISP

Debayer, black level, lens shading, denoise, sharpening, tone mapping, colour
conversion, scaling. On a hardware ISP this is **a few milliseconds** and often
pipelined with readout. In software (libcamera soft ISP, OpenCV) it can be tens
of milliseconds and steals CPU from everything else.

- 3A (auto-exposure/white-balance/focus) runs here and feeds *the next* frame —
  it adds a frame of loop latency to convergence, not to the display path, but a
  slow-converging AE means longer exposures for longer, which loops back to
  stage 1.
- Lever: use the hardware ISP path (V4L2 media controller / libcamera with the
  vendor pipeline handler) rather than a software fallback.

### 6. Application / buffer handoff

The frame becomes a `dmabuf` handed to whatever draws it — a Qt `QVideoSink`, a
GStreamer `appsink`/`glimagesink`, an LVGL canvas, a Weston/Wayland client, a
direct DRM/KMS plane.

- **Every queue between elements is `depth × frame_period` of potential latency.**
  GStreamer's default queue sizes, a `v4l2src` with many buffers, a compositor
  that triple-buffers — each is a place frames wait.
- **Zero-copy or not.** If the frame is memcpy'd (or worse, uploaded to a GL
  texture) at each stage, that is memory-bandwidth time and CPU stalls, a few ms
  each and more under load. `dmabuf` import all the way to the display plane
  avoids it.
- Lever: shortest possible pipeline. A camera frame on its own DRM/KMS overlay
  plane, composited by the display controller, skips GPU composition entirely.

### 7. Composition

If the UI draws the video into a scene (overlays, HUD, controls), the GPU
composites. One frame period at the render rate, plus the GPU's own latency.
Using a hardware overlay plane for the video and only compositing the UI chrome
avoids putting the video through this stage.

### 8. Display refresh and panel response

- The scanout waits for the next **vsync** to start sending the new frame:
  0 to one refresh period of wait (average half). At 60 Hz that is up to 16.7 ms.
- Double/triple buffering adds **one or two more refresh periods** depending on
  whether a frame misses its flip deadline.
- The panel's own response: LCD pixel response and its internal frame buffer /
  overdrive processing, commonly **one frame** for an LCD, sometimes more for a
  panel with a scaler or "image enhancement" it will not let you disable.
- Lever: `PAGE_FLIP` synced to vsync with exactly the buffers you need (often
  double, not triple), and pick a panel/timing controller with a documented,
  low, fixed latency. Disable panel-side processing.

## Adding it up

A representative indoor 30 fps preview, no tuning:

| Stage | Typical |
|---|---|
| Exposure | 15–33 ms |
| Readout | ~16–33 ms |
| CSI-2 + DMA | 1–3 ms |
| Buffer ring wait (4 deep) | 0–100 ms |
| ISP (hardware) | 2–6 ms |
| Pipeline queues (GStreamer defaults) | 30–100 ms |
| GPU composition | ~16–33 ms |
| Vsync wait + buffering | 16–50 ms |
| Panel | 8–33 ms |

That easily reaches **150–250 ms** — the "laggy" complaint — and only ~5 ms of it
is anything a code profiler would show you.

The same pipeline, tuned:

- Sensor at 60 fps, AE capped at 8 ms.
- Two CSI buffers, not four.
- Hardware ISP path.
- No GStreamer queues, or `leaky=downstream max-size-buffers=1`.
- Video on a dedicated KMS overlay plane, UI composited separately.
- Double-buffered page flip synced to vsync, panel processing off.

lands in the **40–70 ms** range, and most of what remains is exposure, readout,
and one refresh period — the irreducible physics.

## How to measure it

Do not trust per-stage estimates; measure glass-to-glass:

- **Photodiode + LED + scope.** Flash an LED in frame, detect it with a
  photodiode taped to the display, measure LED-on to display-brightens on the
  scope. This is the only number that counts.
- **On-screen millisecond counter filmed with a 120–240 fps camera** alongside a
  physical stopwatch/timer in the same shot; count frames of offset.
- **Per-stage**: V4L2 buffer timestamps (`v4l2_buffer.timestamp`),
  `GST_DEBUG` with `GST_TRACERS=latency`, DRM/KMS `PAGE_FLIP` event timestamps,
  and a GPIO toggle at each app-level handoff. Line them up against the
  glass-to-glass number to find the fat stage.

## The trade-offs you are actually making

- **Fewer buffers = lower latency, less tolerance to scheduling jitter.** Drop a
  deadline with two buffers and you get a visible stutter; with four you get lag
  but no stutter. Pick per product — a welding HUD wants low latency, a
  security-review monitor wants no dropped frames.
- **Shorter exposure = lower latency, more noise.** You are trading SNR for
  responsiveness, frame by frame, in the AE tuning.
- **Overlay plane = low latency, less compositing flexibility.** Overlay planes
  have format, scaling, and count limits set by the display controller; a complex
  UI that must blend with the video may not fit and has to go through the GPU.
- **Running the sensor faster = lower latency and easier AE, more MIPI and memory
  bandwidth, more power and heat.** On a thermally constrained enclosure that can
  be the deciding constraint.
