---
title: Hardware Test Traceability from Requirement to Run
description: Hardware test traceability links a requirement to a test, sequence revision, run, measured value and sign-off. Why RTMs drift and what auditors need.
url: https://galoislabs.ai/blog/hardware-test-traceability
author: Alex Hernandez
author_url: https://galoislabs.ai/blog/authors/alex-hernandez
published: "2026-08-22"
topic: Test evidence
publisher: Galois Labs
---

# Hardware test traceability: from requirement to run to sign-off

![An open bound report seen from above, one blue leader line linking a requirement row on the left page to a run record on the right.](https://galoislabs.ai/blog/figures/evidence-1.light.webp)

*FIG. 1 — REQUIREMENT LINKED TO ITS RUN*

Hardware test traceability is the ability to start at any requirement and reach every run that verified it, and to walk back from any measured value to the requirement it proves. The chain is requirement, test case, sequence revision, run, measured value, evidence and sign-off. A spreadsheet often holds the first two links; the rest get rebuilt by hand.

The gap shows up at the worst moment: a design review, a customer audit, a field return, or the day a limit tightens and nobody can say which passed units were judged against the old one. This essay walks the chain link by link, explains why matrices maintained by hand drift, lists what a reviewer needs to reconstruct one result, and describes a record that stays correct without anyone maintaining it.

## What does hardware test traceability cover?

NASA's systems engineering handbook defines traceability as "a discernible association between two or more logical entities such as requirements, system elements, verifications, or tasks," and bidirectional traceability as "the ability to trace any given requirement/expectation to its parent requirement/expectation and to its allocated children requirements/expectations" ([NASA SE Handbook, 6.2](https://www.nasa.gov/reference/6-2-requirements-management/)). For test, the children of a requirement are its verifications, and a trace is only useful if it works in both directions:

- **Forward:** requirement REQ-PWR-012, revision C, is verified by these test cases, which ran on these units with these results.
- **Backward:** this 3.287 V reading came from this step of this sequence revision, on this unit, and it is evidence for REQ-PWR-012 revision C.

In software, a test usually runs against code identified by a commit, so a green build is close to the whole story. Hardware adds links in between. The test runs on a physical unit with a serial number and a hardware revision, through instruments with their own identities and calibration, under a procedure someone revised last week. Each of those is a place the chain can break.

| Link              | What it holds                                             | Identifier that must stay stable                  | How it usually breaks                         |
| ----------------- | --------------------------------------------------------- | ------------------------------------------------- | --------------------------------------------- |
| Requirement       | What the product must do, with a measurable criterion     | Requirement ID and revision                       | Text edited in place, no new revision         |
| Test case         | The method that verifies one or more requirements         | Test case ID, mapped to requirement IDs           | The mapping lives only in a spreadsheet       |
| Sequence revision | The exact steps, commands and limits that ran             | Revision number and commit hash                   | Limits edited in place; "latest" assumed      |
| Run               | One execution on one unit at one station                  | Run ID, DUT serial, operator, station, start time | A re-run overwrites the earlier result        |
| Measured value    | Number, unit, limits and comparison for one step          | Step within the run                               | Copied into a report, rounded, unit dropped   |
| Evidence          | Raw command and response, instrument identity, timestamps | Instrument model, serial and firmware             | A screenshot or summary replaces the original |
| Sign-off          | A named person accepting results for a requirement        | Approval record citing run IDs                    | An "approved" cell with no link to runs       |

Each link holds only if the one before it is pinned to a revision. "Verified by TC-PWR-012-01" says little if that test case has had four revisions and the record does not say which one ran.

## Why do requirements traceability matrices drift?

A requirements traceability matrix (RTM) is the usual tool: requirements in rows, test cases in columns, status in the cells. As a report, it is useful. As the system of record, it fails in predictable ways.

**It is a copy.** Someone types "Pass" after reading a result somewhere else: a log file, a report, a chat message. From then on, the matrix and the source age separately, and nothing flags when they disagree. Every transcription is also a chance to copy the wrong run or the wrong unit.

**It links names, not revisions.** A cell says TC-PWR-012-01 passed. It does not say which revision of the test, against which revision of the requirement, on which hardware revision. When any of those change, the cell still says Pass.

**It stores verdicts, not measurements.** "Pass" is a conclusion. A reviewer wants the number, the unit, the limit and the comparison. When a limit tightens from ±5% to ±3%, a stored measurement can be re-judged in seconds. A stored verdict has to be re-run or taken on trust.

**Changes do not propagate.** NASA's guidance is that "all changes should be subjected to a review and approval cycle to maintain traceability and to ensure that the impacts are fully assessed for all parts of the system" ([NASA SE Handbook, 6.2](https://www.nasa.gov/reference/6-2-requirements-management/)). A requirement change should mark every test and result downstream of it as suspect until someone reviews it. A spreadsheet does not know what is downstream.

**Many-to-many does not fit in a cell.** One sequence verifies several requirements. One requirement needs runs across temperature, several units and several builds. The matrix flattens that into one status per cell, usually the most recent one typed.

**Failures disappear.** A unit fails, gets reworked and passes; the matrix shows the pass. The failure, and the rework that explains it, is often what a reviewer most wants to see.

None of this is carelessness. It is what happens to any record maintained by hand next to the work instead of produced by the work. A stricter spreadsheet template does not fix it, because the problem is where the record comes from.

## What does an auditor need to reconstruct a test result?

To reconstruct a result is to answer every question about it from the record alone, without asking the engineer who ran it. Whether the person asking is a certification body, a customer's quality engineer, or a colleague chasing a field return a year later, the questions are much the same:

1. What was required, at which revision, with what acceptance criterion?
2. Which procedure ran, at which revision? Who approved that revision, and was it changed afterward?
3. Which unit (serial and hardware revision), at which station, with which instruments (model, serial, firmware)?
4. Who ran it, and when?
5. What did the instrument actually return, and how did that become the reported number?
6. What was the limit, which comparison applied, and what was the verdict?
7. Were there other runs of the same test on the same unit, including failures, and why was this one accepted?
8. Who signed off, on which runs, against which revision of the requirement?

FDA's data integrity guidance, written for drug manufacturing, states the principle behind that list plainly. Data should be "attributable, legible, contemporaneously recorded, original or a true copy, and accurate (ALCOA)," and "metadata is the contextual information required to understand data," with the example that "the number '23' is meaningless without metadata, such as an indication of the unit 'mg'" ([FDA, Data Integrity and Compliance With Drug CGMP, December 2018](https://www.fda.gov/media/119267/download)). Swap milligrams for volts and the point carries to a bench: a reading needs its unit, limit, instrument, operator and timestamp to mean anything. The same guidance defines an audit trail as a record "that allows for reconstruction of the course of events relating to the creation, modification, or deletion of an electronic record."

Regulated teams have the same expectation in rule form. For electronic records kept under FDA record requirements, [21 CFR 11.10(e)](https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11/subpart-B/section-11.10) calls for "secure, computer-generated, time-stamped audit trails" and adds: "Record changes shall not obscure previously recorded information." Since February 2, 2026, FDA's device quality rule, the QMSR, incorporates ISO 13485:2016 by reference ([FDA QMSR](https://www.fda.gov/medical-devices/postmarket-requirements-devices/quality-management-system-regulation-qmsr)); [medical device design verification](https://galoislabs.ai/blog/medical-device-design-verification) covers what that means for bench records. Test labs working to ISO/IEC 17025 have their own technical-records clauses, covered in [ISO/IEC 17025 records for automated benches](https://galoislabs.ai/blog/iso-17025-automated-test-records). Under the EU Cyber Resilience Act, technical documentation includes reports of the tests that verify conformity ([CRA test evidence](https://galoislabs.ai/blog/cyber-resilience-act-test-evidence)), and DO-160 qualification of airborne equipment rests on functional checks run before, during and after each exposure ([DO-160 bench evidence](https://galoislabs.ai/blog/do-160-environmental-test-evidence)).

Two items on the list deserve emphasis.

**Raw responses beat derived numbers.** If the record holds only "3.287 V", a reviewer has to trust whatever script parsed it. If it holds the command sent, the raw response string and the parsed value, the conversion can be checked. The session class in [SCPI instrument automation with Python](https://galoislabs.ai/blog/scpi-automation-python) logs every command and response for this reason.

**Instrument identity makes calibration answerable.** Whether a meter was in calibration on the day of a run is answered by joining its serial number against calibration records. If the run record does not carry the serial, that join cannot be made after the fact. [Calibration status in automated test records](https://galoislabs.ai/blog/instrument-calibration-test-records) goes further into that join.

> **Evidence supports certification; it does not grant it**
>
> A complete trace is what a test lab, certification body or regulator reviews. They decide whether it is sufficient. No record format makes that decision for them.

## How do you build test traceability that does not drift?

The pattern holds whatever tools you use: capture the record where the test runs, pin every link to a revision, keep every run, and generate the matrix instead of maintaining it.

### Capture the record where the test runs

The test executor, not a person, writes the run record at the moment each step finishes. That makes the record contemporaneous by construction and removes transcription entirely. Per step, record the command sent, the raw response, the parsed value and unit, the limits and comparison, the verdict, the instrument and a timestamp. Per run, record the unit serial, operator, station, start and end times, and the procedure revision. An illustrative step record:

```json title="run-7f3a/step-04.json"
{
  "run_id": "7f3a",
  "sequence": { "name": "3V3 rail check", "revision": 4, "commit": "9c41e2b" },
  "dut_serial": "SN-0142",
  "operator": "j.doe",
  "step": "Measure 3.3 V rail",
  "instrument": { "id": "dmm-2", "model": "34461A", "serial": "MY54505555" },
  "command": "MEAS:VOLT:DC?",
  "raw_response": "+3.28714000E+00",
  "value": 3.28714,
  "unit": "V",
  "limits": { "low": 3.2, "high": 3.4, "comparison": "GELE" },
  "verdict": "pass",
  "timestamp": "2026-10-05T14:02:11.482Z"
}
```

Store the instrument's full `*IDN?` reply once per session as well, since it carries the firmware revision.

### Pin every link to a revision

Requirements, test cases and sequences each carry an ID and a revision, and every link points at a revision, not a name. An approved sequence revision should be immutable: editing it creates a new revision that needs its own approval. Each run records the exact revision it executed, down to a commit hash, so "which limits applied?" has one answer.

Approval belongs on the procedure before it touches a unit, not only on the results afterward. A reviewer who approves a sequence is approving its limits, and a wrong limit undermines every result judged against it. [Reviewing an agent-drafted test plan](https://galoislabs.ai/blog/review-ai-generated-test-plan) is a checklist for that review, and it applies just as well to sequences people wrote.

### Keep every run

Append, never overwrite. A unit that failed, was reworked and passed has three records, not one. Which run counts toward a requirement is a decision made by a named person, with a reason, and recorded next to the runs it covers. That decision is the sign-off link.

### Generate the matrix from the records

Once runs cite the requirement revisions they verify, the RTM becomes a query rather than a document. It is computed whenever someone opens it, so it cannot fall behind, and it can flag what a spreadsheet cannot: results recorded against an older requirement revision.

```python title="rtm.py"
from dataclasses import dataclass


@dataclass(frozen=True)
class Requirement:
    id: str
    revision: str


@dataclass(frozen=True)
class Run:
    run_id: str
    requirement_id: str
    requirement_revision: str  # the revision this run was verifying
    dut_serial: str
    passed: bool
    accepted_by: str | None    # sign-off: who accepted this run, if anyone


def rtm(requirements: list[Requirement], runs: list[Run]) -> dict[str, str]:
    status: dict[str, str] = {}
    for req in requirements:
        mine = [r for r in runs if r.requirement_id == req.id]
        current = [r for r in mine if r.requirement_revision == req.revision]
        if not mine:
            status[req.id] = "not tested"
        elif not current:
            status[req.id] = "suspect: results are against an older revision"
        elif any(r.passed and r.accepted_by for r in current):
            status[req.id] = "verified"
        elif any(r.passed for r in current):
            status[req.id] = "passed, awaiting sign-off"
        else:
            status[req.id] = "failing"
    return status
```

A production version also joins on test case and sequence revision, but the shape holds: status is computed from records, never typed. The same records feed the report; a [DVT test report template](https://galoislabs.ai/blog/dvt-test-report-template) shows which fields a reviewer expects to find there.

## How does Galois keep the chain intact?

Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record.

Its record covers the middle of the chain, from sequence revision to evidence, and the platform captures it as the test runs. In the product today:

- **Versioned sequences.** A sequence is YAML with explicit limits and comparison operators (GELE, GELT, GTLE, GTLT, EQ, NE). Each has version history and diffs, a version number is unique within a project and sequence name, and production versions can be locked.
- **Approval gates.** A draft sequence cannot run until an engineer approves it. Edit an approved sequence and it must be approved again, because approval applies to the exact version reviewed.
- **A per-step run record.** Every step records the measured value, the limits, the raw command sent and raw response received, the instrument ID, the step duration and a timestamp, tied to the operator and the DUT serial. The platform writes it as the step runs; nobody types it afterward.
- **Reports from the record.** Évariste, the agent in the Galois platform, can draft a report from a run; a report can also be generated from a data-bound LaTeX template or written in the built-in editor, so a number in a report traces to a command and a response.
- **An audit log** of actions with actor, timestamp and resource, described on the [security page](https://galoislabs.ai/security).

When Évariste writes the sequence, the same gates apply: the draft waits for an engineer's approval, and the run record, not Évariste's summary, is the evidence. [Agent-drafted sequences on hardware benches](https://galoislabs.ai/blog/ai-test-automation-hardware) covers that loop in detail, and the [memory layer](https://galoislabs.ai/workplace-intelligence) describes how Évariste answers questions over earlier runs and notes with citations.

To follow this pattern in Galois, open Évariste from the app sidebar (Ctrl+Shift+E) beside the project and state the objective with the requirement's limits: measure the 3.3 V rail on the 34461A DMM against 3.2 to 3.4 V, with REQ-PWR-012 revision C in the step name. Check the draft's limits and comparison, approve it, then run it on SN-0142 through galois-edge. Production-lock the approved version and name it in the report, so the chain from requirement to sequence version is explicit. Any later edit becomes a new version with a diff. After the runs, ask which steps failed or passed close to a limit, or which runs on SN-0142 measured that rail, failures included; Évariste answers with citations to the runs. Record which run counts, and why, in a project note that Évariste can cite later. You no longer write or maintain the code that logs each step. Wiring the meter to the rail, the limits, the review and the acceptance decision stay with the engineer.

Packaged sign-off deliverables compiled from the same record, a requirements traceability matrix, findings and attestation, are in build ([product](https://galoislabs.ai/product)). Teams that need records to stay on their own infrastructure can run the platform as a dedicated single-tenant cloud or fully on-prem and air-gapped ([deployment options](https://galoislabs.ai/deployment)).

## When is a spreadsheet RTM enough?

A hand-maintained matrix holds up in some situations, and replacing it there costs more than it saves:

- Tens of requirements rather than hundreds, one bench, and one person who both runs the tests and updates the sheet.
- Tests that run once per build and are rarely revisited.
- No external reviewer, regulated buyer or customer audit in view.
- An organization that already runs a requirements management tool holding the trace. There, the job is feeding run IDs and revisions into it, not replacing it.

Even then, two habits keep a spreadsheet honest. Put run IDs and revisions in the cells instead of the word Pass, and keep the raw logs at the locations those IDs point to. A reviewer can then follow any cell back to a measurement, which is most of what traceability asks for.

If your current test executive already logs results to a database, [TestStand alternatives](https://galoislabs.ai/blog/teststand-alternatives) compares how each option records evidence. If your bench telemetry lives in a data platform, [hardware test data platforms](https://galoislabs.ai/blog/hardware-test-data-platforms) covers joining it to step records with shared keys. If you are planning a build-phase test campaign, [EVT, DVT and PVT testing](https://galoislabs.ai/blog/evt-dvt-pvt-testing) covers what each phase has to prove. Whichever applies, the question to ask of any setup is the same: pick one number from last month's runs and see whether the record alone can say where it came from.

## Frequently asked questions

### What is hardware test traceability?

It is the ability to follow a requirement forward to every test case, sequence revision, run and measured value that verified it, and to follow any measured value back to the requirement it proves. For hardware, the chain also carries the unit's serial number, the instruments used, the operator and the procedure revision, because the same test can give different results on different units and stations.

### What is a requirements traceability matrix (RTM)?

An RTM maps each requirement to the test cases that verify it and shows their status. It works well as a report generated from run records. Maintained by hand as the system of record, it drifts: cells name tests rather than revisions, store verdicts rather than measurements, and do not change when a requirement does.

### What should a hardware test record include?

Per step: the command sent, the raw instrument response, the measured value and unit, the limits and comparison, the verdict, the instrument's identity and serial number, and a timestamp. Per run: the procedure revision, the unit's serial number and hardware revision, the operator, the station, and start and end times. Keep every run, including failures.

### Is a pass/fail result enough evidence for an audit?

Usually not. A reviewer needs the measured value, the limit it was judged against, which revision of the procedure ran, on which unit and instruments, who approved the procedure, and who accepted the result. A stored verdict cannot be re-checked when a limit changes; a stored measurement can.

### How do you keep an RTM up to date?

Stop maintaining it by hand. Have the test executor write a run record at execution time that cites the sequence revision and the requirements it verifies, then generate the matrix as a query over those records. Flag results recorded against an older requirement revision as suspect until someone reviews them.
