ISO/IEC 17025 technical records for automated test benches

By Alex Hernandez · · 14 min read

View Markdown
Close-up of an instrument's carrying handle with a blank tag tied on and a round seal over one panel screw.
FIG. 1 — CALIBRATION TAG AND SEAL

ISO/IEC 17025:2017 asks that each laboratory activity leave a technical record (clause 7.5) holding the results, the date, who did the work and enough information to identify what affected the measurement, with amendments traceable to the original. An automated bench meets it by writing that record as the test runs: raw data, equipment identity, operator, method revision, uncertainty context.

This guide is for hardware teams whose bench data has to hold up inside an accredited lab. It maps each clause that touches records to fields an automated bench can capture, and marks what stays with the lab's quality system. Clause numbers were checked against the standard and documents from four accreditation bodies. The standard is sold by ISO, and your accreditation body's interpretation governs.

What does ISO/IEC 17025 require of technical records?

ISO/IEC 17025 "specifies the general requirements for the competence, impartiality and consistent operation of laboratories" and "is applicable to all organizations performing laboratory activities, regardless of the number of personnel" (ISO/IEC 17025:2017, clause 1, publisher preview). Accreditation bodies assess labs against it, which is why a lab is accredited to ISO/IEC 17025 rather than certified.

Technical records are clause 7.5, and it has two parts. The SADCAS assessment checklist summarizes them this way (translated from French):

  • 7.5.1: records include the date and the identity of the personnel responsible for each activity, and contain the results, the report and enough information to help identify the factors affecting the measurement results.
  • 7.5.2: amendments to technical records can be traced back to previous versions or to the original observations. Both the original and the amended data and files are kept, with the date of the change and the person responsible for it.

NATA's interpretation of the standard adds two points that matter for automation. Since several people may take part in one activity, the lab identifies the critical steps and records who performed them. And as far as practicable, records are indelible, kept so the original cannot be amended or lost.

The standard prescribes no record format. Its foreword says the 2017 edition made "some reduction in prescriptive requirements and their replacement by performance-based requirements," and more flexibility in documented information. An assessor checks whether the record answers the questions above, and whether the system that produced it, covered by clause 7.11, is under control.

Which ISO 17025 clauses apply to an automated test bench?

Clause numbers and topics come from the standard's table of contents, the SADCAS checklist, PJLA's clause 7.11 webinar and the CALA 2017 crosswalk.

ClauseTopicWhat the bench record should capture
6.2.5, 6.2.6Competence records; authorization for specific activitiesOperator on every run; approver of each procedure revision
6.3.3Environmental conditions monitored, controlled and recordedAmbient readings as steps, where the method or result depends on them
6.4.13Equipment records, including software and firmware versionFull instrument identity string per run, keyed to the equipment register
6.5Metrological traceabilityInstrument serial, so calibration records can be joined to the run
7.2.1.3, 7.2.2Current method version; validation of lab-developed methodsProcedure revision on every run; validation record per release
7.4.2Unambiguous identification of test itemsDUT serial and hardware revision on every run
7.5.1Results, date, personnel, factors affecting the resultRaw responses, parsed values, units, timestamps, operator
7.5.2Amendments traceable to the originalAppend-only storage; corrections as new entries that cite the old one
7.6, 7.8.3.1 c)Measurement uncertainty, reported where applicableRange and settings sent; uncertainty and coverage factor used for the verdict
7.8.6Statements of conformity and decision rulesDecision rule ID stored with each pass/fail verdict
7.11.2 to 7.11.6Data systems: validation, integrity, checked transfersSoftware revision on the run; system failures logged; parse checks

How do you keep raw data retrievable?

"Results" in 7.5.1 means more than the number that reaches the report. For a bench, the original observation is what the instrument returned: the response string or binary block, before parsing, scaling or rounding. Clause 7.11.3 requires the data system to "be safeguarded against tampering and loss" and to "be maintained in a manner that ensures the integrity of the data and information" (PJLA).

Store the raw response next to the parsed value. A parsed value alone cannot be rechecked; the session class in SCPI instrument automation with Python logs every raw response with its command and timing.

Keep full precision; round in the report. NATA's interpretation is that rounding happens only at the final stage of reporting unless the method says otherwise. A record that stored 3.29 V instead of +3.28714000E+00 has already lost the original.

Append, never overwrite. A correction is a new entry that cites the one it corrects, with the date and the person, as 7.5.2 describes. A re-run is a new run; the earlier one stays.

Keep aborted runs. Clause 7.11.3 e) asks that the data system "include recording system failures and the appropriate immediate and corrective actions." A run that stopped on a crashed executor or a lost connection is evidence of what happened, not noise to delete.

Check transfers in code. Clause 7.11.6 reads: "Calculations and data transfers shall be checked in an appropriate and systematic manner." With raw responses stored, a script can run the check every time:

check_transfers.py
import json
import math
from pathlib import Path
 
 
def check_transfers(run_dir: Path, rel_tol: float = 1e-12) -> list[str]:
    """Recompute each stored value from its raw response; return every mismatch."""
    problems: list[str] = []
    for path in sorted(run_dir.glob("step-*.json")):
        step = json.loads(path.read_text())
        raw = step.get("raw_response")
        if raw is None:
            problems.append(f"{path.name}: no raw response stored")
            continue
        try:
            parsed = float(raw.strip())
        except ValueError:
            problems.append(f"{path.name}: response {raw!r} does not parse as a number")
            continue
        if not math.isclose(parsed, step["value"], rel_tol=rel_tol):
            problems.append(f"{path.name}: stored {step['value']} but raw parses to {parsed}")
    return problems

It works on step records shaped like the one in hardware test traceability. Steps that scale or convert units need that calculation repeated in the check. Keep its output with the run, so the record shows the check happened.

Where the data lives matters too. Under 7.11.4, when a data system is "managed and maintained off-site or through an external provider," the lab must ensure the provider "complies with all applicable requirements of this document." For a cloud test-data platform, that is a supplier evaluation the lab must be able to show.

How should a test record identify equipment?

Clause 6.4.13 lists what the lab's equipment records hold, including the equipment's identity and its software and firmware version. Those records live in the lab's equipment register, not in each run. A run's job is to carry a key that joins to the register, and to capture what can change between runs.

Firmware gets updated. A meter goes out for calibration and a loaner of the same model replaces it. The *IDN? reply carries manufacturer, model, serial number and firmware, so query it at the start of every run and store the full string, not a parsed model name.

Calibration status is a join, not a field you type. Clause 6.5 covers metrological traceability, and evidence that a meter was in calibration on the day of a run comes from joining its serial number against calibration records. Without the serial in the run record, that join cannot be made later; instrument calibration and test records covers it. A pre-run check against the due dates in your register stops an out-of-calibration run before it creates a record someone has to disposition.

The bench's own software belongs in the record too. A driver change can change the commands an instrument receives, so record the driver version next to the procedure revision; declarative instrument drivers covers keeping drivers as versioned data.

How do you record who ran and approved a test?

Clause 7.5.1 asks for the identity of the personnel responsible for each activity. On an automated bench, the activities split three ways, and each needs a named person:

  1. Who authorized the procedure revision. Someone approved these limits, settling times and ranges before they touched a unit.
  2. Who ran it. The operator who set up the device under test and started the run.
  3. Who checked the results and authorized the report. Clause 7.8.2.1 o) requires a report to identify the people authorizing it.

Clauses 6.2.5 and 6.2.6 cover the other side: records of training, supervision, authorization and monitoring of competence, and authorization of personnel for specific activities. A run started by someone not authorized for that method is a gap against 6.2.6, and the bench record should make it visible.

Two habits break this. Shared bench logins make every run look like the same person, and runs started by a scheduler or service account name a machine. When a script or an agent starts a run, record the human who scheduled it and the human who approved the revision it executed. Test sequencers vs test agents covers change control when an agent writes the sequence.

Is a test script a method change under ISO 17025?

On an automated bench, the method is the written procedure plus the code that executes it: commands, settling times, ranges and limits. Change a limit or a delay and the method changed, whether or not the document did.

NATA reads 7.2.1.3 as requiring the current version of a standard method unless a legal or regulatory requirement calls for a superseded one. Clause 7.2.2 covers validation of non-standard and laboratory-developed methods; SADCAS expects validation records holding the procedure, requirements, performance characteristics, results and a statement of validity. Clause 7.8.2.1 f) requires the report to identify the method used.

Clause 7.11.2 reaches the code directly. Systems "used for the collection, processing, recording, reporting, storage or retrieval of data shall be validated for functionality" before introduction, and changes, "including laboratory software configuration or modifications to commercial off-the-shelf software," must be "authorized, documented and validated before implementation." The clause's second note says off-the-shelf software "in general use within its designed application range can be considered to be sufficiently validated." A bench script your team wrote is not off-the-shelf.

Whether an assessor treats a sequence as a method, software configuration or both varies. The controls come out the same:

  • Version every sequence, and have every run cite the exact revision it executed, down to a commit hash.
  • Release only approved revisions. Editing an approved revision creates a new one that needs its own approval.
  • Validate before release. Run the new revision on a reference unit, compare with the previous revision's results, and keep that comparison in the validation record.

What measurement uncertainty context belongs in the record?

Clause 7.6 covers evaluating measurement uncertainty, for calibration and testing labs alike. Under 7.8.3.1 c), test reports include it where applicable. When a customer asks for a statement of conformity, 7.1.3 requires the specification and decision rule to be clearly defined, and 7.8.6 requires the rule to be documented and the statement to name the results, the specification and the rule applied.

A bench limit check of 3.2 V to 3.4 V is a decision rule that ignores uncertainty. A reading of 3.399 V with an expanded uncertainty of 4 mV passes it, yet the true value may sit above 3.4 V. NATA's interpretation of 7.8.6.1 takes uncertainty into account: a statement may be made when the result is inside the limits by at least the uncertainty, when the uncertainty is within a maximum the specification sets, when the test specification defines the rule, or, outside regulatory compliance, when customer and lab have agreed one.

The uncertainty budget is the lab's work. The record's job is to carry the inputs to it and the rule that was applied:

  • The range and integration settings actually sent, so the instrument accuracy specification in the budget is the one that applied. Set them explicitly in the procedure; front-panel state left by the last user does not reach the record.
  • Environmental readings where the budget depends on them, which clause 6.3.3 asks labs to monitor, control and record.
  • The expanded uncertainty U and coverage factor k used for the verdict. The GUM notes that k "is typically in the range 2 to 3."
  • The decision rule ID, stored with the verdict.

A guarded-acceptance rule takes a few lines; storing its context makes the verdict reviewable:

decision_rule.py
from dataclasses import dataclass
 
 
@dataclass(frozen=True)
class DecisionRule:
    rule_id: str   # documented in the quality system and cited in the report
    guard: float   # multiple of U taken off each limit; 0.0 is simple acceptance
 
 
def judge(value: float, low: float, high: float, U: float, k: float, rule: DecisionRule) -> dict:
    """Apply a guarded-acceptance rule; return the verdict with the context to review it."""
    w = rule.guard * U
    if low + w <= value <= high - w:
        verdict = "conforms"
    elif low <= value <= high:
        verdict = "no statement: inside guard band"
    else:
        verdict = "outside limits"
    return {"value": value, "low": low, "high": high, "U": U, "k": k,
            "rule": rule.rule_id, "verdict": verdict}
 
 
rule = DecisionRule(rule_id="DR-02 guarded acceptance, w = U", guard=1.0)
print(judge(3.399, 3.2, 3.4, U=0.004, k=2.0, rule=rule)["verdict"])
# no statement: inside guard band

With the guard band equal to U, the acceptance zone is 3.204 V to 3.396 V, so 3.399 V gets no conformity statement, while simple acceptance passes it. Either can be a documented rule, though accreditation bodies may limit simple acceptance, as NATA's conditions above do; 7.8.6 asks that the record and report say which one applied.

How do Galois run records map to ISO 17025?

Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record.

The platform writes the run side of the record as each step executes. Clause numbers show which requirement each field feeds, not a claim of conformance:

  • Per step (7.5.1): measured value, limits and comparison operator, raw command sent and raw response received, instrument ID, step duration and timestamp.
  • Per run (7.4.2): operator, DUT serial and timestamps.
  • Change control (7.2, 7.11.2): every edit to a sequence is a new version with history and diffs. A draft cannot run until an engineer approves it, and one edited after approval must be approved again. Production versions can be locked.
  • Instrument identity (the join to 6.4.13 records): the galois-edge daemon queries each SCPI instrument with *IDN? at discovery and matches the reply against the profiles it has loaded; Galois ships 573 instrument profiles across 135 manufacturers (instrument library).
  • Who did what (7.5.1, 7.11.3): an audit log of actions with actor, timestamp and resource, described on the security page.
  • Reports (7.8): drafted from run data, so a reported number traces to a command and a response.

Sequences drafted by Évariste, the agent in the Galois platform, pass through the same approval gate, and the run record, not Évariste's summary, is the evidence. For the 7.11.4 question, the platform also runs as a dedicated single-tenant cloud or fully on-prem and air-gapped (deployment options). Sign-off packages compiled from the same record (traceability matrix, findings, attestation) are in build (product). The next section builds this guide's bench record with Évariste.

The rest stays with the lab's quality system: the equipment register and calibration records, uncertainty budgets and decision rules, competence and authorization records, validation of each released revision, and record control under clause 8.4. NATA, for example, sets a retention floor of four years, or the maximum recalibration interval for equipment records if longer, unless legislation or a contract says otherwise.

How to capture ISO/IEC 17025 test records in Galois with Évariste

Open Évariste from the app sidebar (Ctrl+Shift+E) beside the bench's project. Agent-driven test automation covers how it builds sequences.

Instruments. Ask "List connected instruments" to confirm the DMM is on the team's edge, then ask Évariste to send it *IDN?. Check the reply against your equipment register; that register string is what the first step will compare. For a meter outside the library, upload its programming manual; Évariste generates a profile and, after your review, deploys it to the edge and binds it.

Objective. State this guide's identity check, settings and decision rule:

Create a sequence for the 3.3 V rail. First query *IDN? on the DMM and fail unless the reply exactly matches the register identity string I paste. Set DC volts, the 10 V range and 10 NPLC explicitly. Measure the rail under DR-02, guarded acceptance with w = U, U = 4 mV, k = 2: limits 3.204 to 3.396 V inside the 3.2 to 3.4 V specification. Name the step with the rule ID, U, k and the specification.

Review and approve. The draft holds a string_value step comparing the whole *IDN? reply, firmware included, with the register string; action steps for function, range and integration; and a numeric_limit step from 3.204 to 3.396 V. Check the limits, range and NPLC against the uncertainty budget, then the step order and whether the budget needs an ambient step. Nothing runs until you approve it (reviewing a generated test plan). Évariste asks you to confirm dangerous commands it sends directly.

Run and interpret. Start the run; galois-edge executes it while Monitor shows the channel live. Each step stores the raw command and response, so the settings sent and a reply such as +3.28714000E+00 stay on the record. The guarded step fails 3.399 V and 3.45 V alike, so a fail is not yet a DR-02 verdict. Ask which steps failed or passed close to a limit; the specification in each cited step's name separates the "no statement" case of judge() from a reading outside the limits.

For clause 7.11.6, each step keeps the raw response beside the stored value. Ask Évariste to show both for the steps you sample and check them yourself, or keep check_transfers.py running on the records. Whether that meets the clause is the lab's call with its assessor.

Validate and report. Before releasing a new version, run it on the reference unit, ask Évariste to compare that run with the previous version's, and keep both in the validation record. After a production run, ask it to "Generate a test report from the last run". Production-lock the released version, and name it and its approver in the report's method section, so each report cites the version that ran. In the report editor, add the decision rule with U and k and the person authorizing the report, and word a guard-band result as no statement of conformity (7.8.2.1, 7.8.6).

You no longer write or maintain the run loop, the code that writes step records, the acceptance test in judge() or a report script. Reviewing, approving and wiring stay with you; the quality-system work listed above stays with the lab.

StepCode path (this guide)Galois with Évariste
Identify equipmentFull *IDN? string per runstring_value step against the register
DriverVersioned driver codeLibrary or generated profile, reviewed
Fix settingsRange and integration set explicitlyaction steps, recorded as sent
Raw dataRaw response beside the valueRaw command and response per step
Check transferscheck_transfers.py on each runRaw response beside each value; your sampled check or check_transfers.py
Decision rulejudge(), DR-02, guard = UGuarded limits named with rule and specification
PersonnelOperator per run; approver per revisionOperator per run; approver per version
Method revisionCommit hash on every runVersions, diffs, approval, production lock
Validate a releaseCompare with the previous revisionÉvariste compares the two runs
ReportYour report toolingGenerated report, edited before sign-off

For the wider chain from requirement to sign-off, read hardware test traceability. To see what the daemon captures on your own bench, start with the quickstart.

Frequently asked questions

What are technical records in ISO/IEC 17025?
Technical records are covered by clause 7.5 of ISO/IEC 17025:2017. Records for each laboratory activity hold the results, the report, the date, the identity of the people responsible, and enough information to help identify the factors that affected the measurement result. Under 7.5.2, any amendment must be traceable to the previous version or the original observation, with both versions kept along with the date and the person who made the change.
Does ISO/IEC 17025 require a specific record format or software?
No. The 2017 edition reduced some prescriptive requirements in favor of performance-based ones and allows more flexibility in documented information. What it does require is control: under clause 7.11, any system used to collect, process, record, report, store or retrieve data must be validated before use, protected from unauthorized access, safeguarded against tampering and loss, and changed only with authorization, documentation and validation.
Can I capture ISO/IEC 17025 test records without writing Python?
Yes. In Galois, describe the sequence to Évariste in plain English: a *IDN? check on the DMM against your equipment register string, explicit range and integration settings, and guarded limits in a step named with the decision rule, expanded uncertainty U, coverage factor k and the specification. The draft runs on the bench through galois-edge only after an engineer approves it. The run records each step's value, limits, pass or fail, instrument and raw command and response, with the operator, DUT serial and timestamps. Évariste then generates the report, where you add the decision rule with U and k and the person authorizing it; uncertainty budgets, method validation and accreditation stay with the lab.
How long must ISO/IEC 17025 records be kept?
Retention is set by the lab's record controls, contracts, regulation and the accreditation body. NATA in Australia, for example, requires records to be kept for at least four years, or for equipment records the maximum recalibration interval if longer, unless legislation or a contract sets a different period. Check your own accreditation body's requirements.
Can a lab be ISO 17025 certified?
Labs are accredited to ISO/IEC 17025, not certified. An accreditation body assesses the lab's competence, impartiality and consistent operation against the standard and its own requirements. Good records are evidence for that assessment; they do not grant accreditation.

Related

Bring Galois to your bench.

The daemon is Apache-2.0, free forever. Enterprise runs in your cloud or on-prem.