---
title: OpenHTF vs OpenTAP vs pytest vs TestStand
description: "OpenHTF vs OpenTAP vs pytest vs TestStand: test model, plugins, operator UI, result storage, license, language and OS compared, with code for each one."
url: https://galoislabs.ai/blog/test-executive-comparison
author: Alex Hernandez
author_url: https://galoislabs.ai/blog/authors/alex-hernandez
published: "2026-10-02"
topic: Comparisons
publisher: Galois Labs
---

# OpenHTF vs OpenTAP vs pytest vs TestStand: choosing a test executive

![A rugged embedded controller drawn as an exploded assembly, three plug-in modules lifted above their slots.](https://galoislabs.ai/blog/figures/ni-3.light.webp)

*FIG. 1 — CONTROLLER AND MODULES, EXPLODED*

OpenHTF, OpenTAP, pytest and TestStand all run hardware tests with pass or fail verdicts, but differ in what a test is and how much you build yourself. TestStand is a commercial Windows executive with operator interfaces, reports and deployment included. OpenTAP is an open-source .NET sequencer extended by plugins. OpenHTF is a Python test framework. pytest is a general-purpose runner.

Choose by the language your team writes, the operating system your stations run, and how much of the operator interface, result storage and deployment you are prepared to own. Each section below describes one tool from its own documentation, with a code sketch.

> **Disclosure**
>
> We build Galois, which is not one of the four tools compared here. Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record. It appears only at the end. Claims about the four link to their own documentation.

## What does a test executive do?

A test executive sits between test code and the production floor. Whatever the brand, it does six jobs:

1. **Model the test:** steps in order, each with a measurement and the limits that judge it.
2. **Reach the hardware:** open instruments and the DUT, and close them when the run ends, even after a failure.
3. **Control flow on failure:** stop, continue, retry or skip.
4. **Face the operator:** serial number, start, abort, pass or fail.
5. **Store the result:** one record per unit, queryable a year later.
6. **Deploy and version:** the same approved test on every station.

TestStand ships all six in one product. The open-source tools cover the first three in different shapes and leave different amounts of the last three to you. That split, more than syntax, decides most choices.

## OpenHTF vs OpenTAP vs pytest vs TestStand at a glance

| Dimension              | OpenHTF                                               | OpenTAP                                                              | pytest                           | TestStand                                            |
| ---------------------- | ----------------------------------------------------- | -------------------------------------------------------------------- | -------------------------------- | ---------------------------------------------------- |
| License                | Apache-2.0                                            | MPL-2.0                                                              | MIT                              | Commercial, per developer and per station            |
| Language               | Python 3.9+ (release 1.6.1)                           | C# on .NET; Python plugin                                            | Python 3.10+ or PyPy3            | Code modules in LabVIEW, C/C++, .NET or Python       |
| Operating system       | Pure-Python package (`py3-none-any` wheel)            | Windows 10+, Ubuntu 24.04+, macOS 12+ on Apple silicon; Docker guide | Linux, macOS, Windows            | Windows 10 and 11                                    |
| Test definition        | Python program of phases                              | XML test plan (`.TapPlan`) of test steps                             | Python modules of test functions | Sequence file (`.seq`): binary, XML or INI           |
| Limits                 | Measurement validators such as `in_range`             | Verdict set in step code                                             | `assert` in test code            | Numeric Limit, Pass/Fail and String Value Test steps |
| Hardware access        | Plugs                                                 | Instrument and DUT resources                                         | Fixtures                         | Code modules called through adapters                 |
| Operator UI and editor | Station server with a web frontend                    | Keysight Developer System, or the open-source TUI                    | Not bundled                      | Sequence editor and operator interfaces              |
| Result storage         | Output callbacks: JSON, console summary, MfgInspector | Result listeners: Text Log bundled; CSV, SQLite, PostgreSQL          | JUnit XML; plugins               | Reports in HTML, XML, ATML, text; database logging   |
| Deployment             | Python packaging                                      | Package manager with versioned packages                              | Python packaging                 | Deployment packages, licensed per station            |

Each cell is sourced below. The concepts line up more closely than the syntax suggests:

| Concept                   | OpenHTF                               | OpenTAP                                | pytest                     | TestStand                                   |
| ------------------------- | ------------------------------------- | -------------------------------------- | -------------------------- | ------------------------------------------- |
| One step                  | Phase                                 | Test step                              | Test function              | Step                                        |
| Measurement and limit     | `Measurement` with validators         | Published result plus `UpgradeVerdict` | Value plus `assert`        | `Step.Result.Numeric` against `Step.Limits` |
| Hardware handle           | Plug                                  | Resource                               | Fixture                    | Code module                                 |
| Before and after each run | `test_start` trigger, plug `tearDown` | `PrePlanRun`, `PostPlanRun`            | Fixture setup and teardown | Process model                               |
| Run record                | Test record                           | Result tables                          | Test report                | Report and database                         |

## OpenHTF: phases, measurements and plugs

OpenHTF calls itself "the open-source hardware testing framework," a Python library meant to remove boilerplate from hardware test setup and execution ([google/openhtf](https://github.com/google/openhtf)). It is Apache-2.0, and its README states: "This is not an official Google product."

A test is a Python program and phases are its steps. Measurements are declared on the phase with a spec, and a value outside it fails the phase. Plugs wrap the DUT or an instrument, and OpenHTF calls each plug's `tearDown()` when the test ends. Output callbacks write the record.

```python title="rail_check.py"
import openhtf as htf
import pyvisa
from openhtf.output.callbacks import json_factory
from openhtf.plugs import user_input
from openhtf.util import units

RM = pyvisa.ResourceManager()


class Psu(htf.BasePlug):
    def __init__(self):
        self._inst = RM.open_resource("TCPIP0::192.168.1.40::inst0::INSTR")

    def set_output(self, volts: float) -> None:
        self._inst.write(f"VOLT {volts}")
        self._inst.write("OUTP ON")
        self._inst.query("*OPC?")

    def tearDown(self):  # called when the test ends, pass or fail
        self._inst.write("OUTP OFF")
        self._inst.close()


class Dmm(htf.BasePlug):
    def __init__(self):
        self._inst = RM.open_resource("TCPIP0::192.168.1.50::inst0::INSTR")

    def measure_vdc(self) -> float:
        return float(self._inst.query("MEAS:VOLT:DC?"))

    def tearDown(self):
        self._inst.close()


@htf.plug(psu=Psu)
def power_on(test, psu):
    psu.set_output(12.0)


@htf.PhaseOptions(timeout_s=10)
@htf.plug(dmm=Dmm)
@htf.measures(htf.Measurement("vout_5v").with_units(units.VOLT).in_range(4.9, 5.1))
def measure_rail(test, dmm):
    test.measurements.vout_5v = dmm.measure_vdc()


if __name__ == "__main__":
    test = htf.Test(power_on, measure_rail, test_name="rail-check")
    test.add_output_callbacks(
        json_factory.OutputToJSON("./{dut_id}.{start_time_millis}.json", indent=2)
    )
    test.execute(test_start=user_input.prompt_for_test_start())  # operator enters the DUT ID
```

**Flow control** lives in the phase. A phase can return a `PhaseResult` such as `REPEAT`, `SKIP`, `FAIL_AND_CONTINUE` or `STOP`, and `PhaseOptions` sets timeouts and repeat limits. A `PhaseGroup` guarantees that its teardown phases run even when a main phase hits a terminal error ([source](https://github.com/google/openhtf/blob/master/openhtf/core/phase_group.py)).

**Operator UI.** The station server "serves an Angular frontend and information about a running OpenHTF test," and a [dashboard server](https://github.com/google/openhtf/blob/master/openhtf/output/servers/dashboard_server.py) lists the stations it finds on the network by multicast ([station_server.py](https://github.com/google/openhtf/blob/master/openhtf/output/servers/station_server.py)). OpenHTF [requires a DUT ID](https://github.com/google/openhtf/blob/master/examples/hello_world.py) on every run; `prompt_for_test_start()` asks the operator for it.

**Results.** The bundled output callbacks are a JSON writer, a console summary and an uploader for MfgInspector ([callbacks](https://github.com/google/openhtf/tree/master/openhtf/output/callbacks)). A database or MES upload is a callback you write. [Hardware test data platforms](https://galoislabs.ai/blog/hardware-test-data-platforms) compares Nominal, Sift and Synnax as destinations.

**Instruments.** The bundled plugs cover operator prompts, serial capture, ADB and fastboot devices, and a Cambrionix USB unit. Bench instruments come from plugs you write, usually over PyVISA as above, or from third-party plugs such as a VISA plug listed in the [plugs README](https://github.com/google/openhtf/blob/master/openhtf/plugs/README.md).

**Maintenance.** Commits continue into October 2026, and the latest release on [PyPI](https://pypi.org/project/openhtf/) is 1.6.1, from March 2025. Pin a version or commit per station.

## OpenTAP: test plans, steps and result listeners

OpenTAP is "an Open Source project for fast and easy development and execution of automated tests," built on "an extendable architecture that leverages .NET" ([opentap/opentap](https://github.com/opentap/opentap)). It is MPL-2.0, with a Keysight Technologies copyright in its license file.

"A test plan is a sequence of test steps and their associated data," stored as XML with the `.TapPlan` extension, and steps nest: a parent step decides whether its children run in sequence or in parallel ([user guide](https://doc.opentap.io/User%20Guide/Introduction/Readme.html)). A step is a C# class whose `Run()` method sets its verdict. Instruments are resources configured per bench, and the `ScpiInstrument` base class supplies a VISA address, `ScpiCommand` and `ScpiQuery` ([developer guide, resources](https://doc.opentap.io/Developer%20Guide/Resources/Readme.html)).

```csharp title="RailCheck.cs"
using System.Collections.Generic;
using OpenTap;

[Display("Bench DMM", Group: "Bench")]
public class BenchDmm : ScpiInstrument   // VisaAddress is set in the bench settings
{
    public BenchDmm() { Name = "DMM"; }

    public double MeasureVdc() => ScpiQuery<double>("MEAS:VOLT:DC?");
}

[Display("Measure 5 V Rail", Group: "Power")]
public class MeasureRail : TestStep
{
    public BenchDmm Dmm { get; set; }   // chosen from the configured instruments

    [Unit("V")] public double LowLimit { get; set; } = 4.9;
    [Unit("V")] public double HighLimit { get; set; } = 5.1;

    public override void Run()
    {
        double vout = Dmm.MeasureVdc();
        Results.Publish("Rail", new List<string> { "Vout", "Low", "High" }, vout, LowLimit, HighLimit);
        UpgradeVerdict(vout >= LowLimit && vout <= HighLimit ? Verdict.Pass : Verdict.Fail);
    }
}
```

```sh title="run the plan headless"
# run headless; --results enables only these configured result listeners
tap run RailCheck.TapPlan --results SQLite,CSV
```

**Flow control.** Every step ends with a verdict (NotSet, Pass, Inconclusive, Fail, Aborted or Error), a plan takes the most severe one, and break conditions, set for the engine and overridable per step, decide whether a failure stops the plan. The [basic steps](https://github.com/opentap/opentap/tree/main/BasicSteps) add delays, if, repeat, parallel, sweep loops, dialogs and a raw SCPI step.

**Results.** "Result listeners are notified whenever a test step generates log output, or publishes results." A Text Log listener ships with OpenTAP; CSV, SQLite and PostgreSQL listeners store results, and the database listeners keep a copy of the test plan so an old version can be rerun ([user guide](https://doc.opentap.io/User%20Guide/Introduction/Readme.html)).

**Plugins.** Hardware support is a plugin by design: "Out of the box, OpenTAP does not provide any resources for hardware control" ([user guide](https://doc.opentap.io/User%20Guide/Introduction/Readme.html)). Plugins ship as versioned packages, and the [OpenTAP Python plugin](https://github.com/opentap/OpenTap.Python) (Apache-2.0) exposes "the full OpenTAP API" to Python.

**Editing and platforms.** Plans are edited in Keysight's Developer System, "available with a commercial or community license," or the open-source TUI ([editors](https://doc.opentap.io/User%20Guide/Editors/Readme.html)). Keysight's [PathWave Test Automation](https://www.keysight.com/us/en/products/software/pathwave-test-software/pathwave-test-automation-software.html) "leverages OpenTAP open source test automation sequencing engine"; our [Keysight comparison](https://galoislabs.ai/compare/keysight) and [PathWave alternatives](https://galoislabs.ai/blog/pathwave-test-automation-alternatives) cover it. The [downloads page](https://opentap.io/downloads) offers Windows 10 or later, Ubuntu 24.04 or later and macOS 12 or later on Apple silicon, and links an installation guide for Docker.

## pytest for hardware testing: fixtures, parametrize and JUnit XML

pytest is a general Python test framework, MIT-licensed, with "modular fixtures for managing small or parametrized long-lived test resources" and more than 1,300 external plugins ([pytest-dev/pytest](https://github.com/pytest-dev/pytest)). It can run in the same CI as firmware tests.

A fixture is the hardware handle: code before `yield` opens the instrument, code after it closes it, and the scope sets how often ([fixtures](https://docs.pytest.org/en/stable/how-to/fixtures.html)). The DUT serial is your own convention; here, a command-line option recorded on the test suite.

```python title="conftest.py"
import pytest
import pyvisa

RM = pyvisa.ResourceManager()


def pytest_addoption(parser):
    parser.addoption("--dut-serial", default="UNKNOWN")


@pytest.fixture(scope="session", autouse=True)
def dut_serial(request, record_testsuite_property):
    serial = request.config.getoption("--dut-serial")
    record_testsuite_property("dut_serial", serial)  # lands on <testsuite> in JUnit XML
    return serial


@pytest.fixture(scope="session")
def psu():
    inst = RM.open_resource("TCPIP0::192.168.1.40::inst0::INSTR")
    yield inst
    inst.write("OUTP OFF")  # teardown runs even if a test failed
    inst.close()


@pytest.fixture(scope="session")
def dmm():
    inst = RM.open_resource("TCPIP0::192.168.1.50::inst0::INSTR")
    yield inst
    inst.close()
```

```python title="test_rail.py"
import pytest


@pytest.mark.parametrize("vin", [10.8, 12.0, 13.2])
def test_5v_rail_across_input(psu, dmm, vin, record_property):
    psu.write(f"VOLT {vin}")
    psu.write("OUTP ON")
    psu.query("*OPC?")
    vout = float(dmm.query("MEAS:VOLT:DC?"))
    record_property("vout_v", vout)  # the measured value, not just pass or fail
    assert 4.9 <= vout <= 5.1
```

```ini title="pytest.ini"
[pytest]
# record_property output fits the xunit1 schema; the default xunit2 warns
junit_family = xunit1
```

Run it with `pytest --dut-serial SN-0001 --junit-xml=results/SN-0001.xml`. The `junit_family` line matters: pytest warns that `record_property` is incompatible with the default `xunit2` family, while `record_testsuite_property` is compatible ([pytest docs](https://docs.pytest.org/en/stable/how-to/output.html)).

**Flow control and the rest.** `-x` stops at the first failure, markers select and skip tests, and [plugins](https://docs.pytest.org/en/latest/reference/plugin_list.html) add reruns and ordering. Units, limits as data, an operator interface and station deployment are yours to build, so a production line on pytest means building an in-house executive around it.

## TestStand: sequences, step types and process models

NI describes TestStand as a way to "interactively build and organize test sequences, then monitor execution to identify bottlenecks, and package sequences for deployment to your manufacturing floor" ([NI, What is TestStand](https://www.ni.com/en/shop/electronic-test-instrumentation/application-software-for-electronic-test-and-instrumentation-category/what-is-teststand.html)). Test code stays in "common programming languages such as NI LabVIEW, C/C++, .NET, and Python."

The model separates measuring from judging. A Numeric Limit Test step calls a code module that returns one measurement value, then compares that value to the limits stored in the step ([NI, Numeric Limit Test step](https://www.ni.com/docs/en-US/bundle/teststand-api-ref/page/tsref/numeric-limit-test-step.html)). The value is stored in `Step.Result.Numeric`, the limits in `Step.Limits.Low` and `Step.Limits.High`, and the comparison type, such as EQ, in `Step.Comp`. Pass/Fail Test and String Value Test steps follow the same pattern ([built-in step types](https://www.ni.com/docs/en-US/bundle/teststand/page/built-in-step-types.html)). A Python code module therefore just measures:

```python title="rail_checks.py"
# Code module for a TestStand Numeric Limit Test step.
# The step owns the limits and the comparison; this function only measures.
import pyvisa

_rm = pyvisa.ResourceManager()


def measure_vout(resource: str) -> float:
    dmm = _rm.open_resource(resource)
    try:
        return float(dmm.query("MEAS:VOLT:DC?"))
    finally:
        dmm.close()
```

The step's module settings bind the function's return value to `Step.Result.Numeric`; each adapter has its own settings panel. Limits live in the sequence file, so changing one is a sequence change, not a code change.

**Around the steps**, a process model defines the standard operations around every test. TestStand ships Sequential, Parallel and Batch models; Parallel and Batch test several UUTs at once ([NI, process models](https://www.ni.com/docs/en-US/bundle/teststand/page/teststand-process-models.html)). Operator interfaces, reports in HTML, XML, ATML and ASCII text, database logging, resource scheduling and deployment packages are all included.

**Platform and format.** NI's download page lists Windows as the supported operating system, including for the 2026 Q3 release ([NI download](https://www.ni.com/en/support/downloads/software-products/download.teststand.html)), and NI's [compatibility table](https://www.ni.com/en/support/documentation/compatibility/18/teststand-and-windows-os-compatibility.html) lists Windows 10 and 11 for recent versions. A `.seq` file can be binary, XML or INI, and NI recommends binary "for the fastest load times" unless the file must be viewable without the TestStand engine ([NI, improving TestStand performance](https://www.ni.com/en/support/documentation/supplemental/08/improving-teststand-system-performance.html)). Binary sequences can live in Git but cannot be reviewed line by line there. In TestStand 2026 Q3, NI's agent, Nigel, can create sequences from a specifications document you provide ([NI, Nigel](https://www.ni.com/en/shop/software-portfolio/nigel.html)); [test sequencers vs test agents](https://galoislabs.ai/blog/test-sequencer-vs-test-agent) covers what that changes.

## What do OpenHTF, OpenTAP, pytest and TestStand cost?

The open-source tools cost nothing to license; the cost is engineering time for the jobs they leave to you. Their licenses differ in what they ask back:

- **Apache-2.0 (OpenHTF) and MIT (pytest, PyVISA)** are permissive: ship the code inside proprietary test software and keep the notices. Apache-2.0 adds an express patent license from contributors ([Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0)).
- **MPL-2.0 (OpenTAP)** is copyleft at the file level. Changes you distribute to OpenTAP's own files stay under MPL-2.0; plugins in your own files can stay proprietary ([MPL 2.0 FAQ](https://www.mozilla.org/en-US/MPL/2.0/FAQ/)).

TestStand is licensed per developer, by subscription, and per station, by one-time deployment license. [TestStand alternatives](https://galoislabs.ai/blog/teststand-alternatives#how-much-does-teststand-cost-and-does-it-run-on-linux) lists NI's US list prices, read October 5, 2026, works through the cost for two developers and ten stations, and gives a migration plan.

## Which test executive fits which team?

**TestStand is the better choice when:**

- Your stations run Windows and your LabVIEW, C or .NET code modules already work.
- You test several UUTs at once and want the Parallel or Batch process model rather than building one.
- Operators need a finished interface, reports and database logging on day one.
- MES systems or customers already parse TestStand reports or ATML.
- Your deployment licenses are paid; Base Deployment is one-time and perpetual.

**OpenTAP is the better choice when:**

- Your team writes C#, or wants a TestStand-shaped sequencer under an open-source license.
- You may later want PathWave Test Automation's tools and Keysight support on the same engine.
- Stations run Linux or Docker headless, with plans executed by `tap run`.
- You want plugins versioned as packages and plans stored as XML that diffs.

**OpenHTF is the better choice when:**

- Your team writes Python and wants limits declared next to the measurement they judge.
- A JSON test record per unit fits your data pipeline.
- You can write the plugs for your instruments and the output callback for your database.

**pytest is the better choice when:**

- A small team owns a few development benches, and tests run in the same CI as firmware.
- Your engineers already write Python. In our [hiring study](https://galoislabs.ai/blog/hardware-test-hiring-study) of 1,021 hardware-test job postings, 54.8% named Python and 1.0% named TestStand.
- Pass, fail and a JUnit file are enough for now.

You can also mix them: pytest on development benches, OpenHTF or TestStand in production. Keep instrument access in one layer so the same driver code serves both; [SCPI instrument automation with Python](https://galoislabs.ai/blog/scpi-automation-python) shows how, and [open-source instrument control software](https://galoislabs.ai/blog/open-source-instrument-control) compares the driver libraries.

If your team has a working executive, keep it. If you are choosing one now, start with the language your engineers write and the station operating system, then count the last three jobs honestly.

## Where Galois fits

Galois overlaps with these tools in two places.

**Under them, as the instrument layer.** The Apache-2.0 galois-edge daemon runs on the bench PC and serves its instruments over the network. PyVISA code switches with one line, `pyvisa.ResourceManager("@galois")`, and each call routes through Galois Cloud to the daemon, so an OpenHTF plug or pytest fixture like those above reaches instruments on another machine without NI-VISA or USB and GPIB drivers where the test runs ([PyVISA backend docs](https://docs.galoislabs.ai/guides/pyvisa/)). Instruments with a matching profile also get typed methods; Galois ships 573 instrument profiles across 135 manufacturers ([instrument library](https://galoislabs.ai/instruments)).

**Beside them, as the executive.** In Galois, agents write the sequence and an engineer approves it. A sequence is versioned YAML built from typed steps, such as Numeric Limit, Pass/Fail, Loop and Condition, and each Numeric Limit step names its comparison:

```yaml title="rail-check.sequence.yaml"
steps:
  - name: "Measure 5 V rail"
    type: numeric_limit
    config:
      instrument_id: "dmm"
      command_name: "measure_voltage_dc"
      low_limit: 4.9
      high_limit: 5.1
      unit: "V"
      comparison: "GELE"   # low <= value <= high
```

A draft cannot run until approved, and an edited sequence needs approval again. Every run stores, per step, the SCPI command sent, the raw response, the measured value, the limits and the instrument, with the operator and DUT serial on the run ([product](https://galoislabs.ai/product)). The [TestStand comparison](https://galoislabs.ai/compare/teststand) goes line by line.

## How do I run this rail check in Galois with Évariste?

Every sketch above reads the 5 V rail on a DMM and passes it between 4.9 V and 5.1 V. The OpenHTF and pytest versions also set the supply, and the pytest version repeats the check at inputs of 10.8 V, 12.0 V and 13.2 V. Évariste, the agent in the Galois platform, builds that three-input check from a plain-English objective. Open it from the app sidebar (Ctrl+Shift+E) beside a project and state the task with the limits from your datasheet:

> Create a rail-check sequence with the supply at TCPIP0::192.168.1.40::inst0::INSTR and the DMM at TCPIP0::192.168.1.50::inst0::INSTR. At inputs of 10.8 V, 12.0 V and 13.2 V, set the supply, turn its output on, measure the 5 V rail and pass between 4.9 V and 5.1 V inclusive. Turn the supply output off at the end.

Ask "List connected instruments" to confirm both addresses; named profile commands such as `source_voltage` and `measure_voltage_dc` take the place of OpenHTF's plugs, OpenTAP's instrument plugins, the pytest fixtures and TestStand's code modules. Évariste then drafts the sequence in the format of `rail-check.sequence.yaml` above:

```yaml title="rail-check-vin.sequence.yaml"
steps:
  - name: "Set input to 10.8 V"
    type: action
    config:
      instrument_id: "psu"
      command_name: "source_voltage"
      parameters: { value: "10.8" }

  - name: "Enable output"
    type: action
    config:
      instrument_id: "psu"
      command_name: "output_on"

  - name: "Measure 5 V rail at 10.8 V"
    type: numeric_limit
    config:
      instrument_id: "dmm"
      command_name: "measure_voltage_dc"
      low_limit: 4.9
      high_limit: 5.1
      unit: "V"
      comparison: "GELE"

  # the 12.0 V and 13.2 V set and measure steps are elided

  - name: "Disable output"
    type: action
    config:
      instrument_id: "psu"
      command_name: "output_off"
```

The draft does not run until an engineer approves it. Check each step's instrument, the three setpoints, both limits against the datasheet, and that the run ends with `output_off`, as the `Psu` plug and `psu` fixture teardowns do; [how to review an AI-generated test plan](https://galoislabs.ai/blog/review-ai-generated-test-plan) lists what else to check. Ask Évariste for changes in conversation or edit in the sequence builder. Every change is a new version with history and diffs, and an approved sequence can be production-locked.

Wiring the supply and DMM to the DUT and bench safety stay with you. Start the run; galois-edge executes it on the bench and Monitor shows the channels live. If you send a command flagged as dangerous straight from the conversation, Évariste asks you to confirm it first.

The operator and DUT serial land on the run, the job `prompt_for_test_start()` and `--dut-serial` do in the sketches, and each step keeps its measured value, limits, verdict and raw command and response, which a result listener or JUnit file would otherwise carry. Ask Évariste which steps failed or passed close to a limit, how the rail held across the three inputs, or how this unit compares with earlier runs; its answers cite the runs and notes they draw on. "Generate a test report from the last run" produces a PDF or HTML report from a LaTeX template, editable in the report editor, and results can be shared to Slack.

Your part is the objective, the limits from the datasheet, the review, the approval, the bench setup and safety. The plugs, fixtures, step classes and code modules, the DUT serial option, the output callback, error handling, logging, and the scripts that turn JSON, CSV or JUnit files into plots and reports are no longer yours to write or maintain. [AI test automation for hardware benches](https://galoislabs.ai/blog/ai-test-automation-hardware) covers where agents fit in the rest of the test workflow.

| Step            | Code path (this guide)                                           | Galois with Évariste                                            |
| --------------- | ---------------------------------------------------------------- | --------------------------------------------------------------- |
| Inventory       | VISA address per plug, fixture, bench setting or module argument | "List connected instruments" across the team's edges            |
| Driver          | `Psu` and `Dmm` plugs, `BenchDmm`, fixtures, SCPI strings        | Library profile, or one generated from the manual and reviewed  |
| Test and limits | `in_range`, step properties, `assert` or `Step.Limits`           | Plain-English objective, same limits; Évariste drafts it        |
| Review          | Code review of Python or C#; `.seq` in the editor                | Draft reviewed, edited with versioned diffs, approved           |
| DUT serial      | `prompt_for_test_start()` or `--dut-serial`                      | Operator and DUT serial recorded on the run                     |
| Run             | `test.execute`, `tap run`, `pytest` or an operator UI            | Run through galois-edge, watched live in Monitor                |
| Record          | JSON callback, result listeners, JUnit XML or TestStand reports  | Per-step value, limits, pass/fail, raw command and response     |
| Interpret       | Your own queries over the stored results                         | Évariste flags failures and near-limit passes and compares runs |
| Report          | Your report script, or TestStand's reports                       | PDF or HTML report, editable, shareable to Slack                |

## Frequently asked questions

### What is the difference between OpenHTF and OpenTAP?

OpenHTF (Apache-2.0) is a Python library: a test is a Python program built from phases, measurements carry their own limits, and plugs wrap the DUT and instruments. OpenTAP (MPL-2.0) is a sequencer engine on .NET: a test plan is an XML file of test steps written as C# classes, instruments and DUTs are resources, and result listener plugins store the results. OpenHTF fits teams that write Python; OpenTAP fits teams that write C# or may want Keysight's PathWave tools on the same engine.

### Can pytest be used for hardware testing?

Yes. Fixtures open and close instruments, parametrize runs one test across input conditions, and --junit-xml writes a results file a CI server can read. Units, limits, DUT serial numbers and an operator screen are not part of pytest's model, so you add them with fixtures, record_property, plugins or your own code. It suits development benches and small teams better than a multi-station production line.

### Does OpenTAP run on Linux?

Yes. The opentap.io downloads page offers OpenTAP for Windows 10 or later, Ubuntu 24.04 or later, and macOS 12 or later on Apple silicon, and links an installation guide for Docker. tap run executes a .TapPlan file from the command line, and the open-source TUI package gives a terminal editor. Keysight's graphical Developer System is a separate install with a commercial or community license.

### Is OpenHTF still maintained?

Commits to the google/openhtf repository continue into October 2026, and the latest release on PyPI is 1.6.1 from March 2025. The README states it is not an official Google product. Pin a release or a commit for each production station so every station runs the same framework code.

### Do I have to write test code to use a test executive?

With OpenHTF and pytest, yes: tests are Python. OpenTAP test plans are built from plugins you write, and TestStand sequences call code modules, which Nigel in TestStand 2026 Q3 can map or stub from a specification. In Galois, Évariste, the agent in the Galois platform, drafts a versioned sequence from a plain-English objective with your limits, such as the 5 V rail between 4.9 V and 5.1 V at inputs of 10.8, 12.0 and 13.2 V; it runs through galois-edge only after an engineer approves it.
