---
title: Run Hardware Tests in CI with GitHub Actions
description: "Hardware-in-the-loop CI in GitHub Actions: a self-hosted runner on the bench PC, pytest instrument fixtures, firmware flashing, JUnit results, safe teardown."
url: https://galoislabs.ai/blog/hardware-tests-in-ci
author: Alex Hernandez
author_url: https://galoislabs.ai/blog/authors/alex-hernandez
published: "2026-06-05"
topic: Instrument automation
publisher: Galois Labs
---

# Run hardware tests in CI with GitHub Actions and a self-hosted bench runner

![The rear panel of a bench instrument close up: LAN, USB and GPIB ports, with a GPIB plug just detached.](https://galoislabs.ai/blog/figures/bench-3.light.webp)

*FIG. 1 — REAR PANEL, THREE INTERFACES*

To run hardware tests in CI with GitHub Actions, register a self-hosted runner on the PC wired to the bench and route a job to it with `runs-on` labels. That job flashes the firmware build, runs pytest against instrument fixtures, publishes JUnit XML as an artifact, and returns every instrument to a safe state, pass or fail.

This guide builds that workflow for one bench: a Rigol DP832 supply powers the board, a Keysight 34461A multimeter reads its 3.3 V rail, and an ST-LINK probe flashes it. The open-source galois-edge daemon runs on the bench PC as a service and owns the instruments; the CI job runs pytest with the Python SDK against it. [A later section](https://galoislabs.ai/blog/hardware-tests-in-ci#how-do-i-run-these-bench-tests-in-galois-with-évariste) runs the same checks with Évariste, the agent in the Galois platform, and the last section compares the daemon with plain PyVISA in the job.

## What does hardware-in-the-loop CI need?

A GitHub-hosted runner is a virtual machine in GitHub's cloud. It can compile firmware, but it cannot reach a USB supply or a GPIB meter in your lab. Hardware jobs need a self-hosted runner on a machine that can. The host must make [outbound HTTPS connections over port 443](https://docs.github.com/en/actions/reference/runners/self-hosted-runners), and the runner supports recent Linux distributions, Windows 10 and 11, Windows Server, and macOS 11 or later, so the bench PC usually qualifies.

The bench sees one job at a time, and only after the build passes:

| Stage             | Runs on                  | Mechanism                                               |
| ----------------- | ------------------------ | ------------------------------------------------------- |
| Build firmware    | GitHub-hosted runner     | Your toolchain, then `actions/upload-artifact`          |
| Reserve the bench | GitHub                   | One runner per bench, runner labels, job concurrency    |
| Drive instruments | Bench PC                 | galois-edge daemon as a service (or PyVISA in the job)  |
| Flash the board   | Self-hosted runner       | OpenOCD or your vendor's programmer CLI                 |
| Run tests         | Self-hosted runner       | pytest with session-scoped instrument fixtures          |
| Publish results   | GitHub                   | JUnit XML, artifacts, a job summary                     |
| Fail safe         | Bench PC and instruments | Fixture teardown, an `if: always()` step, supply limits |

## How do I set up a self-hosted runner on a test bench?

Add the runner from the repository's Actions settings to get a registration token, then configure it on the bench PC with a label that names the bench. The [labels guide](https://docs.github.com/en/actions/hosting-your-own-runners/managing-self-hosted-runners/using-labels-with-self-hosted-runners) shows the `--labels` flag, and the [service guide](https://docs.github.com/en/actions/hosting-your-own-runners/managing-self-hosted-runners/configuring-the-self-hosted-runner-application-as-a-service) installs the runner as a systemd service so it survives reboots:

```sh title="bench PC, in the unpacked runner directory"
./config.sh --url https://github.com/acme/widget-fw --token <REGISTRATION_TOKEN> --labels bench-a
sudo ./svc.sh install    # or: ./svc.sh install USERNAME to run as another user
sudo ./svc.sh start
```

The runner also gets default labels: `self-hosted`, its OS and its architecture. A runner must carry every label in a job's `runs-on` list to be eligible ([using self-hosted runners in a workflow](https://docs.github.com/en/actions/hosting-your-own-runners/managing-self-hosted-runners/using-self-hosted-runners-in-a-workflow)), so `[self-hosted, linux, bench-a]` reaches this bench and no other.

**Treat repository write access as bench access.** Anyone who can change a workflow can drive the supply. GitHub's [security hardening guide](https://docs.github.com/en/actions/security-guides/security-hardening-for-github-actions) says self-hosted runners "should almost never be used for public repositories," that they "can be persistently compromised by untrusted code in a workflow," and that secrets passed as command-line arguments are visible to other jobs through `ps`. Keep the repository private, scope the runner to the repositories that need it, and pass secrets through environment variables.

Next, install the daemon as its own service. `galois-edge install` registers a systemd unit that runs as a `galois-edge` account by default ([CLI reference](https://docs.galoislabs.ai/reference/cli/)), so the runner's account needs the debug probe and nothing else. A CI bench needs few [configuration keys](https://docs.galoislabs.ai/getting-started/configuration/):

```sh title="/etc/galois-edge/config.env"
EDGE_NAME=bench-a
GRPC_PORT=50051
LAN_INSTRUMENTS=TCPIP::192.168.1.50::INSTR,TCPIP::192.168.1.60::INSTR
INBOUND_AUTH_TOKEN=glc_internal_<32 or more random characters>
WS_ENABLED=false
MCP_ENABLED=false
LOG_LEVEL=info
```

Set `INBOUND_AUTH_TOKEN`. The external gRPC port is also bound on `0.0.0.0`, and with the token empty, the network is the only boundary. With it set, every RPC except `Ping` needs the bearer token, stored as a repository secret for the job. The token gates gRPC; the WebSocket and MCP listeners rely on the network boundary, so a CI-only bench turns them off. `galois-edge status` exits 1 when the daemon's engine is unreachable: a one-line preflight.

## What does the GitHub Actions workflow look like?

Two jobs: `build` on a hosted runner, `bench` on the bench runner once the build passes.

```yaml title=".github/workflows/hil.yml"
name: hil

on:
  push:
    branches: [main]
  pull_request:

permissions:
  contents: read

jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - run: make firmware                  # writes build/fw.elf
      - uses: actions/upload-artifact@v7
        with:
          name: firmware
          path: build/fw.elf
          if-no-files-found: error

  bench:
    needs: build
    # pushes and same-repository pull requests only: fork code never reaches the bench
    if: github.event_name == 'push' || github.event.pull_request.head.repo.full_name == github.repository
    runs-on: [self-hosted, linux, bench-a]
    timeout-minutes: 30
    concurrency:
      group: bench-a
      queue: max
    env:
      GALOIS_EDGE: localhost:50051
      GALOIS_EDGE_TOKEN: ${{ secrets.BENCH_A_EDGE_TOKEN }}
      FIRMWARE: build/fw.elf
    steps:
      - uses: actions/checkout@v7
      - uses: actions/download-artifact@v8
        with:
          name: firmware
          path: build
      - name: Bench preflight
        run: galois-edge status             # exits 1 if the daemon is down
      - name: Python environment
        run: |
          python3 -m venv .venv
          .venv/bin/pip install -r test/hil/requirements.txt   # pytest and the Galois Python SDK (see the SDK guide)
      - name: Hardware tests
        # exec: pytest replaces the shell, so a cancel's SIGINT reaches pytest
        run: exec .venv/bin/pytest test/hil -m hardware --junit-xml=results/junit.xml
      - name: Safe state
        if: always()
        timeout-minutes: 2
        run: .venv/bin/python test/hil/safe_state.py
      - name: Summary
        if: always()
        run: .venv/bin/python test/hil/summary.py results/junit.xml >> "$GITHUB_STEP_SUMMARY"
      - uses: actions/upload-artifact@v7
        if: always()
        with:
          name: hil-results-${{ github.run_id }}-${{ github.run_attempt }}
          path: results/
          retention-days: 30
```

Steps without a condition run only while everything before them succeeded, so a failed preflight stops the job before anything is powered; the `if: always()` steps run regardless. `timeout-minutes` keeps a hung test from holding the bench. Keep the runner application current with the actions: `checkout@v7`, `upload-artifact@v7` and `download-artifact@v8` run on Node 24, which needs runner 2.327.1 or later per the [`actions/checkout` README](https://github.com/actions/checkout) and the [`upload-artifact` v6](https://github.com/actions/upload-artifact/releases/tag/v6.0.0) and [`download-artifact` v7](https://github.com/actions/download-artifact/releases/tag/v7.0.0) release notes.

## How do CI jobs reserve bench instruments?

Three layers, each covering a different collision:

| Layer                | Stops                                              | Scope                                  | Mechanism                                        |
| -------------------- | -------------------------------------------------- | -------------------------------------- | ------------------------------------------------ |
| One runner per bench | Two jobs driving one bench                         | Every repository routed to that runner | GitHub assigns jobs only to online, idle runners |
| Concurrency group    | A new push canceling the run waiting for the bench | One repository                         | `concurrency: { group: bench-a, queue: max }`    |
| The daemon           | Two processes opening one instrument               | The bench                              | One process owns each instrument connection      |

**The runner is the reservation.** GitHub sends a job only to ["an online and idle runner"](https://docs.github.com/en/actions/reference/runners/self-hosted-runners) matching its labels and groups, so one runner per bench means one job per bench. With no matching runner online, the job waits, and fails after 24 hours in the queue. Never register a second runner with the same bench label.

**The concurrency group sets the queue policy.** By default a group holds one pending run, and a newly queued run cancels and replaces it, per GitHub's [concurrency docs](https://docs.github.com/en/actions/how-tos/write-workflows/choose-when-workflows-run/control-workflow-concurrency). On a busy bench, commit B's run quietly replaces commit A's while both wait: fine if only the latest `main` matters, wrong if every merge needs a bench result. `queue: max` keeps up to 100 pending runs, first in first out by the time each started waiting, though GitHub does not guarantee the order. It cannot be combined with `cancel-in-progress: true`, which a bench should not use anyway. Groups are scoped to one repository, so when several repositories share a bench, the single runner serializes them.

**People are the remaining collision.** The daemon owns every instrument connection, but a notebook command or a front-panel change between two CI commands still changes the measurement. When someone needs the bench, stop the runner service (`sudo ./svc.sh stop`); queued jobs wait.

## How do I write pytest fixtures for lab instruments?

The fixtures live in `conftest.py`, so test files only state intent. The SDK calls follow the [Python SDK guide](https://docs.galoislabs.ai/guides/python-sdk/), and the DP832 commands come from Rigol's [DP800 programming guide](https://beyondmeasure.rigoltech.com/acton/attachment/1579/f-03a1/1/-/-/-/-/DP800%20Programming%20Guide.pdf):

```python title="test/hil/conftest.py"
"""bench-a: a Rigol DP832 powers the board, a Keysight 34461A reads its rails."""
import os
import signal
import subprocess
import time

import galois
import pytest

EDGE = os.environ.get("GALOIS_EDGE", "localhost:50051")
TOKEN = os.environ.get("GALOIS_EDGE_TOKEN")
FIRMWARE = os.environ.get("FIRMWARE", "build/fw.elf")
CH, V_IN, I_CLAMP, I_TRIP = "CH1", 12.0, 0.25, 0.20  # from the board's power budget


def _interrupt(signum, frame):
    raise KeyboardInterrupt(f"received signal {signum}")


# Python turns only SIGINT into KeyboardInterrupt. Route SIGTERM (a kill, a
# service manager stop) the same way, so fixture teardown runs for both.
signal.signal(signal.SIGTERM, _interrupt)


@pytest.fixture(scope="session")
def edge():
    kwargs = {"auth_token": TOKEN} if TOKEN else {}
    with galois.Edge.connect(EDGE, **kwargs) as edge:
        edge.ping()  # a dead daemon errors every test; nothing is skipped
        yield edge


def one(edge, model, record):
    found = [i for i in edge.instruments() if i.model.startswith(model)]
    if len(found) != 1:
        pytest.fail(f"expected one {model} on the bench, found {len(found)}", pytrace=False)
    inst = edge.instrument(found[0].id)
    record(f"{model}.idn", inst.query("*IDN?").strip())  # serial and firmware into JUnit
    return inst


@pytest.fixture(scope="session")
def dmm(edge, record_testsuite_property):
    return one(edge, "3446", record_testsuite_property)


@pytest.fixture(scope="session")
def psu(edge, record_testsuite_property):
    psu = one(edge, "DP832", record_testsuite_property)
    for command in (f":OUTP {CH},OFF", f":OUTP:OCP:CLEAR {CH}",
                    f":APPL {CH},{V_IN},{I_CLAMP}",
                    f":OUTP:OCP:VAL {CH},{I_TRIP}", f":OUTP:OCP {CH},ON"):
        psu.write(command)  # limits are armed while the output is still off
    yield psu
    psu.write(f":OUTP {CH},OFF")  # runs on pass, fail, error and interrupt
```

Four choices carry the design:

- **Session scope.** Instruments are found and configured once per run and torn down at session end.
- **Fail, never skip.** A missing instrument or a dead daemon fails the run. A fixture that calls `pytest.skip()` when the bench is unreachable turns an unplugged meter into a green check mark.
- **Identity in the record.** `record_testsuite_property` writes each instrument's `*IDN?` string, with serial number and firmware revision, into the JUnit file, so every result names the instruments that produced it: the starting point for [hardware test traceability](https://galoislabs.ai/blog/hardware-test-traceability).
- **Limits before power.** The clamp and the overcurrent trip are set while the output is off. The trip sits below the clamp, so the DP832 turns the output off above 0.20 A and the clamp bounds the current until it does.

The tests stay short:

```python title="test/hil/test_rails.py"
import pytest

pytestmark = pytest.mark.hardware


def test_input_current_at_idle(dut, record_property):
    volts, amps, _watts = (float(x) for x in dut.query(":MEAS:ALL? CH1").split(","))
    record_property("input_a", amps)
    assert amps < 0.150, f"idle input current {amps * 1e3:.1f} mA at {volts:.2f} V"


def test_3v3_rail(dut, dmm, record_property):
    volts = float(dmm.query("MEAS:VOLT:DC?"))  # meter wired to the 3V3 test point
    record_property("rail_3v3_v", volts)
    assert 3.25 <= volts <= 3.35, f"3.3 V rail at {volts:.4f} V"
```

Register the marker in `test/hil/pytest.ini` (`markers = hardware: needs bench-a`) with `--strict-markers`, so a misspelled marker fails instead of deselecting tests; hosted runners run `pytest -m "not hardware"`. The limits are placeholders: derive yours from regulator tolerance and meter uncertainty, as in [test limits from a KiCad netlist](https://galoislabs.ai/blog/kicad-netlist-test-points). The [SCPI automation guide](https://galoislabs.ai/blog/scpi-automation-python) covers the error-queue checks and `*OPC?` synchronization production fixtures add.

## How do I flash firmware in a hardware CI job?

Build where builds are cheap, flash where the board is. The `build` job uploads `fw.elf`, and the bench job downloads it with `actions/download-artifact@v8`. This board takes its power from the supply, so flashing belongs inside the pytest session, after power-up, under the same current limit and teardown:

```python title="test/hil/conftest.py (continued)"
@pytest.fixture(scope="session")
def dut(psu, record_testsuite_property):
    psu.write(f":OUTP {CH},ON")
    time.sleep(0.5)  # dwell for the board's rails, not instrument synchronization
    mode = psu.query(f":OUTP:MODE? {CH}").strip()
    if mode != "CV":  # CC or UR at idle means the board is drawing the clamp current
        pytest.fail(f"supply in {mode} mode after power-up", pytrace=False)
    subprocess.run(
        ["openocd", "-f", "interface/stlink.cfg", "-f", "target/stm32f3x.cfg",
         "-c", f"program {FIRMWARE} verify reset exit"],
        check=True, timeout=120,
    )
    record_testsuite_property("firmware.sha", os.environ.get("GITHUB_SHA", "local"))
    yield psu
```

OpenOCD's `program` command [programs, verifies and resets the target](https://openocd.org/doc/html/Flash-Programming.html), then `exit` ends OpenOCD. ELF, HEX and S19 files carry their addresses; a raw `.bin` needs the flash address as the last argument. Swap in the configs for your probe and part, or your vendor's programmer CLI, and give the runner's account the probe through a udev rule or group.

`timeout=120` keeps a hung probe from holding the bench. Either failure, `CalledProcessError` or `TimeoutExpired`, errors the fixture, and the [pytest fixture docs](https://docs.pytest.org/en/stable/how-to/fixtures.html) define what follows: a fixture that raises before yielding skips its own teardown, "but, for every fixture that has already run successfully for that test, pytest will still attempt to tear them down." The `psu` fixture has already yielded, so the output goes off. That is why power-off lives in `psu`, not `dut`.

`verify` checks flash contents against the file, not that the board booted the right build. If your firmware reports a version over its console, make the first test compare it with the commit SHA embedded at build time.

## How do I publish hardware test results as JUnit and artifacts?

`pytest --junit-xml=results/junit.xml` writes the file CI tools read ([pytest output docs](https://docs.pytest.org/en/stable/how-to/output.html)). Two fixtures add the bench context:

- `record_testsuite_property` is session-scoped, writes suite-level properties, and stays compatible with the current xunit schema. Use it for instrument identity and the firmware SHA.
- `record_property` adds key-value pairs to a single test case, which suits measured values. pytest warns that it "will break schema verifications for the latest JUnitXML schema," which matters only if your CI tooling validates the schema.

The upload uses `if: always()`, so a failed run still publishes its evidence. The artifact name includes `run_attempt` because [artifacts from `upload-artifact@v4` onward](https://github.com/actions/upload-artifact) are immutable and names must be unique; `retention-days` accepts 1 to 90 unless the repository allows more. pytest writes the JUnit file even after a KeyboardInterrupt, when it exits with code 2. It [returns 5 when no tests were collected](https://docs.pytest.org/en/stable/reference/exit-codes.html), which includes `-m hardware` matching nothing. Leave that failure visible; never append `|| true` to the test step.

For a readable result on the run page, write Markdown to `$GITHUB_STEP_SUMMARY`, up to [1 MiB per step](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-commands):

```python title="test/hil/summary.py"
"""Print a Markdown summary of a pytest JUnit file for $GITHUB_STEP_SUMMARY."""
import sys
import xml.etree.ElementTree as ET
from pathlib import Path

path = Path(sys.argv[1])
if not path.exists():
    print(f"No JUnit file at `{path}`: the test step did not start.")
    sys.exit(0)

root = ET.parse(path).getroot()
suites = [root] if root.tag == "testsuite" else root.findall("testsuite")

print("| Bench property | Value |\n|---|---|")
for prop in (p for s in suites for p in s.findall("properties/property")):
    print(f"| {prop.get('name')} | `{prop.get('value')}` |")

print("\n| Test | Result | Seconds |\n|---|---|---|")
for case in (c for s in suites for c in s.iter("testcase")):
    result = next((el.tag for el in case if el.tag in ("failure", "error", "skipped")), "passed")
    print(f"| `{case.get('name')}` | {result} | {float(case.get('time', 0)):.2f} |")
```

Instrument serials, firmware revisions and the commit under test sit above the results. For release reports built from these files, see the [DVT test report template](https://galoislabs.ai/blog/dvt-test-report-template).

## How do I make hardware tests in CI fail safe?

A hardware test can fail with a supply at 12 V into a shorted board. Plan for each way a run can end:

| How the run ends               | What turns the bench off                                   | Its limit                                            |
| ------------------------------ | ---------------------------------------------------------- | ---------------------------------------------------- |
| A test fails                   | Session fixture teardown                                   | None                                                 |
| A fixture fails mid-setup      | Teardown of fixtures that already yielded                  | Limits and power-off must live in an earlier fixture |
| The job is canceled            | SIGINT, raised as KeyboardInterrupt, then fixture teardown | Teardown must finish within 7.5 s                    |
| pytest is killed               | The `if: always()` safe-state step                         | Needs the runner and daemon alive                    |
| The PC, network or daemon dies | The supply's clamp and overcurrent trip                    | Nothing in software helps                            |

**Cancellation is a timed sequence.** GitHub's [workflow cancellation reference](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-cancellation) says the runner sends SIGINT to the step's entry process, which for a `run` step is the shell; "if the process doesn't exit within 7500 ms," it sends SIGTERM, waits 2,500 ms, then kills the process tree. On Linux the runner [signals only that PID](https://github.com/actions/runner/blob/main/src/Runner.Sdk/ProcessInvoker.cs), so the workflow `exec`s pytest to make it the entry process. SIGINT then arrives as `KeyboardInterrupt`; pytest stops the session and still [tears down every active fixture](https://github.com/pytest-dev/pytest/blob/main/src/_pytest/runner.py), so the output-off command runs. That gives teardown 7.5 seconds: a few commands, not a ramp-down. Python [translates only SIGINT](https://docs.python.org/3/library/signal.html) into `KeyboardInterrupt`, so a SIGTERM that arrives on its own, from a `kill` or a service manager, would end the process with no teardown at all; `conftest.py` maps it to the same exception. [Electronic load automation](https://galoislabs.ai/blog/electronic-load-automation) covers the same signal problem for discharge scripts.

**The safe-state step is the second line.** GitHub's `always()` status function makes a step execute and ["returns `true`, even when canceled"](https://docs.github.com/en/actions/reference/workflows-and-actions/expressions), so the step runs after a failure, a cancellation or a crash of the test step. The same page warns that an `always()` task that fails critically can hang the workflow until it times out, so give the step its own `timeout-minutes: 2`. The script assumes nothing about what pytest got through:

```python title="test/hil/safe_state.py"
"""Workflow step with if: always(). Outputs off on every supply, whatever pytest did."""
import os

import galois

SAFE = {"DP832": [":OUTP CH1,OFF", ":OUTP CH2,OFF", ":OUTP CH3,OFF"]}

token = os.environ.get("GALOIS_EDGE_TOKEN")
kwargs = {"auth_token": token} if token else {}
with galois.Edge.connect(os.environ.get("GALOIS_EDGE", "localhost:50051"), **kwargs) as edge:
    for found in edge.instruments():
        for command in SAFE.get(found.model, []):
            edge.instrument(found.id).write(command)
            print(f"{found.model} {found.id}: {command}")
```

If it cannot reach the daemon, the step goes red. Treat that as a reason to walk to the bench, not as flaky noise.

**The instrument is the last line.** No software survives a crashed PC or a dropped network; the clamp and overcurrent trip live in the supply, which is why they are armed first. Slow ramps, such as magnet fields, are marked `requires_sweep` in their daemon profile and run as daemon-side sweeps that continue if the client drops. If your tests start sweeps, stop them in the safe-state step too; `Sweep.cancel()` runs the profile's abort SCPI.

## How do I run these bench tests in Galois with Évariste?

Évariste reaches this bench through the same galois-edge daemon once `galois-edge setup` registers it with Galois, so the daemon setup above carries over; what changes is who writes `conftest.py`, `test_rails.py` and `summary.py`. Open Évariste from the app sidebar (Ctrl+Shift+E) beside a project, ask "List connected instruments" to confirm the DP832 and 34461A on edge `bench-a`, then state the objective with the limits from your power budget and datasheet:

> Create a sequence for bench-a. On the Rigol DP832, turn CH1 off, clear any overcurrent trip, set 12 V with a 0.25 A current limit, and arm overcurrent protection at 0.20 A before the output turns on. Turn CH1 on, wait 500 ms, and fail unless CH1 is in CV mode. Pass if idle input current is below 150 mA and the Keysight 34461A reads the 3.3 V rail between 3.25 V and 3.35 V. Turn CH1 off at the end.

Évariste reads both profiles and drafts a sequence whose steps call named commands, not raw SCPI. For an instrument outside the library, it generates a profile from the uploaded programming manual; after your review, Évariste deploys it to the edge and binds it to the instrument. An excerpt:

```yaml title="bench_a_rails.yaml (excerpt)"
name: "bench-a: idle current and 3.3 V rail"
steps:
  # select CH1, CH1 off, clear OCP, 12 V, 0.25 A limit: action steps elided
  - name: "OCP level 0.20 A"
    type: action
    config:
      instrument_id: "psu"
      command_name: "current_protection_level"
      parameters: { value: "0.20" }

  - name: "Arm OCP"
    type: action
    config:
      instrument_id: "psu"
      command_name: "current_protection_state"
      parameters: { state: "ON" }

  - name: "CH1 on"
    type: action
    config:
      instrument_id: "psu"
      command_name: "output_state"
      parameters: { state: "ON" }

  - name: "Rails settle"
    type: wait
    config: { duration_ms: 500 }

  - name: "Supply in CV mode"
    type: string_value
    config:
      instrument_id: "psu"
      command_name: "output_mode"
      expected_value: "CV"

  - name: "Idle input current"
    type: numeric_limit
    config:
      instrument_id: "psu"
      command_name: "measure_current"
      high_limit: 0.150
      unit: "A"
      comparison: "GELT"

  - name: "3.3 V rail"
    type: numeric_limit
    config:
      instrument_id: "dmm"
      command_name: "measure_voltage_dc"
      low_limit: 3.25
      high_limit: 3.35
      unit: "V"
      comparison: "GELE"

  # final action step: CH1 off
```

The sequence lands as a draft that does not run until an engineer approves it. Check it against the fixtures: CH1 selected first, limits armed while the output is off, the trip below the clamp, the same two limits, CH1 off last. [How to review an AI-generated test plan](https://galoislabs.ai/blog/review-ai-generated-test-plan) lists what else to check. Ask for changes in conversation or edit in the sequence builder; every change is a new version with history and diffs, and a production version can be locked.

Wiring, the probe and flashing stay with you: flash the build before the run, with the board powered under the same CH1 limits. Stop the runner service first, as for anyone taking the bench. Start the run, or ask Évariste to start it with the board's serial; galois-edge executes it and Monitor shows the channels live. Commands sent individually from the conversation ask for confirmation when flagged as dangerous. A failed check does not stop the run, so the final step still turns CH1 off; if a run ends early, switch CH1 off at the bench.

Each step records its measured value, limits, pass or fail, raw command and response, instrument, operator, DUT serial and timestamps: what the JUnit properties and `summary.py` assemble here. Ask Évariste which steps failed or passed close to a limit, or how this board's idle current compares with the last run; answers cite their runs and notes. "Generate a test report from the last run" builds a PDF or HTML report from a LaTeX template, shareable to Slack; add the firmware build you flashed in the report editor.

Your job is the objective, limits, review, approval and physical safety. The fixtures, tests and summary script are no longer yours to maintain; [AI test automation for hardware benches](https://galoislabs.ai/blog/ai-test-automation-hardware) covers the agent's wider role.

| Step                | Code path (this guide)                 | Galois with Évariste                                              |
| ------------------- | -------------------------------------- | ----------------------------------------------------------------- |
| Find instruments    | `one()` over `edge.instruments()`      | "List connected instruments"                                      |
| Driver              | Raw SCPI in fixtures and tests         | Library profiles, or one generated from the manual                |
| Define tests        | `conftest.py`, `test_rails.py`         | Plain-English objective; Évariste drafts the sequence             |
| Limits before power | `psu` fixture, output off              | Action steps before CH1 on                                        |
| Review              | Pull request review                    | Versioned draft, reviewed and approved                            |
| Flash               | OpenOCD in the `dut` fixture           | OpenOCD or vendor tool, before the run                            |
| Start               | A push or same-repository pull request | The engineer, or Évariste on request                              |
| Run                 | pytest on the bench runner             | galois-edge, watched in Monitor                                   |
| Record              | JUnit XML with `*IDN?` and values      | Per-step value, limits, raw I/O, instrument, operator, DUT serial |
| Interpret           | Job summary, assertion messages        | Évariste flags failures and near-limit passes, compares runs      |
| Safe state          | `psu` teardown, `safe_state.py`        | Final CH1-off step; supply clamp and trip                         |
| Report              | `summary.py`, artifacts                | Generated PDF or HTML report, shareable to Slack                  |

## Should the CI job use PyVISA directly or a bench daemon?

PyVISA in the job is a sound choice when one repository owns one bench and nobody else uses the instruments. Replace the `edge` fixture with a `pyvisa.ResourceManager()`, open the two resources, and keep everything else. PyVISA-py needs no vendor VISA for LAN instruments ([PyVISA on Linux](https://galoislabs.ai/blog/pyvisa-linux-instrument-control) covers the other buses), the job depends on one fewer service, and the patterns in the [SCPI automation guide](https://galoislabs.ai/blog/scpi-automation-python) apply directly. The [Galois and PyVISA comparison](https://galoislabs.ai/compare/pyvisa) shows where the line falls.

A daemon pays off when the bench outlives the job. Instrument connections stay under one owner across runs, and the job's environment needs the Galois Python SDK and nothing else: no VISA library, GPIB driver or USB permissions on the runner's account. Between CI runs, engineers reach the same instruments from notebooks with the same SDK.

Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record.

Galois ships 573 instrument profiles across 135 manufacturers in its [instrument library](https://galoislabs.ai/instruments), and instruments without a profile still take raw SCPI, as in the fixtures above. The daemon is Apache-2.0 and also serves agents over Model Context Protocol (MCP); the CI config above turns that listener off, and [MCP for lab instruments](https://galoislabs.ai/blog/mcp-lab-instruments) covers what to weigh before turning it back on. [Galois](https://galoislabs.ai/product) runs on-prem or air-gapped as well as in the cloud ([deployment options](https://galoislabs.ai/deployment)).

Start CI with the checks that decide whether a board moves to the next build, then grow the suite; [EVT, DVT and PVT testing](https://galoislabs.ai/blog/evt-dvt-pvt-testing) covers which checks those are at each stage.

## Frequently asked questions

### Can GitHub-hosted runners run hardware tests?

Not against instruments on your bench. A GitHub-hosted runner is a virtual machine in GitHub's cloud, so it has no path to the USB, GPIB, serial or lab-network instruments wired to your board. Keep firmware builds, unit tests and simulations on hosted runners, and route the jobs that need hardware to a self-hosted runner on the PC next to the bench with runs-on labels.

### Is it safe to use a self-hosted runner on a public repository?

GitHub's security guide says self-hosted runners should almost never be used for public repositories, because any user can open a pull request and compromise the environment. For a bench runner, use a private repository, put the runner in a group limited to the repositories that need it, and guard the bench job so pull requests from forks never reach it.

### How do I stop two CI runs from using the same bench at once?

Register one runner per bench. GitHub assigns a job only to an online, idle runner whose labels match, so a single runner runs one job on the bench at a time. Add a job-level concurrency group named after the bench with queue: max, so a new push waits in line instead of canceling the run already waiting for the bench.

### What happens to instruments when a GitHub Actions job is canceled?

The runner sends SIGINT to the step's entry process, then SIGTERM if it has not exited within 7,500 ms, then kills the process tree 2,500 ms later. For a run step the entry process is the shell, so start pytest with exec. SIGINT then arrives as KeyboardInterrupt, and pytest stops and runs fixture teardown; keep that teardown to a few commands. Map SIGTERM to the same exception for kills that arrive without SIGINT, add an if: always() step that turns outputs off, and arm the supply's current limit and overcurrent protection before any output turns on.

### Can I run the DP832 and 34461A bench checks without writing pytest fixtures?

Yes. In Galois, you give Évariste, the agent in the Galois platform, the objective and limits in plain English: DP832 CH1 at 12 V with a 0.25 A clamp and a 0.20 A overcurrent trip, idle current below 150 mA, and the 3.3 V rail between 3.25 V and 3.35 V on the 34461A. It drafts a versioned sequence that runs only after an engineer approves it. An engineer, or Évariste on request, starts each run, where the CI job starts on a push. The run executes on the bench through galois-edge with every step recorded, and Évariste generates the report. Flashing the build and physical safety stay with the engineer.
