Declarative instrument drivers: why YAML profiles beat hand-written classes

By Alex Hernandez · · 14 min read

View Markdown
The rear panel of a bench instrument close up: LAN, USB and GPIB ports, with a GPIB plug just detached.
FIG. 1 — REAR PANEL, THREE INTERFACES

A declarative instrument driver describes an instrument's commands as data instead of code: each command's name, command string, typed parameters, units, limits, and return type, in a file that a runtime reads. Because the driver is data, it can be diffed line by line, generated from the programming manual, validated before it reaches hardware, and navigated by agents.

Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record. Its drivers are YAML files called instrument profiles. This essay makes the case for that format against the two incumbents, the hand-written class and the IVI driver, and is specific about where the case stops.

The class-per-instrument habit

Most Python benches start the same way. PyVISA solves transport: in its own words, it "enables you to control all kinds of measurement devices independently of the interface," whether that is GPIB, serial, USB, or Ethernet. It does not know what your instrument's commands are. That vocabulary lives in the programming manual, so somebody opens the manual and writes a class:

keysight_34461a.py
import pyvisa
 
class Keysight34461A:
    def __init__(self, address: str):
        self.inst = pyvisa.ResourceManager().open_resource(address)
        self.inst.read_termination = "\n"
 
    def measure_voltage(self) -> float:
        return float(self.inst.query("MEAS:VOLT:DC?"))
 
    def set_range(self, rng: float = 10) -> None:
        if rng not in (0.1, 1, 10, 100, 1000):
            raise ValueError(f"unsupported range {rng}")
        self.inst.write(f"VOLT:DC:RANG {rng}")

Nothing is wrong with this class. At two methods it is the obvious thing to write. The trouble arrives with scale, and it shows up in three ways.

Drift. The manual, the firmware, and the class are three copies of the same facts, and only the class is maintained by your team. A firmware revision adds a range or changes a default; the tuple in set_range stays where it was. Nobody notices until a run fails on a bench, and then somebody has to work out whether the instrument, the class, or the test is wrong.

Duplication. Because the class is code, it is also somebody's style. Teams that share a multimeter model tend to write separate classes, each covering the subset of commands that team needed, each with its own name for the same measurement. Sharing one means adopting its author's conventions, error handling, and threading assumptions along with its command strings.

Review burden. In a class, the facts a reviewer needs to check (the command string, the allowed values, the unit) are interleaved with control flow. To confirm that set_range matches the manual, a reviewer reads the method, finds the f-string, finds the guard, and hopes there is not a second guard somewhere else. Now multiply by a real command set. The profile for Keysight's InfiniiVision 3000 X-Series oscilloscopes in the galois-edge repository defines 92 commands and four sequences. As a class, that is 92 methods to read and keep honest.

What IVI got right, and what it costs

The industry's answer to duplication was the vendor driver. The IVI Foundation defines "an open driver architecture, a set of instrument classes, and shared software components," so a vendor writes a driver once and a test program can interchange instruments that implement the same class. That interchangeability is real, and teams whose test programs depend on it have good reasons to keep it.

The costs sit elsewhere. The classic driver types, IVI-C, IVI-COM, and IVI.NET, all require the IVI Shared Components, which the foundation publishes as Windows installers. The foundation is modernizing: in 2025 it simplified its driver specifications and added .NET 6+ support, a set it calls IVI Generation 2026, and in February 2026 it completed an IVI Python specification. Even so, an IVI driver is code that someone else owns. If the command you need is missing, you wait for a vendor release or bypass the driver with raw I/O, and at that point you are back to writing command strings by hand.

What a profile looks like

A Galois instrument profile holds the same facts as the class above with the control flow removed. This is adapted from the example in the instrument profiles reference, with the reading command named as in the shipped 34461A profile:

keysight_34461a.yaml
instrument:
  manufacturer: "Keysight"
  model: "34461A"
  class: dmm
  description: "Truevolt Series Digital Multimeter"
 
identity:
  query: "*IDN?"
  patterns:
    - "(?:Keysight|Agilent).*34[4-7][0-9][0-9]A"
 
interfaces:
  - type: gpib
    default_address: 22
  - type: usb
  - type: ethernet
    port: 5025
 
settings:
  timeout_ms: 5000
  terminator: "\n"
  opc_query: true
 
commands:
  measure_voltage_dc:
    scpi: "MEAS:VOLT:DC?"
    type: query
    returns:
      type: float
      unit: V
    description: "Measure DC voltage"
 
  set_range:
    scpi: "VOLT:DC:RANG {range}"
    type: write
    params:
      range:
        type: float
        unit: V
        options: [0.1, 1, 10, 100, 1000]
        default: 10

Read it top to bottom and every line is a fact you could check against the manual: the regex that matches the instrument's *IDN? response, the transports and default GPIB address, the timeout and terminator, and then the commands. set_range carries its SCPI template, its parameter's type and unit, the values the manual lists, and a default. The galois-edge daemon reads the file, matches it to a connected instrument by its identity string, and does the I/O.

The diff story follows directly. When a firmware revision adds a range, the change is one line in options, and the pull request shows exactly that line. No reviewer has to reason about whether a guard elsewhere still agrees.

Profiles are trees, and agents walk them

A profile is a tree, and the YAML nesting is the tree. The instrument sits at the root. Under it, commands holds each command; each command holds its params; each parameter holds its type, unit, limits, and options. Read the example above as a path and you get commands → set_range → params → range → options, which is the route an engineer or an agent takes from "this multimeter" to "the values this setting accepts."

To be precise about the shape: every command sits as a sibling under commands, and grouping comes from naming and from the SCPI path itself, which is hierarchical by design: VOLT:DC:RANG reads as subsystem, function, setting. Commands that belong together share a name prefix and a SCPI root.

The tree matters most to agents. The galois-edge daemon embeds an MCP server, and an agent does not need every command of every instrument in its context to use one of them. It can descend the tree:

  1. list_instruments returns what is connected, one root per instrument.
  2. get_capabilities for one instrument returns that instrument's commands, sequences, and settings.
  3. The agent calls one command, either through the generic execute_command or through the typed tool the daemon generates for it, named <profile_key>__<command_name>, such as keysight_34461a__measure_voltage_dc.

Each typed tool's input schema is built from the parameter level of the tree. An enum's options become a JSON Schema enum, min and max become minimum and maximum, and the unit is appended to the description. Out-of-range values are rejected before any SCPI reaches the wire. When an instrument is plugged in, the daemon matches its profile and tells connected agents that the tool list changed; when it is unplugged, its tools are removed. MCP for lab instruments follows that loop from first connection to a reviewed test sequence.

Compare an agent working from a driver class. To learn what set_range accepts, it has to read the method body and find the guard, and the guard may be incomplete. From a profile it reads one node.

The same tree serves Python. With the pyvisa-galois backend, profile commands appear as keyword-only methods on an ordinary PyVISA resource, so smu.set_voltage(voltage=1.5) works next to smu.query("*IDN?"), and a misspelled parameter name fails at call time instead of sending the wrong command. SCPI instrument automation with Python walks through that workflow end to end. Évariste, the agent in the Galois platform, works from the same tree without a script; generating a driver with Évariste builds the 34461A profile above from its manual and then uses it on a bench.

Generated from the manual

A programming manual is already a declarative document. Its command reference lists each command's syntax, parameters, ranges, units, and query form, usually in tables. Writing a driver class means translating those tables into control flow by hand. Writing a profile means transcribing them into another table, and a transcription can be checked against its source line by line.

That is how Galois adds instruments it does not already cover: you upload the instrument's programming manual to Évariste and it generates a profile (PDF-to-driver). The output is the same YAML a person would write, so reviewing it means reading a diff, not auditing generated code. Check the SCPI templates against the manual's command reference, check ranges and units, and flag anything that can damage hardware (more on that below). A wrong range is one line in one place, and so is the fix.

The format is also a target for other converters. For CAN devices, the galois-edge repository includes a dbc2galois script that turns a vendor DBC file into a profile, so an ECU's message and signal layout becomes typed commands the same way an oscilloscope's SCPI command set does. Motor drive test automation shows where it fits for drives commanded over vendor CAN.

Generated profiles are not exempt from judgment. A model reading a long manual can misread a table, and no file format makes the manual itself correct. What the format does is put every claim the generator made where a reviewer can see it, in the same shape every time.

Validation and safety live in the data

Because a profile is data, the daemon can check it before anything runs. Every profile is validated when it loads, and failures are logged with the offending command or sequence name. The checks include:

  • an identity with neither pattern nor patterns, or a regex that does not compile;
  • a command with no way to send it, meaning none of scpi, getter, setter, sdk_call, or can;
  • a parameter with an unknown type, or an enum with no options;
  • an unknown return type, or a sequence with no steps.

A hand-written class gets equivalent checks only if someone writes tests for them, and they run only when someone runs the tests.

Safety metadata lives in the same file. Mark a command is_dangerous: true and its MCP tool carries a destructive hint, so clients that ask a person before a destructive call can do so; on the relay path, the daemon refuses dangerous calls from any token that lacks danger_allow. The other safeguards an agent-driven bench needs are covered in failure modes and layered controls. Mark a command requires_sweep: true and the daemon refuses to run it as a single command at all. It has to go through start_sweep: a ramp that runs on the daemon at a requested rate, polls a status query until the instrument reports it is done, and has an abort command if someone stops it. That is the right path for magnets and temperature controllers, and because the ramp lives on the daemon, it does not depend on the agent's session staying connected. Lab automation for university research labs puts a Lake Shore 336 setpoint ramp on that path.

Here is an illustrative excerpt, not taken from any real instrument's command set:

ramp_excerpt.yaml
commands:
  set_field:
    scpi: "FIELD:TARG {value}"
    type: write
    is_dangerous: true
    requires_sweep: true
    params:
      value:
        type: float
        unit: T
    sweep:
      rate_param: sweep_rate
      command: "FIELD:TARG {value};RATE {sweep_rate}"
      check_command: "STATE?"
      check_idle_match: "HOLDING"
      stop_command: "PAUSE"
      poll_interval_ms: 1000

A reviewer can see in under twenty lines that the command is hazardous, that it cannot be fired directly, how the ramp is checked, and how it stops. In a class, the same guarantees tend to be spread across a method, a helper, and a polling thread.

How to generate an instrument driver from its manual in Galois with Évariste

The same driver, and a bench check that exercises it, can be built in the Galois app without writing the class or the YAML. Open Évariste from the sidebar (Ctrl+Shift+E) beside your project. Galois runs in the cloud or on-prem (deployment options), and Évariste reaches the bench through the same galois-edge daemon. The 34461A already ships in the library, so on a real bench the daemon would already match it. This walkthrough generates it anyway, because the class and profile above give a known answer to check the draft against.

Describe the driver. Upload the 34461A manual as a PDF and state what the class encodes:

Write a profile for the Keysight 34461A from this manual. Match the *IDN? reply on Keysight or Agilent and models 34400A through 34799A. Add a DC voltage measurement that returns a float in volts, and a DC voltage range setting with the options 0.1, 1, 10, 100 and 1000 V and a default of 10 V.

Évariste drafts the profile in the format above:

Évariste draft: keysight_34461a.yaml (excerpt)
identity:
  query: "*IDN?"
  patterns:
    - "(?:Keysight|Agilent).*34[4-7][0-9][0-9]A"
 
commands:
  set_range:
    scpi: "VOLT:DC:RANG {range}"
    type: write
    params:
      range:
        type: float
        unit: V
        options: [0.1, 1, 10, 100, 1000]
        default: 10

Review and approve. Check each SCPI template against the command reference, each options list against the manual's range table, the units, and the identity regex against your meter's actual *IDN? reply. Neither command here is hazardous; for a supply or a magnet controller, the prompt also names the commands to mark is_dangerous or requires_sweep, the review confirms the flags, and Évariste asks you to confirm before it sends a dangerous command. Ask for fixes in conversation, then approve. How to review an AI-generated test plan applies the same discipline to sequences.

Deploy and read. Ask Évariste to deploy the profile to the bench's galois-edge and bind it to the multimeter. "List connected instruments" now shows the 34461A with measure_voltage_dc and set_range, and "Measure voltage on the DMM" returns one reading.

Run a check with limits. A reading is not a test, so ask for one: "Create a DMM check sequence: with the inputs shorted, set the 10 V range, measure DC voltage and expect -0.001 to 0.001 V." The limit is yours. The sequence lands as a draft; confirm it calls set_range and measure_voltage_dc with the values you stated, then approve it. Later edits become new versions with diffs. Short the inputs and start the run; it executes through galois-edge while Monitor shows the channel live.

Results and report. Each step is recorded with its measured value, limits, pass or fail, the raw command and response, the instrument, operator, DUT serial and timestamps. Ask Évariste which steps failed or passed close to a limit. After a firmware update, rerun the sequence and ask it to compare the two runs; if the release notes add a range, upload the revised manual and check the new options list before you approve and redeploy. "Generate a test report from the last run" produces a PDF or HTML report you can edit in the report editor and share, for example to Slack.

StepCode pathGalois with Évariste
Describe commandsTranscribe the manual into Keysight34461A or YAMLUpload the PDF; Évariste drafts the same YAML
Encode rangesThe set_range guard, or optionsoptions from your prompt, checked against the manual
ValidateTests you write; schema check at loadSchema check at load, plus your review and approval
InstallImport the class, or copy the YAML and restartÉvariste deploys and binds it
Take a readingmeasure_voltage() from a script"Measure voltage on the DMM"
Check with limitsA script with assertionsApproved draft sequence, run through galois-edge
ResultsYour logging and plotsPer-step record, near-limit passes, run comparisons
ReportYour report scriptGenerated PDF or HTML report
Firmware adds a rangeEdit the tuple and its tests, or optionsUpload the revised manual; review options

With Évariste, you no longer write the class or type the YAML, and you do not maintain the PyVISA session setup, a script that calls the driver, logging for readings, or a report script. Stating the models and ranges from the manual, setting the limits, reviewing the draft, approving it, shorting the inputs and keeping the bench safe stay your job.

Hand-written class vs IVI driver vs declarative profile

ApproachDiffableGenerated from manualValidatedAgent-navigableCross-platformEffort to add a command
Hand-written class (PyVISA)Yes, but facts mix with logicNo, written by handOnly by tests you writeAgent reads source codeYes, with a VISA backendA method, plus its tests
IVI driverVendor-owned; you track versionsNo, written by the vendorVendor's responsibilityThrough a wrapper you writeClassic types: Windows. Generation 2026: .NET 6+ and Python specsA vendor release, or raw I/O
Declarative profile (Galois)Yes, one fact per lineYes, then reviewedSchema-checked at loadYes, tree to typed MCP toolsDaemon on Linux, Windows, Raspberry PiA few lines of YAML

Source: IVI Foundation, Shared Components

When code is still the right answer

A profile describes an instrument. It is not a program, and some driver work is a program.

Branching, stateful logic. As documented, a profile sequence is an ordered list of steps: call a named command or send raw SCPI, fill in arguments, capture a result. Logic that branches or repeats belongs in Python: a calibration routine that adjusts based on the last reading, a bisection search for a threshold, or a retry policy keyed to a specific instrument error. That code should call the profile's commands rather than re-encode them, so the vocabulary stays in data and the logic stays in code.

Vendor SDKs. Some instruments do not speak SCPI and ship a Python SDK instead. A profile can still describe them: a top-level sdk block names the package and class and says how to connect, and a command's sdk_call maps profile parameters onto SDK method arguments. The SDK itself is still code, though, and when its API does not map cleanly onto named methods with keyword arguments, a thin data mapping is the wrong tool. galois-edge handles those cases with Python wrapper modules that declare their own typed tools, for example for the FNIRSI DPS-150 power supply and the Digilent Analog Discovery 3. We use code where data does not fit, too.

Bespoke binary protocols. Profiles handle IEEE binary blocks, ASCII arrays, regex and split parsers, and CAN signal layouts. A proprietary framing with checksums and stateful decoding is better written, and tested, as code.

An interchangeability program that already works. If your test programs are built on IVI class drivers so that one vendor's instrument can replace another's without code changes, that investment is doing its job. A declarative layer can sit beside it; there is no reason to tear it out.

The rule of thumb: describe the instrument in data, and put the test on top of it, in code when it branches. A profile is the vocabulary; a test is what you say with it.

Profiles in Galois today

Galois ships 573 instrument profiles across 135 manufacturers in its instrument library. The galois-edge daemon is open source under Apache-2.0. To add your own, follow the step-by-step SCPI profile guide: drop a YAML file in the profile directory and restart the daemon, which logs every profile it loaded and every validation error. Or upload the programming manual to Évariste and review the profile it generates. The daemon runs on Linux, Windows, and Raspberry Pi.

For what agents do once instruments are described this way, from generating test sequences to running them and writing up the results, read AI test automation for hardware benches. For the rest of the platform, see the product overview.

Frequently asked questions

What is a declarative instrument driver?
A driver written as data rather than code: a file that lists an instrument's commands, their command strings, typed parameters with units and limits, and return types. A runtime reads the file and handles the I/O, so the driver can be diffed, validated, and generated from the programming manual.
Why is a declarative driver easier to generate than a driver class?
A programming manual's command reference is already a set of tables: syntax, parameters, ranges, and units. Generating a profile turns one table into another, so each generated line can be checked against one line of the manual. Generating a class produces control flow, which a reviewer has to audit as code.
When should a driver still be code?
When the work is a program rather than a vocabulary: branching calibration routines, adaptive searches, retry logic, vendor SDKs whose APIs do not map onto named methods, and proprietary binary protocols. Keep the instrument description in data and write that logic in Python that calls it.
How do I review a driver generated from a programming manual?
Check it like a transcription: each SCPI template against the manual's command reference, each range and options list against its table, the units, the identity regex against your instrument's real *IDN? reply, and that every command that energizes an output is flagged is_dangerous. In Galois, Évariste, the agent in the Galois platform, drafts the profile from the uploaded manual; after you approve it, Évariste deploys it to galois-edge and binds it to the instrument.

Related

Bring Galois to your bench.

The daemon is Apache-2.0, free forever. Enterprise runs in your cloud or on-prem.