---
title: Can an LLM Safely Drive Lab Instruments?
description: "LLM instrument control safety: how an agent can hurt a bench (wrong range, output enable, runaway sweeps, stale state) and the layers that contain it."
url: https://galoislabs.ai/blog/llm-instrument-safety
author: Alex Hernandez
author_url: https://galoislabs.ai/blog/authors/alex-hernandez
published: "2026-06-27"
topic: Agents
publisher: Galois Labs
---

# Can an LLM safely drive lab instruments? Failure modes and layered controls

![A small safety interlock box in isometric with a guarded emergency-stop button and a key switch on top.](https://galoislabs.ai/blog/figures/agents-2.light.webp)

*FIG. 1 — INTERLOCK: STOP BUTTON AND KEY*

Yes, if the limits that matter sit outside the model. A language model can drive a power supply or a source-measure unit safely when every command passes through controls it cannot rewrite: typed ranges checked before anything is sent, approval before a sequence runs, ramps executed by the daemon, permissions scoped per call, and interlocks wired in hardware.

A prompt that says "be careful" is not one of those controls, and neither is a tool annotation. This essay covers four ways an agent can hurt a bench, the layer that contains each, and the risk that remains.

Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record. The examples use the galois-edge daemon because we built it. The failure modes apply to any agent with a path to an instrument.

## What can go wrong when an LLM sends commands to an instrument?

A failed software test leaves a red badge; a command sent to a source puts volts on a pin. None of the failures below need exotic model behavior. They are a tired engineer's mistakes, made by a process that types faster and never feels a board get warm.

### Wrong range

A unit error such as `VOLT 3300` for 3.3 V exceeds the instrument's range, and [SCPI-99](https://www.ivifoundation.org/downloads/SCPI/scpi-99.pdf) (section 7.2) does not define what the instrument then sets. A decimal slip is worse: `VOLT 33` where 3.3 was meant is a legal command on a 60 V supply, and the supply will do it. The instrument checks its own limits, not your board's.

Current limits fail by omission. An agent that sets the voltage but never the current limit leaves the supply at whatever limit it last had, and a shorted rail draws up to that.

### Output enable

Enabling an output is the moment a mistake reaches the board. It applies whatever setpoint the instrument holds right now, which may come from the previous session, the front panel, or a different test. The safe order is limits first, setpoint second, enable last. An agent issuing calls one at a time can get that order wrong, or retry an enable after an error that should have ended the session.

SCPI helps at the start of a session: the [SCPI-99 command reference](https://www.ivifoundation.org/downloads/SCPI/scpi-99.pdf) (section 15.12) says that at `*RST` the output state "shall be set to OFF." But a reset also wipes configuration someone set on purpose, and nothing forces an agent to send one.

### Runaway sweeps

Ask an agent to ramp a supply from 0 to 12 V in 0.5 V steps and it may write the ramp as a loop of tool calls. The rate is then the model's turn latency, and the ramp's end depends on the session surviving. If the API call times out or the network drops at 7.5 V, the output sits at 7.5 V with nobody watching. If the model loses count, it can overshoot. Magnet supplies and temperature controllers have rated ramp rates; a ramp paced by inference latency respects none of them.

### Stale state

An agent's picture of the bench is its context window: tool results it read some turns ago. Meanwhile someone pressed a front-panel key, another script reconfigured the meter, a protection circuit tripped, or one unit was swapped for another of the same model. "The output is off" in the model's context is a memory, not a measurement, and a compacted context can lose the turn where the state changed.

| Failure mode  | Typical trigger                                  | Layer that contains it                                     |
| ------------- | ------------------------------------------------ | ---------------------------------------------------------- |
| Wrong range   | Decimal slip, missing current limit              | Profile ranges, hardware limits                            |
| Output enable | Stale setpoint, wrong order, retry after a fault | Dangerous-command flags, approval, cleanup on a clean stop |
| Runaway sweep | Ramp written as a loop of tool calls             | Daemon-side sweeps                                         |
| Stale state   | Context out of date with the bench               | Read-back, locked sequences, fail-closed tool names        |

## Why MCP tool annotations are hints, not controls

The Model Context Protocol lets a server annotate each tool with `readOnlyHint`, `destructiveHint`, `idempotentHint`, and `openWorldHint`, and it is explicit about what those are. The [MCP schema](https://github.com/modelcontextprotocol/modelcontextprotocol/blob/main/schema/2025-11-25/schema.ts) says "all properties in ToolAnnotations are **hints**. They are not guaranteed to provide a faithful description of tool behavior," and that "clients should never make tool use decisions based on ToolAnnotations received from untrusted servers." The [tools page of the specification](https://modelcontextprotocol.io/specification/2025-11-25/server/tools) adds that clients must treat annotations as untrusted unless they come from trusted servers, and that there should always be a human in the loop with the ability to deny tool invocations.

galois-edge maps a profile's `is_dangerous: true` to `destructiveHint: true`, and Claude Desktop turns that into a confirmation prompt ([client configuration](https://docs.galoislabs.ai/agents/clients/)). That prompt is worth having, but it is not the control:

- **The client decides what the hint means.** A different client, or a script calling the server directly, can ignore it.
- **Prompts decay.** A dialog that offers "Always allow" is one click from not being a dialog.
- **The hint describes the tool, not the bench.** A prompt to enable an output looks the same whether the setpoint behind it is 3.3 V or 33 V, because the setpoint was written by an earlier, harmless-looking call.

OWASP's guidance on [excessive agency in LLM applications](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/) states the general rule as complete mediation: "Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not." On a bench, downstream means the daemon and the hardware.

## What controls keep an LLM inside safe limits?

Each layer catches some of what the one above lets through. None is sufficient alone.

| Layer                          | Enforced by                        | Contains                                              | Leaves                                        |
| ------------------------------ | ---------------------------------- | ----------------------------------------------------- | --------------------------------------------- |
| Profile ranges and flags       | Daemon, before SCPI is sent        | Out-of-range values, calls to hidden setters          | In-range values that are wrong for this unit  |
| Approval gates                 | Platform, before a sequence runs   | Wrong limits, wrong order, missing cleanup            | Interactive tool calls, habitual approvals    |
| Daemon-side sweeps             | Daemon and the instrument's ramp   | Loop-paced ramps, ramps stranded by a dropped session | A bad target or rate chosen by the caller     |
| Scoped permissions             | Daemon, per call on the relay path | Tools and dangerous commands outside the grant        | Anything a trusted network can reach directly |
| Physical limits and interlocks | Hardware                           | Whatever every layer above missed                     | Wiring mistakes inside the limits             |

### Profile ranges and dangerous-command flags

In galois-edge, typed instrument commands come from a YAML profile that declares each parameter's type, unit, minimum, and maximum ([profile schema](https://docs.galoislabs.ai/reference/instrument-profiles/)). The daemon publishes each command as a typed MCP tool whose input schema carries those limits, and it rejects out-of-range values before any SCPI reaches the instrument ([MCP server reference](https://docs.galoislabs.ai/agents/mcp-server/)). An illustrative excerpt for a bench supply powering a 3.3 V board:

```yaml title="psu_excerpt.yaml"
settings:
  cleanup_commands: ["OUTP OFF"]   # sent when the daemon stops cleanly and disconnects

commands:
  source_voltage:
    scpi: "VOLT {value}"
    type: write
    params:
      value:
        type: float
        unit: V
        min: 0
        max: 3.6                   # what the board survives, not what the supply can do

  set_current_limit:
    scpi: "CURR {current}"
    type: write
    params:
      current:
        type: float
        unit: A
        min: 0
        max: 0.5

  output_on:
    scpi: "OUTP ON"
    type: write
    is_dangerous: true             # published as destructiveHint; refused without danger_allow on the relay

  output_off:
    scpi: "OUTP OFF"
    type: write                    # no flag: turning the output off never waits on approval

  set_ovp_level:
    scpi: "VOLT:PROT {level}"
    type: write
    enabled: false                 # hidden from clients: a person sets protection
    params:
      level:
        type: float
        unit: V
```

The most important number in that file is `max: 3.6`. A model sees it before calling anything, because the tool's input schema publishes it. And when a call is refused, the MCP specification expects input-validation errors to go back to the model so it can "self-correct and retry with adjusted parameters." A determined model will find the ceiling and use it, so set the maximum to what the device under test survives, not to the instrument's rating. None of the 573 profiles in the [instrument library](https://galoislabs.ai/instruments) can know your board; narrow the ranges before an agent uses one.

Two more lines address the other failure modes. `cleanup_commands` turns the output off when the daemon stops cleanly and disconnects from the supply. A crash, a `kill -9`, or a power cut sends nothing, which is one reason a hardware limit sits under every source. `enabled: false` keeps the protection setter out of every client's tool list, so the trip level stays where a person put it. [Declarative instrument drivers](https://galoislabs.ai/blog/declarative-instrument-drivers) argues why this metadata belongs in data rather than in a driver class.

Stale state gets two structural answers. A profile sequence runs every step under the daemon's per-instrument lock ([daemon API](https://docs.galoislabs.ai/reference/daemon-api/)), so no other client's commands land between its steps. And when a second instance of the same profile appears, the daemon re-emits the first instance's tools under a suffixed name, so a call to the old name fails instead of reaching a unit the agent did not pick. A one-for-one swap keeps the old name, so query the instrument before acting on what it was.

### Approval gates

Galois treats a generated sequence as a draft. The platform refuses to run a draft until an engineer approves it, and a sequence edited after approval needs a new approval. The sequence is text, so the reviewer reads a diff: every limit, the order of limit, setpoint, and enable, and whether the last step turns the output off. [How to review an AI-generated test plan](https://galoislabs.ai/blog/review-ai-generated-test-plan) covers what to check.

Approval gates a sequence run from the platform. In an interactive session, where an agent calls tools one at a time, the per-call gate is the client's confirmation prompt, so the layers below carry more weight. [Keysight's MCP Server for Instrument Control](https://helpfiles.keysight.com/kmsic/English/keysight_mcp_for_instrument_control/Content/overview.html) draws the line elsewhere: every command sequence needs explicit engineer approval before it reaches hardware. Finer approval catches more and tires reviewers faster.

### Daemon-side sweeps

A profile can mark a command `requires_sweep: true`. The daemon then refuses it as a one-shot call; it has to go through `start_sweep`, which sends the target and rate to the profile's ramp command, returns a handle, and polls a status query until the instrument reports it is idle. `stop_sweep` either holds at the current value or runs the profile's abort command.

That moves the ramp's pace from the model's loop to the instrument's own ramp. The sweep also continues if the agent's session drops, so it finishes at the rate it was given; measurement streams, by contrast, end on disconnect. A sweep that outlives its agent still needs an owner, and anyone holding the sweep ID can check or stop it. The caller still chooses the target and rate, so keep `start_sweep` out of any permission set not meant to ramp.

### Scoped permissions

The galois-edge MCP server has two ways in, and they trust the caller differently.

On the direct path, a client dials the daemon's MCP port and there is no per-call auth: the supervisor exposes the port on the tailnet address and on `0.0.0.0`, so network reachability is the boundary. Every tool is reachable there, including `send_scpi`, which passes a raw string to the instrument and bypasses profile validation. Treat direct access like a raw SCPI console and keep it on a network you would trust with one.

On the relay path, each call carries a signed per-call token, checked against tool permissions. The token is a JWT minted for one user and one bench that expires after five minutes. Its `tools_allow` list names permitted tools, optionally narrowed to a single command, and with `danger_allow` false the daemon refuses any command flagged dangerous regardless of the list. A first session can be read-only:

```json title="relay token claims for a read-only session (excerpt)"
{
  "aud": "edge:<edge_id>",
  "sub": "user:<user_id>",
  "tools_allow": [
    "list_instruments",
    "get_capabilities",
    "execute_command:example_psu__measure_voltage",
    "get_sweep_status"
  ],
  "danger_allow": false
}
```

The list leaves out `send_scpi` and `start_sweep`. The [security page](https://galoislabs.ai/security) covers the wider model, and [MCP for lab instruments](https://galoislabs.ai/blog/mcp-lab-instruments) covers the protocol side.

### Physical limits and interlocks

Many instruments carry their own protection, and SCPI standardizes the commands. In SCPI-99, `SOURce:VOLTage:PROTection:LEVel` "sets the output level at which the output protection circuit will trip," and `OUTPut:PROTection:TRIPped?` reports whether it has. Use them, but the trip level is itself a bus command. Until the setter is hidden and raw SCPI is out of reach, over-voltage protection is one more setting an agent can change.

The bottom layer is whatever no command can change. A supply whose maximum output is below what the board survives. A fuse or current-limit resistor sized to the rail. A fixture lid switch wired to the instrument's interlock input, where it has one. An emergency stop that removes power from the device under test. These keep working when the model, the daemon, and the network are all wrong at once.

After the fact, the record checks everything above it. MCP calls run through the same command path as the daemon's gRPC API, so one audit trail covers agents and scripts, and each sequence step keeps the SCPI sent and the raw response received ([hardware test traceability](https://galoislabs.ai/blog/hardware-test-traceability)).

## What risk remains with every layer in place?

Layers reduce the space of possible mistakes. They do not empty it.

- **In range but wrong.** Ranges belong to a profile, not to a net. A value legal for one rail can be wrong for the rail the lead is actually on.
- **The agent cannot see the bench.** A swapped lead, a probe on the wrong pin, or the wrong unit in the fixture passes every software check.
- **Profiles can be wrong.** Load-time validation checks a profile's structure, not whether its command strings and ranges match the manual. Generated and hand-written profiles both need review.
- **Approval decays.** Reviewers approve what looks familiar, and the hundredth sequence gets less attention than the first.
- **The direct path trusts the network.** Anyone who can reach the port can call any tool.
- **Text can carry instructions.** An agent that reads manuals, datasheets, or instrument responses can meet [text written to steer it](https://genai.owasp.org/llmrisk/llm01-prompt-injection/). Permissions bound what that text can make the agent do.

Galois publishes no reliability figure for model behavior, and the argument here needs none: the layers are designed to hold when the model is wrong. So the honest answer is the one a lab gives any new operator: yes, on a bench set up so the worst plausible mistake is survivable, with a person who knows which mistakes those are.

## How to set up a bench for an AI agent

1. **Start read-only**: discovery and measurement tools, nothing that sources.
2. **Narrow every source profile** to the device under test, voltage and current both.
3. **Flag output enables** with `is_dangerous: true`, leave the off command unflagged, and route ramps through `requires_sweep`.
4. **Hide protection setters** with `enabled: false`; set trip levels at the front panel.
5. **Turn outputs off on a clean stop** with `cleanup_commands`; step 7 covers crashes and power cuts.
6. **Keep `send_scpi` out** of agent permissions, and the direct MCP port on a trusted network.
7. **Put a hardware limit under every source**: a lower-rated supply, a fuse, or an interlock.
8. **Watch the first sessions**, then read the per-step record, not the agent's summary.

Évariste, the agent in the Galois platform, reaches the bench through galois-edge, and this is the setup to give it. Upload the supply's programming manual and Évariste generates the profile. Before Évariste deploys it to an edge and binds it to the supply, check its command strings against the manual and apply steps 2 to 5: ranges narrowed to what the board survives (3.6 V and 0.5 A in the excerpt above), output enable flagged, the protection setter hidden, and the output turned off on a clean stop. Then start from the prompt "Create a power rail validation sequence" with the 3.3 V rail passing from 3.2 to 3.4 V, the window from the board's datasheet. The sequence arrives as a draft you approve before it runs, and each edit becomes a new version with a diff. A dangerous command Évariste sends ad hoc waits for your confirmation, a prompt that sits above the hardware limit in step 7, not in place of it. Every run step records the measured value, limits, raw command, and response, which is the record step 8 says to read. You no longer write the profile, the sequence, or the logging by hand; the limits, the review, the wiring, and the interlocks stay yours.

[AI test automation for hardware benches](https://galoislabs.ai/blog/ai-test-automation-hardware) describes the full loop these controls sit inside, and [SCPI instrument automation with Python](https://galoislabs.ai/blog/scpi-automation-python) covers the same instruments without an agent. To try the daemon on a bench, start with the [quickstart](https://docs.galoislabs.ai/getting-started/quickstart/) and the [agents overview](https://docs.galoislabs.ai/agents/).

## Frequently asked questions

### Can an LLM damage lab equipment?

Yes, through ordinary bench mistakes made quickly: a decimal slip that stays inside the instrument's range, an output enabled with a setpoint left from an earlier session, a ramp paced by a chat loop, or a command based on state that has since changed. Protection comes from limits outside the model: profile ranges, approval gates, daemon-run ramps, scoped permissions, and hardware interlocks.

### Are MCP tool annotations like destructiveHint enforced?

No. The MCP specification defines tool annotations as hints that are not guaranteed to describe a tool faithfully, and says clients must treat them as untrusted unless they come from a trusted server. A client can use destructiveHint to show a confirmation prompt, but enforcement has to happen in the server and in the hardware behind it.

### How do I let an AI agent control a power supply safely?

Narrow the profile's voltage and current ranges to what the board survives, flag output enable as dangerous, hide protection setters from clients, turn the output off at daemon shutdown, keep raw SCPI tools out of the agent's permissions, start with read-only access, and add hardware limits that no command can change: a supply rated below the board's maximum, a fuse, or an interlock.

### Does human approval make LLM instrument control safe?

It helps, and it is one layer. Approval catches a wrong limit or a bad order of operations in a reviewable sequence. It does not catch a swapped lead, a value that is in range but wrong for this unit, or a reviewer who approves out of habit. Keep hardware limits underneath it.
