How to review an AI-generated test plan: a checklist worked on an LDO rail

By Alex Hernandez · · 14 min read

View Markdown
A mechanical drum sequencer in side elevation, pins driving four lever followers, one pin marked by an empty callout circle.
FIG. 1 — SEQUENCE DRUM, ONE STEP FLAGGED

To review an AI-generated test plan, check six things in order: each requirement maps to steps at its worst-case conditions; each limit cites a requirement or datasheet line; instrument ranges and waits fit the measurement; sequencing ends in a safe state on every path; a dead board fails the plan; and approval binds to the revision you read.

A machine-written plan reads the same whether a number came from a datasheet or from nowhere. That is the reviewer's problem. This guide works the checklist through one power rail. The method applies to any plan; the closing sections show the same work done in Galois and how it handles approval.

The worked example: a 3.3 V LDO rail

Take a board whose 3.3 V rail comes from a TLV75533P, the fixed 3.3 V version of TI's TLV755P 500 mA LDO, in the SOT-23-5 (DBV) package, fed from a 5 V input. EN is tied to IN, as the datasheet suggests when shutdown is not needed. Every device figure below comes from the TLV755P datasheet (SBVS320D, revised September 2024).

The board's designer wrote three requirements:

  • R1. 3V3 stays within 3.3 V ±3% from 0 to 300 mA, for inputs from 4.5 V to 5.5 V.
  • R2. 3V3 is in regulation within 2 ms of the input reaching 4.5 V.
  • R3. A short from 3V3 to ground damages nothing, and the rail recovers when it is removed.

Here is an illustrative plan of the kind an agent drafts from the schematic and the datasheet:

StepConditionsDrafted limit
Power-upWait 500 msNone
Vout5.0 V in, 1 mA3.267–3.333 V
Vout5.0 V in, 500 mA3.267–3.333 V
Load regulation1 to 500 mAChange ≤ 30 mV
Line regulation3.8 to 5.5 V inChange ≤ 2 mV
Dropout500 mA≤ 215 mV
Current limitOutput shorted≥ 500 mA
Ripple500 mA, scope≤ 20 mV peak to peak
PSRR100 kHz≥ 46 dB
Shutdown currentEN low≤ 1 µA
Input voltageSupply readback4.75–5.25 V

It looks thorough. Most of it does not test the requirements.

What to check in an AI-generated test plan

Ask the plan's author, human or agent, for evidence rather than reassurance:

CheckQuestionEvidence
Requirement coverageIs every requirement tested at its worst case, and does every step trace to one?Requirement-to-step matrix
Limit provenanceWhere did each number come from?Requirement ID, or datasheet row, column and conditions
Ranges and settlingCan each instrument make this measurement, and how long does the physics take?Range and mode per step; a source for every wait
Sequencing and safe statesWhat order is power applied and removed, and where does every exit leave the bench?Power order; end state after pass, fail, error, abort
Do-nothing scoreWhich checks pass on a board that never powers up?The plan scored against a dead board; target zero
ApprovalWhat exactly is approved, and by whom?Named reviewer, pinned revision, rule for edits

Coverage comes first because it changes everything after it; approval comes last because it records the other five.

Does the test plan cover every requirement?

Build the matrix before reading a single limit. For each requirement, list the conditions that make it hardest to meet, then find the steps that exercise them.

RequirementWorst-case conditionsDraft stepsFinding
R14.5 V and 5.5 V in, at 0 and 300 mAVout at 5.0 V inNo corner tested; 500 mA exceeds the requirement
R2Input and output on a scopeWait 500 msA fixed wait measures nothing
R3Short, bounded dwell, remove, re-measureCurrent limitA different quantity; no recovery check
NoneRegulation, dropout, PSRR, ripple, shutdown currentTrace to a requirement or cut

The draft tests the part; the requirements describe the board, and the board's own questions (corners, timing, recovery) are missing. Read the steps against the schematic too: shutdown current needs EN driven low, and on this board EN is tied to IN, so that step cannot run as written. Untraced steps cost bench time and their own limit review.

For building a plan from requirements, see power rail validation plans; for keeping the links once it exists, see hardware test traceability.

Where did each test limit come from?

Every limit has one of three origins: a requirement, a datasheet, or nothing. A datasheet limit counts only if it is a minimum or maximum, not a typical value, at the conditions the datasheet states.

Drafted limitOriginFinding
Vout ±1% at 1 mADatasheet: ±1%, junction −40 °C to 85 °C, DBVStated at VIN = 3.8 V; the board test should use R1's ±3%
Vout ±1% at 500 mAThe same row, outside its conditionsLoad regulation alone takes about 30 mV of a ±33 mV window
Load and line regulationDatasheet typical (0.060 V/A, 2 mV)No maximum exists to test against
Dropout ≤ 215 mVDatasheet maximum at 500 mA, junction to 85 °CValid; the plan must define the dropout point, which the table does not
Current limit ≥ 500 mARated output current: a guessDatasheet: 560 to 865 mA, at specific conditions
Ripple ≤ 20 mVNothingNo requirement and no matching datasheet figure
PSRR ≥ 46 dBDatasheet typicalTypical, and untraced

The electrical characteristics table holds IOUT at 1 mA and VIN at VOUT(NOM) + 0.5 V (or 2.0 V, whichever is greater) unless a row says otherwise. Reusing the ±1% row at 500 mA ignores load regulation, 0.060 V/A × 0.5 A or about 30 mV typical, so a healthy part can fail.

The datasheet specifies the current limit as 560 mA minimum, 720 mA typical and 865 mA maximum with the output pulled to 90% of nominal and the input at VOUT + VDO(MAX) + 0.25 V. The draft shorts the output instead. Below 0.4 × VOUT(NOM) the regulator folds its limit back, and a dead short draws 355 mA typical (section 6.3.3). A healthy part fails the drafted limit.

Two habits keep provenance alive. Put the source in the step itself, for example "Vout at 4.5 V in, 300 mA (R1, ±3%)", so it lands in the run record. And compare every window with what eats into it. At 500 mA, 50 mΩ of wiring and contact resistance drops 25 mV, three quarters of a ±33 mV window, unless the meter senses at the board.

Are the instrument ranges and settling times right?

Setpoints against both sets of limits. An instrument driver's ranges protect the instrument and know nothing about the board. The TLV755P's absolute maximum input is 6.0 V and its recommended maximum 5.5 V, so input setpoints stay at or below 5.5 V, with the supply's overvoltage protection below 6.0 V as a backstop.

Limits and modes that measure the board, not the bench. If the bench supply's current limit is 0.6 A, a current-limit test measures the supply. It must clear 865 mA plus ground current. The datasheet's condition holds the output at 2.97 V, which means an electronic load in constant-voltage mode; a constant-current load set above the limit drags the output into foldback. Fix the meter's range (10 V for a 3.3 V rail) rather than autoranging, and count its integration time in the step's budget, as SCPI automation with Python shows.

Waits that come from physics. Each wait should name the process it covers and the source of its number. Startup time is 550 µs typical, so the 500 ms wait is about 900 times longer and still says nothing about R2, which needs a scope capture with the 2 ms limit applied to the measured delay.

Thermal settling is the wait this draft misses. At R1's hardest corner, 5.5 V in and 300 mA out, the regulator dissipates (5.5 − 3.3) V × 0.3 A = 0.66 W. The datasheet gives the DBV package 100.8 °C/W on TI's evaluation board and 231.1 °C/W on the JEDEC test board: a junction rise of about 67 °C to 153 °C. At the high end, from a 25 °C bench, the junction would pass the 165 °C thermal shutdown point. The output drifts as the junction heats, so the plan must say when the reading is taken, and the corner goes to the designer before it goes to the bench.

Is the power sequencing safe on every exit path?

Power-up order should come with a reason for each step:

  1. Set the supply's voltage, current limit and overvoltage protection with its output off.
  2. Leave the electronic load's input off, then enable the supply and confirm the rail is up.
  3. Enable the load. The datasheet is explicit: "For constant-current loads, disable the output load until the output rises to the nominal voltage." Foldback during startup can hold the output down.

Power-down order is a datasheet question, not a convention. The reverse-current section (7.1.4) lists three conditions: the output biased while the input is not established, the output biased above the input, and a large output capacitance with the input collapsing at little or no load current. So nothing that can source, such as a second supply or an SMU, stays on the output when the input is off. With large output capacitance, drop the input while the load still draws current, then disable the load. The 120 Ω active discharge works only while there is enough input voltage, so it cannot be counted on once the input is gone.

The short-circuit test needs a stated dwell. At 5 V in, a dead short dissipates about 5 V × 355 mA ≈ 1.8 W, and the datasheet says a sustained short cycles the part between current limit and thermal shutdown. R3 asks whether the board survives that: apply the short for a stated time, remove it, re-measure Vout against R1.

Then write the end state for every way a run can stop (pass, fail, error, operator abort, lost connection) and verify it with a measurement: supply off, load off, rail below a stated voltage. Electronic load automation covers load modes and safe defaults.

What would a do-nothing run score?

Score the plan against a board that never powers up: an empty fixture, a supply never enabled, a meter on the wrong channel. Any check that passes proves nothing about the hardware. The target is zero, and it runs on paper:

dead_board.py
"""Score a test plan against a board that never powers up."""
from dataclasses import dataclass
 
 
@dataclass
class Check:
    name: str
    low: float | None   # None: no lower bound
    high: float | None  # None: no upper bound
    dead: float         # what this check reads when the rail never comes up
 
 
def passes(check: Check, value: float) -> bool:
    return (check.low is None or value >= check.low) and (
        check.high is None or value <= check.high
    )
 
 
PLAN = [
    Check("vout_1ma_v", 3.267, 3.333, dead=0.0),
    Check("vout_500ma_v", 3.267, 3.333, dead=0.0),
    Check("load_reg_delta_v", None, 0.030, dead=0.0),  # 0 V minus 0 V
    Check("line_reg_delta_v", None, 0.002, dead=0.0),  # 0 V minus 0 V
    Check("ripple_mv_pp", None, 20.0, dead=0.0),       # no output, no ripple
    Check("current_limit_ma", 500.0, None, dead=0.0),
    Check("shutdown_current_ua", None, 1.0, dead=0.0),  # nothing draws current
    Check("vin_v", 4.75, 5.25, dead=5.0),               # VOLT? returns the setpoint
]
 
passed = [c.name for c in PLAN if passes(c, c.dead)]
print(f"dead board passes {len(passed)} of {len(PLAN)}:")
for name in passed:
    print(f"  {name}")
dead board passes 5 of 8:
  load_reg_delta_v
  line_reg_delta_v
  ripple_mv_pp
  shutdown_current_ua
  vin_v

Five of eight checks pass with nothing working. They fall into three families.

Upper bound only. Ripple and shutdown current have no floor, so zero passes. Give limits two sides where physics allows; where it does not, gate the check on a step that proves the rail is alive.

Differences between readings. Regulation subtracts two readings, and 0 V minus 0 V is within any limit. Assert each reading before the difference.

Readback instead of measurement. Under SCPI-99, the query form of a command returns "the current setting associated with the command" (Volume 1, section 6.2.3), while MEASure? configures the instrument, takes a measurement and returns the result (Volume 2, chapter 3). VOLT? passes whenever the setpoint is right; MEAS:VOLT? reads what the output is doing.

If a known-bad unit exists, run the corrected plan against it on the bench too. A plan that cannot fail a broken board is not ready to pass a good one.

How to approve a test plan

Approval should bind to one artifact: the exact revision (commit hash or document version), the limits with their origins, the bench configuration (which physical instruments the plan's names resolve to), and the safe-state behavior. Record the answers to the six checks with it. Any later edit, including one the agent makes, needs a new approval. If the agent also drafted an instrument driver, review it as code; declarative instrument drivers explains why a constrained format keeps that review small.

How to draft, review and run this plan in Galois with Évariste

Évariste, the agent in the Galois platform, can draft, run and report on this plan; the review stays yours. Open it from the app sidebar (Ctrl+Shift+E) beside the project and extend the starter prompt "Create a power rail validation sequence" with this board's requirements:

Create a power rail validation sequence for the 3V3 rail (TLV75533P, 5 V in, EN tied to IN) using psu, dmm, eload and scope. R1: 3.3 V ±3% at 0 and 300 mA, at 4.5 V and 5.5 V in. R2: in regulation within 2 ms of the input reaching 4.5 V. R3: a short from 3V3 to ground damages nothing, and the rail recovers when it is removed. Keep the input at or below 5.5 V with overvoltage protection below 6.0 V, the load off until the rail is up, and the supply off before the load. Put the requirement ID and limit in every step name.

Évariste lists the instruments connected to the team's galois-edge daemons and reads their profile commands. For a load outside the library, it generates a profile from the programming manual you upload and, after your review, deploys it to the bench and binds it. The sequence lands as a draft that cannot run until you approve it.

Review it like any other plan. The reviewed version below has setup steps with outputs off, a gate on the rail, a two-sided numeric_limit step per R1 corner (3.201 to 3.399 V), a scope capture for R2 and a timed short with recovery for R3. Build the coverage matrix from the requirement IDs in the step names and trace each limit to its source. Ask Évariste which SCPI command each step sends, so a VOLT? readback cannot pass for MEAS:VOLT?, and look for steps with one bound. Check R3's dwell, when each R1 reading is taken, and where every branch leaves the bench. Ask for fixes in conversation or edit in the sequence builder; every change is a new version with a diff. Approve the version you read.

Run the approved sequence on an empty fixture, and a known-bad unit if you have one; both should fail. Then run it on the board through galois-edge, watching the rail live in Monitor. Dangerous commands sent from the conversation wait for your confirmation.

Every step records its value, limits, verdict, and raw command and response. Ask Évariste which steps failed or passed close to a limit, such as the 5.5 V, 300 mA corner as the junction heats, or how the board's run compares with the empty-fixture run: a step that passed both proves nothing. Answers cite their runs and notes. "Generate a test report from the last run" produces a PDF or HTML report from a LaTeX template; add the six answers in the report editor and share it to Slack.

You keep the requirements, limits, six checks and approval, plus the wiring, sense leads, short fixture and bench safety. The driver classes, sequencing loop, error handling, command logging and report script are no longer yours to write or maintain.

StepCode path (this guide)Galois with Évariste
DraftAn agent's plan tableR1 to R3 in plain English; Évariste drafts the sequence
DriversInstrument classes you maintainLibrary profiles, or one generated from the manual
Coverage and limitsMatrix and provenance table by handThe same tables, built from step names
Ranges and readbacksRange, mode and query per stepRange in the YAML; Évariste shows each step's SCPI
Do-nothing scoredead_board.py, then a known-bad unitOne-bound steps, readbacks, empty-fixture and known-bad runs
ApprovalSign-off on a commit or document versionVersioned draft with diffs, approved by revision
RunYour test scriptgalois-edge on the bench, watched in Monitor
RecordYour loggingPer-step value, limits, verdict, raw I/O, operator, DUT serial
InterpretRead the logÉvariste flags failures and near-limit passes, compares runs
ReportYour report scriptEditable PDF or HTML report, shared to Slack

How Galois handles AI-drafted test sequences

Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record.

A Galois sequence goes through five stages:

  1. Drafted. Évariste writes the sequence as YAML, and it lands as a draft.
  2. Refused while a draft. A draft cannot start a run.
  3. Approved. The reviewer reads the YAML as a diff. Approval records the approver and the time and applies to that exact version.
  4. Re-approved after edits. If the sequence changes after approval, runs are refused until it is approved again. A sub-sequence it calls must also be approved and in the same project. Production revisions can be locked.
  5. Recorded. Each step records the command sent, the raw response, the measured value, the limits and the instrument; the run records the operator and the unit's serial number.

Steps name logical instruments such as psu and dmm, which the project's bench configuration resolves to physical ones. After the power-up steps, the reviewed plan for the worked example continues like this:

ldo_3v3_rail.yaml (excerpt)
name: "3V3 rail, TLV75533P (R1 to R3)"
steps:
  # action steps first: psu at 4.5 V with limits set and output off, then output_on
  - name: "Gate: run R1 to R3 only if the rail is up"
    type: condition
    config:
      instrument_id: "dmm"
      command_name: "measure_voltage_dc"
      parameters: { range: "10", resolution: "0.0001" }
      operator: ">="
      threshold: 3.0
      on_true:
        - name: "Vout at 4.5 V in, 0 mA (R1, 3.3 V ±3%)"
          type: numeric_limit
          config:
            instrument_id: "dmm"
            command_name: "measure_voltage_dc"
            parameters: { range: "10", resolution: "0.0001" }
            low_limit: 3.201
            high_limit: 3.399
            unit: "V"
            comparison: "GELE"
        # three more R1 corners, then R2 and R3, each set up by action steps
      on_false:
        - name: "Rail absent at 4.5 V in (fails the gate)"
          type: numeric_limit
          config:
            instrument_id: "dmm"
            command_name: "measure_voltage_dc"
            parameters: { range: "10", resolution: "0.0001" }
            low_limit: 3.0
            high_limit: 3.6
            unit: "V"
            comparison: "GELE"
 
  # end state: psu output_off before eload input off (reverse current, datasheet 7.1.4)
  - name: "Disable supply output"
    type: action
    config:
      instrument_id: "psu"
      command_name: "output_off"
 
  - name: "Disable load input"
    type: action
    config:
      instrument_id: "eload"
      command_name: "input_state"
      parameters: { state: "OFF" }

On a dead board, the gate takes its on_false branch, that check fails, and nothing that needs a live rail runs.

The do-nothing check applies directly: each numeric_limit step carries low_limit, high_limit and a comparison that defaults to GELE, and an omitted bound is disabled. For characterization, a measure step records a reading with no verdict; runs count those separately from passes and label a measurement-only run as recorded with no assertions.

Beneath the sequence, the instrument profile bounds what a step can send through it: declared minimum and maximum parameter values are checked before a command reaches the instrument, and ramps must run through the daemon's sweep path. LLM instrument safety covers the guardrails, and the MCP reference covers bench access.

The full loop is in AI test automation for hardware benches, and the product overview shows the sequence lifecycle. To try it on your own bench, start with the quickstart.

Frequently asked questions

How do you review an AI-generated test plan?
Check six things in order. Every requirement maps to steps at its worst-case conditions. Every limit cites a requirement or a datasheet minimum or maximum at matching conditions. Instrument ranges, modes and waits fit the measurement. Power sequencing reaches a safe state on every exit path. A board that never powers up fails the plan. Approval binds to the exact revision you read.
What is a do-nothing run in test plan review?
It is the plan scored against a board that never powers up. Every check that passes in that state proves nothing about the hardware. Look for limits with only an upper bound, differences between two readings, and queries that return an instrument's setpoint instead of a measurement. Fix them with two-sided limits, a gate that confirms the rail is alive, and MEASure queries.
Can I use datasheet typical values as test limits?
Not as pass/fail limits. A typical value is not guaranteed. TI's TLV755P datasheet, for example, gives line regulation (2 mV) and load regulation (0.060 V/A) as typical values with no minimum or maximum. Use the board's requirement, or a datasheet minimum or maximum at the conditions the datasheet states, and record which one each limit came from.
Should an AI agent run the test plan it wrote?
Not before an engineer approves it. In Galois, a machine-authored sequence lands as a draft and the platform refuses to run a draft. Approval records who approved it and applies to the exact version reviewed. If the sequence is edited afterward, it must be approved again before the next run.
Can the agent fix what the review finds?
Yes. In Galois, ask Évariste, the agent in the Galois platform, for each fix in conversation, or edit the sequence in the sequence builder. Every change is a new version with a diff, and an edited sequence must be approved again before it runs. Apply the same six checks to the new version, including a do-nothing run on an empty fixture.

Related

Bring Galois to your bench.

The daemon is Apache-2.0, free forever. Enterprise runs in your cloud or on-prem.