Skip to content

Manual TUI testing — the driver (#1161)

For contributors investigating interactive TUI behavior and maintainers recording a manual play-test. This page explains the driver that sends keys and waits for observable state.

Read Container testing first: the driver belongs inside that isolated sandbox. For browser behavior, use the web client selftest.

Driving the real multi-pane TUI by hand for a play-test gate used to be error-prone: keys landed in a live pane as literal text (#1156), hand-rolled tmux harnesses died on $TMUX/TMUX_TMPDIR collisions (#1155), and every run re-derived send-keys/capture-pane/sleep N plumbing from scratch — where the blind sleeps were the main flake source.

scripts/tui-driver.sh is the fix: a sourceable driver library of self-synchronizing functions that drive the live TUI and assert on it. It builds ON the #1130 container sandbox (which already solves isolation — throwaway home, mock repo, private tmux, pids/memory caps; see container-testing.md). The container is the where; this driver is the how.


1. The interaction model — read this first

The #1156 mis-drive happened because the driver didn't know which mode the TUI was in. Get this right and the rest follows.

Mode What the keyboard does Menu bar shows How you got here
Nav Keys are TUI commands (n new, s open pane, t new tab, j/k move the tree cursor, …) context hints (n new · …, ↵ interact · …) default; Ctrl-] from interactive
Interactive Every key — including Tab — forwards to the focused pane's shell/agent only ctrl+] nav mode Enter on a focused pane
Attached (full-screen) tmux owns the terminal; the TUI chrome is gone the tmux status line ([af_…) o on a selected instance

The single most important primitive is af_ensure_nav — it sends Ctrl-] (a no-op in nav, the escape hatch from interactive) so a scenario can never mistake a live pane for the host. Call it before any nav action. This is the one-line fix for the entire #1156 class.

The focus ring

In nav mode, Tab/Shift-Tab cycle ring focus across regions: the instances tree → each open pane → the automations strip → the projects section → back. The menu bar is context-sensitive and follows focus:

  • tree focused → n new · … (plus instance verbs when a row is selected)
  • pane focused → ↵ interact · o attach · ← prev pane · → next pane · … · s open pane · x hide pane
  • automations focused → enter manage · …
  • projects focused → ↵ switch · / search │ tab focus · ? help · q quit

af_focus_tree walks the ring (checking before it presses, so it never Tabs off the tree) until the n new hint proves the tree has focus.

The instances tree

The left rail is a tree: each instance row carries its tabs as children. The display-selected instance auto-expands (arrow , its tab children shown) while every other instance collapses (arrow ). That arrow is the reliable, text-greppable selection signal the driver asserts on:

  ▾ alpha          ●      ← selected (expanded); ● = ready
       └ 1 ◆ Agent *
  ▸ beta           ●      ← not selected (collapsed)

On a cold boot the cursor sits on the section header (no instance selected — menu shows the plain n new · / search · ? help · q quit set); the first j moves the cursor down onto the first instance. af_select is robust to any starting position: it anchors at the top (k is idempotent there), then steps down with j until the target row shows and the tree cursor is actually on it.

Display-selected ≠ cursor-on-row (#1174 / #1199). The sticky is only a display signal. A single auto-selected instance renders while the cursor still sits on the section header — there GetSelectedInstance() is nil, so o/D/attach silently no-op even though the row looks selected (a false pass waiting to happen). af_select therefore also requires the menu to advertise a row-scoped verb (D delete session, or the legacy D kill the driver still accepts) — present only when a real instance is under the cursor — and keeps pressing j past the header until it appears. af_attach/af_open_pane inherit the fix because scenarios call af_select first.

Which key goes where (defaults, from keys/keys.go)

Key Nav action Key Nav action
n new instance s open selected tab as pane
Enter interact (enter pane) x hide focused pane
o attach full-screen t new tab
Ctrl-] exit interactive → nav w delete tab (asks first; y confirms)
Tab cycle focus ring 19 jump to tab
j/k,/ move tree cursor m tasks overlay
/ switch focused pane [/] previous/next section
D kill instance / search
a archive r restore
q quit
e hooks editor
Ctrl-P switch project Ctrl-U/Ctrl-D preview scroll
Ctrl-W detach (full-screen)

While naming a new instance the form owns the keyboard and its keys are fixed: Tab opens the program picker, Shift-Tab opens the initial-prompt field (#1936), Ctrl-R opens the backend field (#1933), Enter submits, Esc/Ctrl-C cancel the create. Inside the prompt field Enter is a newline — Tab/Esc save the prompt edit and return to naming. Ctrl-C discards only the current prompt edit and returns to the naming form with the previously saved prompt and all other pending values intact. Ctrl-C from naming cancels the create.

Enter, Tab, Shift-Tab, Esc, Ctrl-], and 19 are reserved and cannot be rebound ([keys] config). Ctrl-W is the configurable detach key (detach_keys); the driver reads it from AF_DRIVER_DETACH_KEY.

The detach key is recognized in every encoding a terminal may report it in, not just the legacy control byte: an agent CLI that turns on the kitty keyboard protocol or modifyOtherKeys makes the terminal send Ctrl-W as an escape sequence instead (#1832). af_detach takes an optional raw sequence so a scenario can drive those encodings without that terminal.

Detaching also drops those negotiated modes back to legacy, so the terminal you land back in reports Ctrl keys normally again. On a real kitty terminal this is the check the driver cannot make for you: attach to a claude session, detach, and confirm Ctrl-] and friends still work in the TUI — a terminal left upgraded needs a manual reset to recover. / switch panes only when a workspace pane has focus; with tree focus they keep the tree's collapse/expand behavior.


2. The wait-not-sleep principle

Never synchronize on a fixed sleep. A sleep 2 is a bet that the TUI finished repainting in under two seconds — it flakes when it's slow and wastes time when it's fast. Instead, every action helper returns only once the screen shows its completion marker:

af_send n                      # request a new instance
af_wait_for 'submit name'      # ← block until the name prompt actually appears

af_wait_for <regex> [timeout] [label] polls capture-pane (a short poll interval, not a synchronization sleep) until the screen matches, and dumps the last screen on timeout so a failure is debuggable. af_wait_gone is the inverse. Every scenario helper is built from these, so a scenario is a straight-line list of intent with no timing knobs.


3. Driver command reference

Source the library inside the container, then call the functions. State lives in the running TUI (a private tmux session, default name drive), so calls compose across docker exec invocations.

Anti-flake core

Function What it does
af_wait_for <re> [t] [label] poll until the screen matches <re> (no blind sleeps)
af_wait_gone <re> [t] [label] poll until <re> is absent
af_ensure_nav Ctrl-] → force nav mode (fixes the #1156 class)
af_focus_tree Tab the ring until the instances tree has focus

Scenario helpers (each self-synchronizes on its marker)

Function Keys Completion marker
af_reset_sandbox wipes instances/branches for a deterministic rerun (sandbox-scoped, fails closed)
af_boot launch af Agent Factory frame + the session rail painted — Sessions ( or the scrolled ▲ N more the rail shows in its place (#2148)
af_new_instance <name> n,text,Enter the row shows <name> … ● (ready)
af_select <name> k×,j× <name>'s row shows
af_open_pane s pane-focus menu (x hide pane)
af_hide_pane x visible-pane set changes (pane header row)
af_enter_interactive Enter interactive menu (ctrl+] nav mode)
af_exit_interactive Ctrl-] interactive menu gone
af_send_to_pane <text> marker+text,Enter short delivery marker echoes in the pane (then wait for output yourself)
af_attach o TUI chrome gone (full-screen)
af_send_line <text> [t] text,Enter the WHOLE line echoed on the attached screen. Fails closed (#2147): it clears and re-pastes up to AF_DRIVER_SEND_LINE_ATTEMPTS times, and if the line still never lands complete it returns non-zero having sent no Enter — a submitted partial leaves an unbalanced quote and drops the shell into >
af_detach [raw_seq] Ctrl-W (once) TUI chrome back and the attach client is reaped (guards #1157)
af_new_tab t tab-child count rises
af_close_tab w → wait for “Delete tab” → y dialog must appear before consent; then tab-child count falls
af_open_tasks / af_close_tasks m / Esc rounded dialog frame containing the task manager's pinned footer appears / gone
af_click <x> <y> / af_click_instance <name> SGR mouse injects a left click at a cell / on an instance row
af_scroll <up\|down> [x] [y] SGR wheel injects a wheel event
af_set_config <toml> + af_relaunch rewrites config.toml (canonical since #1030) and reboots the TUI
af_resize <cols> <rows> pins the session to an exact geometry that STICKS (window-size manual + resize-window), for tiny-size gates
af_quit q back to a shell prompt

Assertions & introspection

Function Pass condition
af_assert_screen <re> / af_refute_screen <re> screen matches / does not match
af_expect_selected <name> <name>'s row carries
af_tmux_ls prints the tmux sessions (introspection)
af_ps prints the daemon + tmux attach/new-session process tree
af_assert_no_orphan_clients no tmux attach-session reparented to init (the #1155/#1157 leak signature); the daemon's own monitor clients are parented to the daemon and excluded

Configuration (env vars)

AF_DRIVER_SESSION (drive), AF_DRIVER_COLS/ROWS (100/30), AF_DRIVER_REPO ($HOME/sandbox/mock-repo), AGENT_FACTORY_HOME, AGENT_FACTORY_AUTO_UPDATE (driver sets to false to prevent mid-test self-updates that would timeout instance creation), AF_DRIVER_TIMEOUT (25s), AF_DRIVER_POLL (0.25s), AF_DRIVER_DETACH_KEY (C-w), AF_DRIVER_BIN (auto-resolved), AF_DRIVER_HELP_SEEN (15).

Set AF_DRIVER_COLS/ROWS before af_boot to launch at a non-default size — af_boot pins it so it sticks (see the tiny-size gate below). Change the size mid-run with af_resize <cols> <rows>.

Driving a first-run overlay

af_boot writes help_screens_seen before launching, and the default 15 marks every one-time overlay seen — that suppression is what makes ordinary scenarios deterministic, but it also means no scenario could reach a first-run screen at all. Clear the bit for the overlay you want:

bit overlay
1 general help (?)
2 instance-start help (n)
4 instance-attach help (o)
8 interactive-pane help (enter on a live pane)
export AF_DRIVER_HELP_SEEN=7   # everything seen EXCEPT the interactive help

The bits are app/help.go's mask() methods. Re-run af_reset_sandbox before a second boot in the same container — it wipes state.json, so the overlay is genuinely "first run" again; without it the first boot has already marked it seen. scripts/tui-2413-scenario.sh is a worked example.


4. Running it

The self-test (acceptance proof + bitrot guard)

make tui-driver-selftest

Boots a dedicated container sandbox (af-driver-selftest, so it never disturbs a drive/playtest container you have open), then runs the exact scenario that failed in #1156, now deterministic:

reset → boot → create two instances → select each (assert selection) → open a pane → enter interactive → type into the pane → exit → attach full-screen → detach → assert selection preserved → assert no orphan clients.

Green means the driver drives the TUI reliably. Any failure prints the step and the offending screen.

In CI

The TUI driver selftest workflow (.github/workflows/tui-driver-selftest.yml) runs the same scripts/testbox.sh selftest entry point on pushes to master that touch the TUI, its tmux session layer, or the container/driver harness. Auto Gate merges with GITHUB_TOKEN, which suppresses those push events, so its master-side workflow list re-dispatches the selftest after every later merge. Manual and gate-raised dispatches always run regardless of changed paths. The roughly ten-minute run on each merge is intentional while the signal is measured; the runner is free for this public repository. The daemon, TUI, stand-in sessions, and private tmux server remain inside the ephemeral testbox container.

This is the signal-only first step from #4000: it deliberately has no pull_request trigger and does not gate merges. Its purpose is to identify the exact master commit that makes the scenario red while runtime is measured. Failed or cancelled runs upload the full console stream, including the driver's last-screen dump, as the tui-driver-selftest-transcript artifact.

One scenario script (a per-fix real-TUI gate)

scripts/testbox.sh scenario scripts/tui-2413-scenario.sh

Runs a single scenario script in the same ephemeral, uniquely-named sandbox the self-test uses, then tears it down. The path is repo-relative (the repo is mounted read-only at /src); pin AF_SELFTEST_NAME to reuse a container instead.

Use this for a regression scenario that belongs to one fix. Do not bolt such a case onto tui-driver-selftest.sh: that scenario is the shared acceptance proof, and destabilizing it costs more than the bug the new case guards.

Driving by hand

make tui-driver          # boots af via the driver, then attaches you to the
                         # live session (detach with your tmux prefix + d)

Or drive over docker exec against a detached sandbox. The container name is unique per run (#1171); pin it with AF_PLAYTEST_NAME so your docker exec targets it:

export AF_PLAYTEST_NAME="af-playtest-$$"
make playtest-container-detached
docker exec "$AF_PLAYTEST_NAME" bash -lc '
  source /src/scripts/tui-driver.sh
  af_boot
  af_new_instance alpha
  af_new_instance beta
  af_select beta && af_expect_selected beta
  af_open_pane && af_enter_interactive
  af_send_to_pane "echo hi"
  af_exit_interactive
  af_assert_no_orphan_clients
'
docker rm -f "$AF_PLAYTEST_NAME"   # teardown

Everything runs inside the container — the host tmux server, the real ~/.agent-factory, and this repo are all untouched.


5. Gate-recipe library

To gate a visible-TUI PR, run the scenario for its class and assert the markers. All of these are driver calls; each already self-synchronizes.

Any TUI-visible change → the smoke gate

make tui-driver-selftest

The self-test is the baseline gate for any PR that touches startup, the sidebar/tree, panes, interactive mode, or attach/detach. If it isn't green, stop.

Create-form changes (the #1936 class)

The naming form is a multi-field form behind a single row, so gate every field plus the paths that leave it:

af_boot
af_ensure_nav; af_focus_tree
af_send n; af_wait_for 'submit name'
af_wait_for 'initial prompt'                  # the field is advertised
af_send Tab; af_wait_for 'Select program'     # sibling field still opens
af_send Escape; af_wait_for 'submit name'
af_send BTab; af_wait_for 'enter newline'     # shift+tab opens the field
af_send_literal 'first line'; af_send Enter   # enter is a NEWLINE here…
af_wait_for 'enter newline'                   # …so the field is still open
af_send Tab; af_wait_for 'initial prompt ✓'   # closes, hint confirms it stuck
af_send BTab; af_wait_for 'first line'        # reopening shows the text
af_send Escape                                # esc keeps the text (field, not dialog)
af_wait_for 'initial prompt ✓'

The backend field (#1933) is the third field and needs a wide terminal: its hint sheds below ~93 columns (ui/menu.go hintDropOrder), so drive this at 120+ or the marker will legitimately be absent.

af_boot                                       # 120 cols or wider
af_ensure_nav; af_focus_tree
af_send n; af_wait_for 'backend'               # the field is advertised
af_send C-r; af_wait_for 'Select backend'      # daemon round trip, then the list
af_wait_for 'Repo default'                     # first row names the resolved default
af_send Escape; af_wait_for 'submit name'      # esc backs out of the field only
af_send C-r; af_wait_for 'Select backend'
af_send Down; af_send Enter                    # pick the row below Repo default
af_wait_for 'backend ✓'                        # hint confirms a non-default backend

A backend the repo cannot use is listed with — unavailable (or — cannot check) and refuses the pick with the daemon's own reason — the same sentence af sessions create --backend <that one> prints. Selecting it must leave the form open with the hint back to backend and no . Then finish a create on a backend the repo CAN use and confirm the session actually comes up there: the round trip is the whole point and no marker stands in for it.

The account field (#3844) is the fourth, and it needs a wider terminal still: its hint sheds first, below 112 columns. Register an account on the daemon host before driving it (af accounts add claude <name>), or the field will correctly report that there is nothing to pick:

af_boot                                       # 112 cols or wider
af_ensure_nav; af_focus_tree
af_send n; af_wait_for 'account'              # the field is advertised
af_send C-o; af_wait_for 'Select claude account'   # the title names the AGENT
af_wait_for "Use the agent's own login"       # no configured default in this fixture
af_send Escape; af_wait_for 'submit name'     # esc backs out of the field only
af_send C-o; af_send Down; af_send Enter      # pick a registered account
af_wait_for 'account ✓'                       # hint confirms a scoped create
af_send Tab; af_send Down; af_send Enter      # change the program…
af_wait_for 'submit name'

The last two lines are the load-bearing leg. An account belongs to ONE agent, so changing the program must drop the pick: the hint has to go back to account with no , and reopening the field must list the NEW agent's accounts. Then finish the create and confirm the session actually runs as that account — af sessions list --json reports it on the session's account field, and a session the daemon created WITHOUT the account raises an error naming both identities rather than reporting the one you picked. That last case is version skew and only an old daemon produces it; the marker to look for is the phrase did not apply account.

scripts/tui-3844-scenario.sh automates every leg above except that final create: the sandbox points claude at a bash stand-in, and the account boundary refuses to launch an agent whose command it cannot prove is a direct invocation of that agent. The scenario turns that into evidence rather than skipping it — the refusal comes from the daemon and NAMES THE ACCOUNT, which is only possible if the picked value rode CreateSessionRequest.Account across the wire. A session that actually RUNS as the account needs a real agent binary, which is why that step stays here.

The project default (#3386) is the same field arriving pre-filled. Set it for the project first, with af's own verbs rather than by hand — the point is that the documented gesture produces it:

af projects add <repo>                               # once
af config set default_accounts.claude <name> --project <repo>
af_boot; af_ensure_nav; af_focus_tree
af_send n; af_wait_for 'submit name'
af_wait_for 'account ✓'          # NO ctrl+o, NO pick — this is the whole feature
af_send C-o; af_wait_for '<name>.*project default'   # and it says WHY
af_send Escape
af_send Tab; af_send Down; af_send Enter             # change the program…
af_wait_gone 'account ✓'                             # …and the claude default drops
af_send Tab; af_send Up; af_send Enter               # back to claude…
af_wait_for 'account ✓'                              # …and it returns, per agent

The account ✓ with no keypress is the assertion that distinguishes this feature from the daemon quietly applying a default on the create: the session is identical either way, and what #3386 adds is that the user sees which identity it will run as before pressing enter. The appear → disappear → reappear sequence on that one marker is also what makes the check non-vacuous — a marker that were always painted would fail the middle leg.

Set the key to an account that is NOT registered to see the other half: the row is still offered, appended last and labelled project default · not registered, because hiding it would leave the form reporting the ambient identity while the config says otherwise. That create is refused by the daemon naming the key and the file it is set in.

scripts/tui-3386-scenario.sh automates all of the above, with the same stand-in caveat as #3844's: the create is refused by the account boundary, and the refusal naming the account is the evidence that a value nobody typed was resolved from the project's config and carried across the wire.

enter newline is the overlay's own hint row, used as the marker rather than its Initial prompt title: the status-bar hint underneath says initial prompt too, so the title alone cannot tell "field open" from "field advertised".

Then finish the create by hand and confirm the agent receives the prompt as its first input — that round trip is the whole point of the feature and no marker can stand in for it. Re-run n afterwards and confirm the hint is back to initial prompt with no : a prompt must never leak into the next session.

Test the form in a repo that declares a non-local backend, too (#2599). The naming row used to be built by provisioning the create's runtime, so in a repo whose .agent-factory/config.json says backend = "docker" (or ssh/hook) pressing n ran a real provisioner and the form never opened at all. Every create-form gate above passes in a local repo while that is broken, which is why this is its own step:

mkdir -p "$AF_DRIVER_REPO/.agent-factory"
printf '{"backend": "docker"}\n' >"$AF_DRIVER_REPO/.agent-factory/config.json"
af_boot; af_ensure_nav; af_focus_tree
af_send n; af_wait_for 'submit name'          # the form opens at all
af_send_literal 'declared'; af_send Enter
af_wait_for 'docker'                          # the DAEMON refuses, naming docker

The second wait is the load-bearing one. A create that succeeds here as a local session means the placeholder's backend was pinned local and the repo's declared backend went with it — which passes "the form opens" and silently gives the user a session their repo did not ask for. The refusal has to come from the daemon and has to name the backend. scripts/tui-2599-scenario.sh automates all three legs (form opens, backend honored, ctrl+rlocal still creates).

Usage-limit account switching (the #3127 class)

scripts/tui-3127-scenario.sh drives the whole opt-in flow against a real daemon. It is not a TUI-rendering gate — the decision under test is the daemon's — but the driver is what makes it reachable: every unit test in daemon/ and session/ swaps a seam, and none of them proves a real daemon polling a real tmux pane reaches the decision at all.

Two things about it are worth stealing for any account gate.

The stand-in has to be PROVABLE. configure-playtest-agent.sh installs its bash stand-in as af-playtest-standin, and the account boundary refuses to scope a command it cannot prove is a direct invocation of the agent — so that default can never be account-scoped. This scenario installs its own at $HOME/bin/claude and, crucially, sets no program_overrides at all: trustBase (session/program_resolution.go) admits an override only when it is byte-for-byte af's own detected built-in, which here is the quoted path plus --dangerously-skip-permissions. Writing the bare path instead — the obvious thing — refuses every candidate with "could not be proven to be a direct claude invocation".

A successful swap is a TRANSIENT state. Park, decide, tear down and replace takes about four seconds, so polling for liveness_name=limit-reached on the path that MOVES is a race the scenario loses. Assert the durable facts instead — account, account_auto_selected, and the daemon's own auto-resumed limit-blocked session … on claude account "…" line — and keep the limit-state assertion for the paths where af must NOT move, where the wall is where the session stays. The stand-in walls every identity except the swap target for exactly that reason.

For "the replacement really runs as the account", ask the LIVE process rather than matching its startup line: enter the pane and echo $CLAUDE_CONFIG_DIR. A startup line proves what the process was told at exec; the echo proves what it still has — and the in-session switch notice is long enough to push that startup line out of the visible pane anyway.

Tree / selection / focus changes (the #1156, #1084 class)

af_boot
af_new_instance a; af_new_instance b; af_new_instance c   # (cap: 3)
af_select a; af_expect_selected a
af_select c; af_expect_selected c
af_select b; af_expect_selected b
# after confirmed tab deletion, selection must not silently drift:
af_new_tab; af_open_pane; af_close_tab; af_expect_selected b

Pane / interactive changes (the #1088, #1089 class)

af_select a; af_open_pane
af_enter_interactive
af_send_to_pane 'echo PANE_OK'; af_wait_for 'PANE_OK'
af_exit_interactive
af_refute_screen 'nav mode'          # cleanly back in nav
af_hide_pane                          # pane hides, nothing killed

Attach / detach changes (the #1155, #1157, #1159 class)

af_select a
af_attach                             # full-screen
af_detach                             # syncs on the attach client being reaped
af_assert_no_orphan_clients           # the hard leak check
af_expect_selected a                  # selection survives the round trip

Tabs (the #930 class)

af_select a
af_new_tab; af_new_tab                # add two shell tabs
af_close_tab                          # waits for Delete tab, sends y, then waits for removal
# Manual equivalent: w opens consent; y accepts, n/Esc cancels.
# Hiding a pane is reversible and does not delete its tab.

Config / keymap changes (the #1030 class)

af_set_config "$(cat <<'TOML'
default_program = 'claude'
[program_overrides]
claude = 'bash'
[keys]
new = ['c']
TOML
)"
af_relaunch
af_ensure_nav; af_focus_tree
af_send c; af_wait_for 'submit name'  # the rebound 'new' key works

Mouse (the #1143 class)

af_new_instance a; af_new_instance b
af_click_instance a; af_expect_selected a
af_click_instance b; af_expect_selected b
af_scroll down; af_scroll up

Tiny geometry / responsive-layout changes (the #1174-item-2 class)

A detached tmux session defaults to window-size latest, which snaps the window back to the last-attached client (80x23) and ignores new-session -x/-y. So a naive small-size boot silently ran at 80x23 and never exercised the tiny layout. Boot with the size preset (af_boot pins it), or resize mid-run with af_resize:

AF_DRIVER_COLS=60 AF_DRIVER_ROWS=15 af_boot   # boots pinned at 60x15
af_new_instance a
af_select a; af_expect_selected a             # selection still works when narrow

af_resize 40 10                               # squeeze to 40x10 mid-run
af_assert_screen 'Sessions'                   # header must survive the squeeze

6. Gating a branch cut BEFORE #1166 (the driver isn't in the tree yet)

scripts/tui-driver.sh landed in #1166. A branch cut before that commit has no driver to source, so source /src/scripts/tui-driver.sh fails with No such file. Two ways to gate such a branch — pick one before you boot:

  • Rebase the branch onto master (preferred when the branch is yours and rebasing is clean) — this pulls the driver + self-test into the tree naturally, and you gate exactly what will merge.
  • Copy the driver in from master when a rebase is noisy or the branch is an external PR you don't want to rewrite. Inside the running sandbox container:
# From the host, into the sandbox container (name is unique per run, #1171):
docker cp scripts/tui-driver.sh          "$AF_PLAYTEST_NAME":/src/scripts/
docker cp scripts/tui-driver-selftest.sh "$AF_PLAYTEST_NAME":/src/scripts/
# then drive as usual — the driver is pure harness, so master's copy gates
# any older product tree without changing what you're testing.

Because the driver only sends keys and reads the screen — it carries no product code — master's copy is safe to run against an older af build; it asserts on the same on-screen markers regardless of the branch under test.


7. Isolation & box safety (inherited from the container)

Every rule from the tui-playtest skill is satisfied structurally by running inside the container: private tmux server, throwaway AGENT_FACTORY_HOME, pre-built mock repo, pids/memory caps, teardown is docker rm -f. The driver reinforces this:

  • It only ever kills its own named session and (in af_reset_sandbox) the sandbox's af_* sessions — never kill-server.
  • af_reset_sandbox fails closed: it refuses to wipe anything unless AGENT_FACTORY_HOME and the mock repo are sandbox paths, so it can never touch a real ~/.agent-factory.
  • Sessions run the cheap bash program (the sandbox's config.toml override), never a real agent or an unbounded generator.

Performance measurements

make perf-container drives this library against the same 1,000-session fixture as the browser, recording three full-frame and key-to-render samples. See perf-baselines.md for the measurement boundaries and budgets.