SuperQode
WorkOrders
Jev
Monty
Pi Durable
Harness Engineering

SuperQode: WorkOrders, Jev Engineering, Monty and Pi Durable

October 4, 2026
10 min read
By Shashi Jagtap
SuperQode: WorkOrders, Jev Engineering, Monty and Pi Durable

SuperQode β€’ WorkOrders β€’ Jev β€’ Monty β€’ Pi Durable

SuperQode: WorkOrders, Jev Engineering, Monty and Pi Durable

SuperQode 2.5.5 is the release this post is about. Recently we released SuperQode 2.5.0 with a polished terminal interface, Herdr status, and a clearer PiPy surface next to the Pi coding agent. Developers could connect a harness, an agent, or a provider and stay in the TUI. Looking at that release against Pi Durable, the WorkOrder path was still not in a state we were happy to compare. Admission, evidence, and recovery existed, but they were not as durable, or as easy to inspect, as the checkpointed runs people were reading about in Pi Durable.

So we kept going. The releases after 2.5.0 were aimed at making WorkOrders closer to Pi Durable: retry-safe admission, durable evidence references, predecessor evidence reuse, and opt-in PiPy recovery after a process interruption. This post is about that work, and about two neighbouring questions that came up while we were doing it.

The first is Monty. We use it for optional, restricted Python programs inside a WorkOrder: compose admitted tools, process evidence, and resume a suspended program without repeating calls that already committed. The second is Jev engineering for coding agents, which had just been published. We wanted an honest read of where SuperQode stands against that guidance, not a claim that we already match it. This post covers both, alongside the WorkOrder and Pi Durable comparison.

Open a WorkOrder with :work view WORK_ORDER_ID. The inspector stays open beside the work and rereads committed state every two seconds. Overview shows task dependencies, attempts, and acceptance commands. Evidence shows what a later task was assigned from an earlier one. Recovery shows each operation's state. Review shows acceptance output and the candidate diff. You can queue, run, recover a stale lease, resume, check acceptance, and prepare a candidate without leaving the view.

SuperQode remains the harness layer. Recovery, evidence reuse, and context selection stay opt-in.

Evidence with provenance

Dependent workers can retrieve predecessor evidence instead of repeating the same investigation. The inspector keeps two labels separate. Freshness says whether the source files and lineage still match: current, stale, or unknown. Verification says what kind of claim the text is. An investigator report starts as reported even when its sources are current. A tool receipt is observed_output: it records which bytes came back, and its freshness stays unknown until something checks them.

Pages are bounded. Each page is checked against current permissions. Reuse stays useful, and an unchanged report stays a reported finding rather than a proven fact.

Recovery you can read

Opt-in PiPy recovery reconstructs a coding run and reuses committed model and tool outcomes after a process interruption. The TUI is where that distinction becomes visible. A committed outcome is a finished call on record. A reuse admission is a later decision that this same call may be skipped. The Recovery panel shows committed, reuse admitted, in flight, retry eligible, retry authorized, and reconcile needed. Private model and tool results stay out of the view.

An interrupted operation with an unknown effect needs reconciliation. A command may have changed a file, or a provider may have accepted a request, before the worker could commit its outcome. Recovery leaves that unknown side effect unresolved until a person records a verified result or explicitly authorizes a retry. Interrupt stops the action process that this inspector launched. After a restart you can read the persisted rows, recover a stale lease, and resolve anything uncertain before continuing.

Candidate review is the human gate on the delivery candidate. Approval asks you to open the complete candidate, which has to match its stored digest, and to enter a review reason. Identity and digest are checked again at decision time, so a replacement candidate cannot take an approval that belonged to the one you opened. Acceptance checks still have to pass. Approving the candidate leaves merging as its own action.

That gate approves a delivery candidate. It is separate from an interactive prompt on every PiPy tool call. Hosted PiPy checks current call and result policy, including recovered results. It still has neither an interactive tool-approval prompt nor an OS sandbox. Standalone PiPy runs with the permissions of the launching process.

Jev engineering in SuperQode

While this WorkOrder work was landing, Jev engineering for coding agents was published. The sources we read are the Jev engineering page and the Jev Engineering for Coding Agents notes. We read it as a comparison point for context, evidence, and execution, and we checked SuperQode against the parts we had already chosen to build.

SuperQode's approach to Jev engineering focuses on better context, reusable evidence and reliable execution. Durable context references keep retained tool evidence addressable, while a shared Core and PiPy policy selects bounded excerpts when needed. Selection decisions can be reused while their inputs remain valid, reducing unnecessary context churn. WorkOrders extend this approach through predecessor evidence with provenance and freshness checks. These controls remain opt-in, with shadow mode preserving baseline behavior and enforcement enabled explicitly. This release implements selected Jev principles. Live evaluations are still needed to establish their effect on coding correctness, context size, completion time and spend.

:context evidence is the inspector for native Core and hosted PiPy context decisions. It shows the selection mode, the selector, proposed excerpts, actual and proposed character counts, and whether a cached decision was reused. You can open a bounded page of the retained original under current permissions. Shadow mode keeps the baseline prompt and shows the proposal beside it. Enforce mode is what substitutes excerpts. Context selection defaults to off. Character counts are not a measured token saving. Model routing and price optimization stay outside this work.

Where this stands against Pi Durable

Pi Durable is a separate framework from the Pi terminal coding agent. It checkpoints tasks through the agent loop, and it supports multiple conversations, background work, and application state. Automatic wake-up through durable alarms belongs to its Cloudflare hosting integration, rather than to every local Pi Durable process.

We treated that design as the bar. Opt-in PiPy recovery now rebuilds a run after a process interruption and reuses committed model and tool outcomes. Committed tool results are reused, explicitly safe operations can retry, and an uncertain effect waits for reconciliation. Around that sits the WorkOrder delivery path: dependencies, isolated worktrees, evidence with provenance, acceptance commands, a reviewed diff, human approval, and a guarded merge.

Pi Durable is still ahead on the runtime around the loop. It resends an interrupted model request and keeps an aborted partial answer. We block an unknown model outcome for reconciliation, which is more conservative and less autonomous. Its worker hosting can wake an interrupted agent. Our worker service recovers stale leases while it is running, and a full restart still needs supervision such as systemd, launchd, or Kubernetes. Pi Durable also compacts in the background and offers more general checkpointed tasks, timers, ownership trees, and typed documents. We have compaction, durable evidence references, and opt-in context selection. We have not shown that broader runtime yet.

We tried to bring SuperQode WorkOrder recovery as close to Pi Durable as these releases could honestly support. It is real recovery for opt-in PiPy WorkOrders, and it is not full Pi Durable compatibility. The next few releases will harden WorkOrders toward that compatibility: interrupted-request handling, supervised restart, and durable background execution. Matched live failure tests, and independently graded coding runs, are what would support a stronger comparison. Until those exist, we are not claiming better recovery, correctness, or speed.

Monty inside a WorkOrder

Monty powers optional, restricted Python programs that compose admitted tools and process evidence. Inside a WorkOrder we persist a suspended interpreter and its host-call intent, commit the outcome, and resume without repeating committed calls. Those checkpoints cover the program. They do not checkpoint a shell process, recreate dependencies, or restore a lost filesystem. Model and tool recovery comes from the separate PiPy recovery driver. A change to code, capabilities, policy, workspace, or the Monty runtime can block continuation. That prevents stale execution. It also means a checkpoint is not freely portable across an upgrade or a changed repository.

Monty implements a Python subset. Current Monty supports simple classes and some async features. It does not import third-party packages or the project environment, so NumPy, pandas, pytest, and application imports belong in the normal project execution tools. Our PiPy python_program path is further limited to reads: it can discover and compose admitted read tools. The regular coding loop is what runs shell commands, edits files, writes project files, uses the network, or starts nested orchestration. The native python_repl bridge has a different capability profile, and the two should not be treated as the same tool.

Programs stay small on purpose. Current integration caps are 32 host-call units, 30 seconds per attempt, 32 MiB of interpreter memory, and an 8 MiB checkpoint. A parallel batch allows up to eight trusted native reads. Returned text and structured evidence are bounded, and images are omitted. An MCP tool has to be explicitly admitted. Marking a tool read-only does not establish that repeating it is harmless or free.

Monty's own snapshots can already move between processes or machines, and compatible dump formats can cross releases. Our exact worker-binary fingerprint is a stricter SuperQode safeguard. The time limit, the call cap, the read-only program profile, which MCP capabilities we admit, worker supervision, and workspace restoration are SuperQode choices.

This hosting case is concrete enough to discuss upstream. These are proposals, not shipped features and not submitted requests:

  • Snapshot inspection and a compatibility preflight, so a host can read dump kind, format version, stored limits, and compatibility through a supported API before it tries to load the interpreter.
  • A caller-defined serialization limit, so a dump can stop when the host's own budget, such as our 8 MiB cap, is exceeded.

Preserving consumed suspension counters across restoration, with remaining counters a host can inspect, would also help. So would a durable async recovery example that restores pending futures from externally persisted call IDs and committed outcomes. Monty already documents versioned dumps, preserved accumulated execution time, manual resolution of restored futures, a 256 MiB dump cap, and a suspension count that resets on restoration. There is an existing suspension-limit discussion. We would read that thread before opening a duplicate. We would not ask Monty for unrestricted package imports, shell access, or full CPython compatibility. The restricted interpreter is intentional. Those capabilities stay behind controlled host tools.

Try this release

Install:

Shell
curl -fsSL https://superqode.dev/install.sh | sh
sq

Already installed:

Shell
superqode update
# or: uv tool upgrade superqode

Install SuperQode, then run the small coding demo. It sets up an investigator and a dependent implementer around a failing acceptance test. Setup is offline. Execution uses the provider and model you name. The sample demo enables recovery with --recovery. Recovery can also be enabled through a PiPy HarnessSpec. See the WorkOrders docs, the PiPy docs, and tool composition and recovery.

What we want to hear now

Tell us where :work view already makes a WorkOrder understandable while it runs, and where evidence, recovery, or candidate review still hide the state you need. We also want to hear whether the Monty hosting limits match how you would use a restricted program inside a coding WorkOrder, and which Pi Durable behaviours you want WorkOrders to match first.

πŸ“š Our blogs are also published on

Follow along wherever you already read

πŸ’‘ Found this helpful? Share it with your network and help others discover these insights!

Try it

Try this release

Install or update SuperQode. Recovery is on only when you pass --recovery.