ToWow

Harness engineering: six practices for multi-day agent work

Harness engineering is the work of designing everything around a model that decides what it reads, what it may do, how its output gets checked, and when it stops. Birgitta Böckeler, writing on Martin Fowler's site, notes that the term has become shorthand for "everything in an AI agent except the model itself" (2026-04-02). OpenAI describes the engineer's job as "to design environments, specify intent, and build feedback loops that allow Codex agents to do reliable work" (OpenAI, 2026-02-11). Prompt engineering covers the instructions in one request. Context engineering, in Anthropic's definition, covers "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference" (Anthropic, 2025-09-29). Choosing an orchestration library such as LangGraph decides how steps are wired together. The harness is the work that remains whichever library you pick.

Zhang Chenxi (Nature), who builds agent systems for manufacturers and distributors. . Drafted with AI assistance.

This article lists six practices from Flowness, our own multi-agent framework, which has been running for over six months on real delivery work. Each answers a question that appears when many agents work for days: how much new work to start, how many times to retry a failure, what the next agent reads, when to stop gathering information, how many alerts a person sees, and what a final report leaves out. Five sections describe decision rules in our code as of September 2026, with teaching examples. The sixth comes from a dated run record. Each section labels its evidence.

Existing guides focus on coding agents in one repository

Five pieces cover this ground. OpenAI's report on a product built with no hand-written code covers the repository as the record of knowledge, linters and tests that enforce architecture, and background tasks that clean up drift. Böckeler's article sorts controls into guides that steer before an agent acts and sensors that observe afterward, and into computational and inferential checks. Google's developer blog argues for behavioral evaluations that work like integration tests for a harness (Google, 2026-09-09). Anthropic's account of long-running agents describes a progress file and git history that let a fresh session see the state of work (Anthropic, 2025-11-26), and a later piece separates the agent doing the work from the agent judging it (Anthropic, 2026-03-24).

They focus on a coding task in one repository, worked by one agent or by a few agents in sequence; Anthropic's later piece uses a planner, a generator, and an evaluator. Running many agents at once for days, on separate tasks that share usage limits, alert channels, and the people who answer them, adds other questions, and the six practices below address those.

1. Admit new work by forecast, not by current headroom

A usage window that is 60 percent used says little by itself. One hour to the reset and four hours to the reset leave different amounts of room, and a rate of 10 points per hour differs from a rate of 30. Our scheduler reads the share used, the time to reset, and the consumption rate. For the rate it takes the largest of three estimates: the smoothed recent rate, the token rate converted to quota terms with historical ratios, and total use divided by elapsed time. The largest is the cautious choice, and it keeps a brief quiet spell from making the planner add sessions.

The rule then turns room into a concurrency target. Take a teaching example: the short window is 60 percent used, the reset is one hour away, and the target ceiling is 90 percent. The sustainable rate is 30 points per hour. With four sessions running at a predicted total of 30 points per hour, the target stays at four. If the predicted total is 15 points per hour (say the window has run for four hours, so the average-use estimate is also 15), the target is four times 30 divided by 15, which is eight. The same planner also checks that its inputs are fresh (300 seconds by default) and that the session list is complete. When the forecast reaches the 95 percent warning line within 15 minutes, it narrows the entrance to new work.

The first change is to the admission plan. A target that drops from four to two does not kill two running sessions, because waiting, finishing, or migrating belongs to the execution layer. The 90 percent target, the 95 percent warning line, the 300-second freshness limit, and the 15-minute lead time are configuration defaults in our code as of September 2026. The usage, reset time, rates, and session counts in the example are made up for teaching. The planner runs in shadow mode by default, and automatic migration is off by default, so this describes a decision rule, not an action that already runs on its own. This is a form of backpressure: it slows the entrance before the pressure arrives. You would test it by comparing unplanned interruptions, completed work, and waiting time under the same workload (Flowness write-up on quota admission, in Chinese).

2. Give recovery its own budget and keep the stop

A failed task usually deserves another attempt, since a dropped connection or a temporary service fault passes on its own. The trouble is a scheduler that scans for unfinished tasks on every pass and hands the same failure to a new agent each time. Suppose a task keeps failing because it lacks an access permission. A fifth attempt will not create the permission, and each attempt still costs a session, logs, and money.

Our dispatcher counts self-healing redispatches per task, separately from the first launch, and compares the count with a cap. The default cap is 3. At the cap, the task leaves the automatic dispatch pool, the system writes a persistent stop marker, and it emits an event saying the budget is exhausted. Later scans see the marker and neither redispatch nor repeat the alert, so the act of stopping does not become a background chore of its own.

Three boundaries matter. First, this is closest to a circuit breaker, yet nothing probes the task again after a wait, and a person or a repair process clears the marker. Clearing it without changing the failing condition sends the task around the same loop. Second, a cap and a wait between attempts address different costs. The wait spreads out attempts, and the cap ends a path that cannot recover. Third, the cap limits launches only. An action with side effects still needs its own deduplication, as the replay section of our LangGraph alternative article explains. The default of 3 is an engineering choice, and you would judge it by counting which retry succeeds, by failure type (Flowness write-up on bounded redispatch, in Chinese).

3. Rebuild each handoff's context from sources

When one agent hands work to another, someone prepares what the next agent reads. A summary is cheap, and a summary can shrink an acceptance rule to "finish the main features" and drop the exceptions. Sending every file crowds out what matters. Anthropic's progress file is one answer, and it lives in the repository. Our handoff adds a check on the package itself.

For each role (executor, fixer, reviewer, planner), a program lists what must be read, and a caller cannot downgrade it. A known counterexample stays mandatory even when the caller marks it optional and puts it last. Each item carries its source and the range to read, and reopening a stored package rereads the source and recomputes the selection. A replaced passage that still cites a real source fails that check. A record separates what was delivered in full, what was replaced by a short account of a query receipt (what the query reported, what it covered, and what it still leaves open), and what is still pending. The code allows that replacement for one narrow receipt format only. If nobody searched for counterexamples upstream, the record says "not searched" and does not suggest the search came back empty.

Here is a teaching example. The budget is 4,096 bytes, and the task packet alone is 5,000 bytes of selected text. A normal delivery does not trim the packet's tail and call it complete. It reports what could not be shown. A delivery that requires complete mandatory reading includes the whole packet and marks the budget as exceeded. The code measures bytes of the final rendered text, which is not the same as tokens. Our test files include cases for long acceptance criteria that stay whole, a low-ranked counterexample that still becomes mandatory, and a stored package that is checked against its sources when reopened. We have not tied this to a task success rate (Flowness write-up on handoff context, in Chinese).

4. Decide when to stop with checks first and scores second

Question count and transcript length do not tell you whether requirements are ready for design. A user can describe many scenarios and still not say who approves the launch. Our stop rule runs four required checks before any score: whether every gap on a hard requirement is closed (in the example below, who may see customer files), whether the direction is stable, whether the information can support a design, and whether the parts only the user can judge are covered. If one check fails, the result is "not enough", even when every quality score is 0.9.

If all checks pass, four quality dimensions are compared one by one with a threshold, 0.6 by default: breadth, specificity, usability in conversation, and overall quality. One low dimension yields "enough with conditions", so a high score on one dimension cannot hide a low score on another. Here is a teaching example for a customer file upload system. Access permissions are confirmed and all four checks pass. The scores are 0.8, 0.7, 0.55, and 0.9. The third is below 0.6, so the result is "enough with conditions", and the next stage knows which dimension is still weak. If the third is later rated 0.65 with evidence, the result becomes "enough".

The scores come from an assessment step. They are not a measured quantity such as information entropy, and 0.65 does not mean 65 percent reliability. The dimensions borrow the idea of information power from qualitative interview research (Malterud, Siersma, and Guassora, 2016): the more relevant information the material holds, the less you need to collect. The four checks, the four dimensions, and the 0.6 threshold are our own rules, not taken from that paper. The final decision still belongs to the main session and the user. Treat 0.6 as a parameter to calibrate: if tasks judged "enough" keep finding key gaps during design, revisit the check definitions and the threshold (Flowness write-up on the interview stop rule, in Chinese).

5. Group repeated alerts under the open problem they repeat

A health check that runs every few minutes reports one unresolved fault with changing numbers and timestamps. Create a ticket per report and a to-do list shows one problem as a dozen. Remembering only the last message fails when two messages alternate. Our alert path asks a different question: is there already an unresolved alert from this machine about this problem?

A match needs the same machine identity, the same normalized problem fingerprint, and an open alert younger than 14 days. The fingerprint replaces long hexadecimal IDs and digits and keeps the first 40 characters of the question. Sources differ in what they need. One health-check source reports sampled values in its reason field, so using the reason would split one problem into many, and its fingerprint uses the question alone. A relay monitor asks the same sentence about rejected writes, stale heartbeats, or a backlog of messages, so its fingerprint adds a 10-character digest of the full normalized reason.

In a teaching example, a health check raises alert A, and two later checks report the same problem with different samples. Both attach to A, and the recurrence count goes from 0 to 2 while the list still shows one item. If A has been resolved, the next report opens a new alert, because grouping does not stand in for a fix. The recurrence record keeps a count, a last-seen time, and the latest reason sample of up to 300 characters. It does not preserve every reason. Only machine sessions take part, and two named worker sessions that send the same words stay as two items, because each is waiting on its own decision (Flowness write-up on alert grouping, in Chinese).

6. Give every kind of cost a place in the final report

A report written by an agent can be accurate and still hide a cost. In a July 2026 run, an executor finished a change to a completion rule. Its independent review could not start, because a function raised a TypeError on an argument it did not accept. The executor followed the written procedure: log the infrastructure problem, degrade explicitly, and say so in the report. The defect was fixed about 32 minutes after it was logged.

The same session had a different interruption. A request to an adviser sent at 08:00:10 ended at 08:03:49 when the connection dropped. The next message arrived at 10:01:16, typed by a person (the Chinese word for "continue"). Between them the session record is empty for 1 hour, 57 minutes, and 27 seconds. The final report explained the review failure and said nothing about the gap. Review failure had a path into the report, and an adviser outage and a manual resume did not. The timestamps show the gap in the record. They do not show what the model knew while it was idle.

The practice is to decide in advance where each kind of interruption, recovery, and human intervention appears in the final report, and to take those facts from records rather than from the agent's recollection. The run log can supply when a request started, failed, and resumed, and a source field can show whether a resume message came from a person or from the runtime. When we investigated this run, we cross-checked the session record against the event ledger and the git history, since a single self-report sees only part of the work. The practice comes from one recorded run. The way to test it is to compare final reports item by item with the run record and check whether each clear interruption is kept and each human resume is identified as one (Flowness write-up on reporting blind spots, in Chinese).

Handoff sheets with per-task ownership, a read-back from the target before a step counts as done, storing dedup evidence in the same save as the result and moving the progress marker last, generation numbers that stop a stale worker from writing, and a pause that says what stops are covered in LangGraph alternative for long-running agent work. The study behind the read-back practice is When "done" did not happen. For how we use the word harness, see agent-to-agent harness.

If you run agents for days and have hit a case these practices do not cover, write to us at hi@towow.ai. Send the failure as it happened, with timestamps if you have them.

If you want a system like this built around your own workflow, see AI agent systems for manufacturers and distributors or write to hi@towow.ai.