Our story · Built from experience on ARVO

We built a system
that could fix bugs.
Then we taught it
to keep moving.

Good agents can do good work. A useful autonomous system also needs to know when to wait, when to act and when to question its own answer.

What a night waiting for review taught us

The turning point

The process was alive.
Progress had stopped.

On ARVO, we built a Reliability Factory that could investigate production problems, produce repairs, get them reviewed, release them and check the result.

But a PR waiting for an external review held the active slot almost all night. Each scheduled check found the same job. Nothing had crashed. Nothing useful could happen next.

Being busy is not the same as making progress.

One review missingThe same job keeps the slot
CheckWaitCheckWait
Useful work still waiting behind it

ARVO incident, retold. Illustration, not a live log.

The rule that changed the loop

Keep its priority.
Free its place.

A critical job can be the most important thing in the project and still have no permitted action available right now.

Save its cause, review, evidence and return condition. Let another job run. When the review, traffic or decision arrives, bring the waiting job back.

Important and actionable are two different questions.

Critical repair · waiting for trafficPriority and evidence stay intact.
Execution capacity is free
Next permitted job · workingNo reminder from a human needed in the target model.
New evidence → ready for selection again

From Backlog Zero to Operational Zero

A healthy project.
Not just an empty list.

An open item can be understood, safely waiting and still important. The goal is to leave nothing urgent, unexplained or forgotten.

Unexplained signals

Know what each signal means, or the plan to find out.

Urgent work left ready

Handle the critical work that can be acted on now.

Outdated conclusions

Recheck decisions when their evidence gets old.

Waits without a limit

Every wait has a reason and a way back.

Expired experiments

End the test, make a decision and clean up.

Overdue release checks

Verify within the project’s agreed time limit.

Six things we aim to bring to zero. Design targets, not current live measurements.

The exact targets

UNEXPLAINED = 0

ACTIONABLE_P0_P1 = 0

STALE_DECISIVE_JUDGMENTS = 0

UNBOUNDED_WAITS = 0

EXPIRED_EXPERIMENTS = 0

UNVERIFIED_RELEASES_BEYOND_SLA = 0

The next lessons

Even a good answer
can grow old.

New errors kept arriving while real repairs shipped. Some reports shared a cause. Others only looked alike. Old “expected behavior” judgments no longer matched new evidence.

We learned to group only what we can prove, revisit old conclusions and end experiments at a declared deadline.

See the Critic in the loop
“Expected behavior”

Is that still true?

A new release. A different error.
A reason to look again.

Confirm it · Recheck it · Reopen it

Production taught us another rule

A deploy is where
the watch begins.

A release on ARVO accidentally left out a dependency. Chat requests started failing. The normal backlog should not have stayed ahead of that regression.

Severe new failures must move first. Recovery can mean a fix or a rollback, followed by a real check. Waiting for enough traffic must free the next job to run.

Released. Then observed.

  1. NowRelease health
  2. ~1 hourEarly traffic
  3. LaterSustained behavior
  4. 24 hoursReconfirm

Project-configured windows. A new severe regression changes the priority.

More agents.
One shared history.

The Critic questions it. The Lab adds evidence. The reviewer records a verdict on an exact revision. None becomes a second source of truth.

Live sourcesCanonical historyWork & views

After a restart, the system must recover from that durable history. Not from what the last chat session remembered.

Meet the building blocks

The lessons we carry into Orchestro.

ARVO is the experience. Orchestro is where we generalize the rules for different projects.

Running is not the same as progressing.

A process can keep waking up without doing useful work. Measure durable progress and verified results separately from uptime and activity.

Important does not always mean ready.

A critical repair waiting for traffic keeps its priority, cause and evidence. Selection also asks whether a permitted action is possible now.

Waiting work makes room.

Save the reason and return condition, release execution capacity and select another actionable job. Waiting is not resolved; a review, new evidence or a due date can bring it back.

Every required gate has a way forward.

A required external dependency must be controllable or parkable. An unavailable reviewer cannot hold the whole system. A separate reviewer checks the exact code revision; a changed revision needs a fresh check.

Yesterday’s answer has an expiry.

The Critic revisits expected behavior, duplicates, declined fixes, previous cures and measurement gaps. New evidence can confirm the conclusion, require another check, reopen the same problem or leave an honest unknown.

Similar errors need proof of a shared cause.

A matching failure signature, route and release context can suggest a family. Only evidence of the same mechanism justifies joining causes. Keep plausible groups, distinct causes, platform noise and unknowns separate.

A new production failure goes first.

Watch after a release. A severe regression takes priority over ordinary work. Recovery or rollback still needs the project’s authority and its own verification.

Shipped is a milestone, not the finish.

Merge, deployment, distribution and cure carry separate evidence. Check immediately, after an initial traffic window, later that day and at 24 hours where configured. A mobile build is not proof that users received a working release.

Every temporary test has an ending.

Declare the question, baseline, start, due date, sample requirements, success/failure/inconclusive criteria, extension limit, cleanup and final decision authority. Adopt, reject or allow a bounded follow-up. At the limit, record insufficient evidence or ask for an explicit new investment decision.

More specialists. One shared history.

Live sources feed the canonical ledger; selection and views derive from it. The Critic challenges conclusions, the Lab adds evidence and the reviewer adds a revision-bound receipt. None creates a competing scheduler, backlog or source of truth.

A restart must not erase the work.

Reconstruct waiting and ready work from durable history. The current chat session is not the memory of the system. ARVO’s parking and restart tests informed this requirement; they are not proof of a fully adopted unattended production loop.

Health is more than a smaller task count.

Ask what remains unexplained, urgent, stale, waiting without a limit, experimentally expired or released without timely verification. An explained open item can be legitimate. A quiet dashboard can still hide unresolved risk.

The lesson is real.
The next proof still matters.

This account comes from our work on ARVO, shared on 2 October 2026. Parking, releasing the active slot and reconstructing work after restart had been implemented and tested there.

Full production adoption was still incomplete: database migration, GitHub wake events, production release and parts of experiment and post-release monitoring remained. The next test was a 36–72-hour unattended run after that adoption.

That is not a completed autonomy benchmark, or a claim that all these capabilities ship in Orchestro today.

What Orchestro offers today Explore the illustrative loop

The goal is useful progress that keeps its proof.

What would you
put on a loop?

Bring your project