How Orchestro works

Every job
keeps its proof.

A continuous loop coordinates many pieces of work. Each one keeps its cause, permissions, history and evidence from first signal to verified result.

Watch a day in the loop

Operating model and illustrative examples. Check current availability.

The building blocks

One core.
Your kind of work.

The core keeps the rules. A project pack brings your goals, tools and checks. The agents do the work.

Each part adds a function. All share one history.

Core

The part every project shares.

The core coordinates the loop: read evidence, choose permitted work, keep its history and require a check. The kernel is the central machinery inside that core.

Why it matters

So each new project inherits the rules instead of inventing another scheduler.

Select a checkout repair only when the evidence and permissions allow it.

Private foundations. The full continuous operating model is still being developed.

Read every block and its examples

Core: the shared rules

The core coordinates the loop: read evidence, choose permitted work, keep its history and require a check. The kernel is the central machinery inside that core. So each new project inherits the rules instead of inventing another scheduler.

Product: Select a checkout repair only when the evidence and permissions allow it.

Marketing: Apply the same work-and-evidence rules to a campaign, with the project’s own approvals.

Private foundations. The full continuous operating model is still being developed.

Project pack: your project’s rules

A pack defines the project’s sources, goals, tools, checks, release surfaces and permissions. The core reads this configuration; it does not guess how your business works. So different projects can share a core without sharing credentials, budgets or authority.

Product: Runtime errors, a repository, checkout tests and web or mobile release checks.

Marketing: Campaign signals, approved creative, spending limits and a mature measurement window.

Configured per project. Marketing is a design-partner example, not a ready-made pack marketplace.

Agents & tools: do the work

A supported agent receives a bounded task and the tools the project allows. Providers and subscriptions can change without becoming the owner of the work’s history. So the system can use your tools while keeping scope, cost and permissions explicit.

Product: Investigate the reproduced failure and propose a focused repair.

Marketing: Prepare an approved creative variant or repair the landing page within scope.

Depends on configured adapters, an available host and your agent quota.

Independent checks: prove the result

The agent’s report is not the verdict. A separate review checks the exact revision; release verification checks what users actually receive. A required check must be controllable or allow the job to wait safely. So a finished task cannot quietly turn into an unproven success claim.

Product: The original checkout failure passes on the released version, with enough real traffic.

Marketing: The approved creative is delivered and its destination works. Conversion improvement needs separate evidence.

Checks and review authorities must be configured. A demo receipt is not live proof.

Shared history: remember what is true

Sources and durable receipts record what happened. Ready work, waits and views are reconstructed from that canonical history, rather than the memory of one agent session. So another agent or a restarted process does not have to invent the state again.

Product: Keep the same cause, review, release evidence and traffic wait together.

Marketing: Keep the approval, delivery receipt and measurement window tied to the same campaign.

Durable history is a foundation. ARVO’s demonstrated parking recovery informs Orchestro’s broader target.

Specialists: question and improve

The Governor handles operating limits and stalls. The Critic questions old conclusions. The Lab runs tests with an end date. Each adds a function to the same loop. So learning and supervision improve the system without creating shadow backlogs.

Product: Reopen an old classification when new errors contradict it; end an inconclusive test.

Marketing: Recheck yesterday’s winning creative against a changed audience; retire an expired experiment.

Roles in the operating model. Complete continuous learning and unattended recovery remain roadmap.

An explanation of the architecture, not an installation screen. How project connections work · Current availability

The path behind the progress.

A waiting item leaves the active slot. An urgent regression changes the priority. The whole portfolio keeps moving.

  1. 01

    Notice

    Errors, feedback, QA and post-release monitoring supply the evidence.

  2. 02

    Understand

    Group reports by cause. Explain noise. Give unknowns a bounded evidence plan.

  3. 03

    Choose

    Select actionable work. A serious production regression takes priority.

  4. 04

    Investigate & repair

    Reproduce the failure and fix its cause within the project’s permissions.

  5. 05

    Review independently

    A real finding needs reproduction, a fix and another review. No self-certification.

  6. 06

    Release

    Record merge, deployment and distribution separately for each release channel.

  7. 07

    Watch

    Check the release immediately and over time. Recover or roll back within authority if it regresses.

  8. 08

    Verify

    Require released behavior and the agreed evidence window before declaring a cure.

  9. 09

    Choose again

    Resume the next useful job. Retain evidence and watch for recurrence.

Park. Wake. Resume.

Review unavailable, traffic still maturing or a decision pending: save the work and its wake condition. Return it to selection when it becomes actionable.

Keep questioning the answer.

The Critic can challenge a duplicate, an expected behavior classification or a claimed cure when new evidence contradicts it.

Why these rules exist: what running a Reliability Factory on ARVO taught us. The technical lifecycle below uses the product example.

Every signal has a next step.

First decide what the evidence means. Only actionable work enters the repair path; the other routes keep their context and conditions.

Expected or platform noise

Explain and observe

Keep the evidence. Do not invent a product fix.

Duplicate or known cause

Join its history

Attach the signal to the canonical cause. If cured, check recurrence first.

Unknown

Plan the evidence

Declare baseline, sample target, due date and success, failure or inconclusive criteria. Park until due.

Product bug

Investigate and repair

An actionable cause can enter the repair loop within existing authority.

Feature gap

Resolve the product decision

Get a direction decision where needed before starting implementation.

Operations or cost

Check the operational gate

Respect spending, infrastructure and recovery limits.

Owner or policy

Ask the right person

Record a real permission decision. Keep independent work moving.

200 events do not mean 200 fixes.

150 occurrences of one timeout + 30 platform-noise events + 15 occurrences of another regression + 5 unknowns become two identified causes, explained noise and five open questions. Keep every event linked to its evidence.

Release starts the watch.

Use the observation windows configured for the project. A serious regression can interrupt ordinary work at any point; waiting for a sample frees the active slot.

  1. Immediately

    Check release health

  2. Around 1 hour

    Evaluate early traffic

  3. Several hours

    Check sustained behavior

  4. 24 hours

    Reconfirm the outcome

Verification chooses the way forward.

Each outcome returns to selection. A cure stays open to future evidence; the Critic can challenge any conclusion.

Cured
Record independent proof, keep watching, select the next job.
Same cause returns
Reopen the original boundary and keep its repair history.
New cause
Open a distinct boundary; retain the evidence linking the symptoms.
Failed
Return to remediation, review and the relevant release path.
Inconclusive
Create a bounded follow-up within the extension limit; otherwise record insufficient evidence.
Expected or not a product issue
Change the diagnosis when evidence supports it. The Critic may challenge it later.

From an active task to a verified result.

Inspect a checkout repair in Web, Chat + MCP or CLI. Follow its evidence through a technical wait, release and independent verification. All data is illustrative.

Example
Mobile checkout One job. Explore it in three surfaces.
Ready to investigate

Two reports. One piece of work.

QA and user feedback point to the same checkout failure. Keep the reports together and investigate the cause.

QA finding
Checkout cannot be tapped on mobile Safari.
User feedback
“I fill in my address, but I can’t pay.”
Governor No decision needed

Work is within the project’s approved scope. Investigation can start.

1 of 5

Interactive product example. No live customer data. Availability

Try the difficult moments.

Four product-recovery examples. Choose an event to see how the operating model handles it. These controls only change the demonstration.

What if a release fails?

A failed release stays failed, even when code is merged. Use a bounded retry, an authorized rollback or a technical repair. If a new release harms production, rollback is a first-class recovery option. Then select the next useful job.

Code merged YesRelease succeeded NoRecovery verified No

The affected release stays failed. Independent work can continue.

When the retry limit is exhausted, repair or escalate explicitly. A recovery action never certifies its own success.
What if the sources disagree?

If the work history, repository and production disagree, stop changes to that affected item. Record the authority conflict, repair the evidence and keep other independent work moving.

Repository MergedWork history Claims releasedProduction Old version

Changes to this job are blocked until the sources agree. The rest of the portfolio remains available.

This example corrects a false release claim; it does not invent a deployment.
What if a solved problem returns differently?

Investigate whether the new failure shares the original cause. Reopen the same causal boundary when it does; create a new one when it does not. Similar symptoms alone are not proof of a shared cause.

The checkout was cured under P-B. A new packaging error now appears on the same route. A similar location alone does not identify its cause.

Investigate first. Keep the previous cure and the new evidence visible.

What if there is never enough evidence?

An inconclusive test can have at most one bounded follow-up in these examples. At the limit, record not measurable or insufficient evidence, or request an explicit owner decision about further investment. Remove temporary instrumentation.

Example reliability contract: baseline recorded on 2 October, at least 50 samples, decision by 9 October. Adopt below 1% failure; reject above 3%. The middle band stays inconclusive.

Choose the evidence and evaluate the contract. Elapsed time alone never proves success.

Bounded follow-ups used: 0 / 1. Changing the sample does not reset this limit.

Verification belongs
to the result.

The coding agent’s report is not the final authority. Check the released system against the original failure and agreed acceptance criteria.

Web

Merged is not deployed.Check the actual server, migration or edge-function release and the affected user path.

Mobile

Built is not distributed.Record the app version, distribution channel and device evidence. A store submission is not a verified user result.

Value

Recovery is not business impact.Measure conversion, cost or engagement with defined cohorts and mature data. Keep unproven outcomes explicit.

What can you use today?

Execution depends on the configured project. The complete presentation is an operating model, not a claim that every integration ships today.

Private core
Structured intake, conservative grouping, issue history, bounded remediation, evidence plans and separate merge, release and verification records.
Project configuration
Readable sources, agents, test commands, release surfaces, permissions and independent verification. An available host and agent quota are required.
Model & roadmap
The full unattended loop, general event-driven wakeups, autonomous recovery and continuous learning remain planned. Growth loops are for design partners.
Read the availability reference

Your goals. Your permissions. Your evidence.

Start with a project
you know.

Request early access