Documentation · illustrative operating examples

One kernel.
Every domain keeps its proof.

The kernel manages persistent work, evidence, authority and independent verification. Each domain supplies its own work model, tools and definition of done.

Watch a day in the loop

Operating model and illustrative examples. Check current availability.

Explore the longer product and campaign stories Illustrative operating models

Illustrative product operating model

A healthier product.

A review is pending. An old bug needs a fix. A new release is about to surprise you. Follow the work through one illustrative morning.

Explore the interactive events
07:00

Waiting is a state. Work keeps moving.

The search fix needs an independent review. Orchestro saves its context, priority and evidence, then frees the active slot. The next useful job can start immediately.

Search fix: waiting for review. Checkout bug: now active.

07:01

37 reports. One thing to fix.

Different error messages point to the same checkout timeout. Orchestro proves which reports share the same mechanism, then coordinates one fix, independent review and release. Similarity alone is not enough.

One underlying cause gets one investigation and one repair.

08:15

When production breaks, priorities change.

After the release, monitoring spots a burst of server errors. A new regression takes priority over routine backlog work. Orchestro investigates the cause and prepares a repair within the project’s permissions.

Urgent recovery moves first. The other work keeps its place.

09:00

The fix is live. Proof takes time.

The regression fix is released, but needs 60 minutes of real traffic before it can be called resolved. Orchestro parks that verification and resumes the backlog. A release alone is never proof of recovery.

Waiting for traffic frees the next job to run.

10:00

The moment it’s ready, it’s back.

The search review arrives. Its saved work becomes actionable again: checks pass, then merge, release and verification. Separately, the regression’s traffic window can now be evaluated against its recovery criteria.

An event wakes the right job. It never had to hold up the morning.

10:30

Even yesterday’s answers get checked.

An old error was classified as expected behavior. Twelve new occurrences look different. The Critic challenges that conclusion, reopens the investigation and checks whether this is the same cause or something new.

A past decision stays open to new evidence.

2–9 Oct

Every experiment has an ending.

A test has a baseline, at least 50 samples and a seven-day deadline. In this example, failure below 1% means adopt; above 3% means reject. The middle band or too little evidence means inconclusive, with at most one bounded follow-up.

When the follow-up ends, decide or record insufficient evidence. Clean up temporary instrumentation.

The result

A healthier product. And the next useful job.

The checkout works again. The urgent regression has passed its traffic checks. The reviewed search fix has its own release evidence. Open questions remain visible, and Orchestro selects the next actionable job.

The goal is to keep your product healthy, continuously.

Illustrative scenario. Scroll to follow the work.

Illustrative campaign operating model

A better campaign.

A creative is waiting for approval. A landing page needs work. Results need time to mature. The same loop, with different tools, evidence and limits. Design-partner example.

Explore the interactive events
09:00

Approval pending. Progress continues.

A new creative needs your approval. Its context stays saved while the team investigates a landing-page problem within the already approved scope.

Waiting for a human decision never grants permission to publish.

09:15

Many weak ads. One broken destination.

Several ads show the same drop after a click. The evidence points to a broken mobile form on their shared landing page. One verified repair can address the shared cause.

Investigate the common cause before rewriting every ad.

10:00

Protect the budget first.

A broken campaign link appears while ads are spending. The loop prioritizes it, pausing the affected delivery only where that permission was granted. New spending still needs your decision.

Act within agreed budget and publishing limits.

11:00

Let results mature. Keep working.

The repaired form passes a separate functional check. Measuring conversion needs a mature cohort and the agreed attribution window. That measurement waits while another approved task moves forward.

A working form and better conversion are separate conclusions.

14:00

Approved. Ready to move.

Your creative approval arrives. The saved work resumes with a quality check, publication within the approved budget and a separate delivery check.

The approval wakes that campaign, with its original limits intact.

Next day

A winner stays a question.

Yesterday’s best creative starts attracting a different audience. The Critic checks cohort, spend, tracking and attribution before deciding whether the old conclusion still holds.

Observed improvement is not automatically a causal lift.

Day 7

Close the test. Keep the learning.

The experiment reaches its declared deadline. Evaluate the target metric, sample requirement and guardrails. Adopt, reject or record inconclusive, with at most one bounded follow-up. More time never appears by default.

No endless test, no silent spend increase, no invented winner.

The result

A campaign you can account for.

The form works. The creative is approved and live. Budget decisions have an owner. Conversion remains a measured question until comparable, mature evidence supports a conclusion.

Every action has a reason. Every result has evidence.

Illustrative scenario. Scroll to follow the work.

What if a release fails?

A failed release stays failed, even when code is merged. Use a bounded retry, an authorized rollback or a technical repair. If a new release harms production, rollback is a first-class recovery option. Then select the next useful job.

What if the sources disagree?

If the work history, repository and production disagree, stop changes to that affected item. Record the authority conflict, repair the evidence and keep other independent work moving.

What if a solved problem returns differently?

Investigate whether the new failure shares the original cause. Reopen the same causal boundary when it does; create a new one when it does not. Similar symptoms alone are not proof of a shared cause.

What if there is never enough evidence?

An inconclusive test can have at most one bounded follow-up in these examples. At the limit, record not measurable or insufficient evidence, or request an explicit owner decision about further investment. Remove temporary instrumentation.

Compare Reliability and Content Interactive simulation
orchestro / mission controlInteractive simulation · sample data
Same kernel. Same rules.
Mission

Keep checkout working.

Checkout service · approved repair scope
  1. 1Observe
  2. 2Understand
  3. 3Select
  4. 4Strategize
  5. 5Execute
  6. 6Verify
  7. 7Learn
  8. 8Next work
Durable work Context stays with the job
ActiveRepair checkout timeoutKernel stage: Observe
WaitingSearch fix · review pendingSaved context · wakes when evidence arrives
Later in this cycleDiscover the next useful jobThe mission continues beyond this task
1 / 8 · Observe

The system finds work.

Monitoring spots failed checkouts. The mission stays active even when nobody is writing a prompt.

Observed failureExample
  • Checkout request → timeout
  • Customer cannot complete an order
  • Source: example application monitor
Simulated receipt

Signal saved with source and affected behavior.

Test the boundaries
Read both domain walkthroughs

Software reliability: Keep checkout working.

  1. Observe: The system finds work.

    Monitoring spots failed checkouts. The mission stays active even when nobody is writing a prompt.

  2. Understand: A symptom becomes a question.

    The investigator reproduces the timeout and links it to an exhausted connection pool. Similar-looking reports still need evidence.

  3. Select: Pick what can actually move.

    Checkout is actionable. The search fix keeps its priority and history while waiting for an independent review.

  4. Strategize: The agent chooses the method.

    It plans a focused repair and a regression test. The kernel checks the affected service, allowed changes and run budget.

  5. Execute: Make the change. Keep the proof.

    The worker prepares the fix. In this scenario, green review gates permit a release. Deployment is recorded; the job remains open.

  6. Verify: A separate check decides.

    The verifier checks the released checkout against the original failure and the agreed observation window. The worker cannot certify itself.

  7. Learn: Keep the lesson, with its limits.

    The repair history records the retry failure and its regression check. A project lesson is not automatically a universal kernel rule.

  8. Next work: One job ends. The mission continues.

    A slow order confirmation becomes new work. The repaired checkout keeps its evidence; the search review still has its place.

Content operations: Run an evidence-led publication.

  1. Observe: The system finds material.

    Evidence mining finds a lesson in project notes. The mission is to run a publication, not just generate one article.

  2. Understand: Build a thesis you can defend.

    The domain checks novelty, source quality and relevance. A promising idea becomes work with explicit editorial acceptance criteria.

  3. Select: Choose the next useful article.

    This article has enough source material to proceed. A previous article waits for its measurement window without blocking new work.

  4. Strategize: Same rules. A different plan.

    The agent chooses the outline, examples and internal links. The domain adds voice, factual and adversarial review requirements.

  5. Execute: Create, review, then publish.

    The draft passes its editorial gates. The publishing adapter acts within existing authorization. A publication receipt is still not proof of a good outcome.

  6. Verify: Read what actually went live.

    A separate check reads the published page and verifies the text, sources and links. Traffic and commercial impact remain a later question.

  7. Learn: Learn without inventing a winner.

    Record the editorial lesson now. Evaluate audience response only when a defined measurement window supplies enough evidence.

  8. Next work: The article creates the next question.

    An unanswered reader question becomes research work. The article keeps its sources, reviews and publication evidence in the same history.

A guided model of the two domains, not a live dashboard. No code is deployed, no content is published. Switch domains at any stage to compare the same kernel.

What is Orchestro?

The mission changes.
The machinery doesn’t.

A runtime for autonomous operations: persistent work, evidence and memory let agents maintain a mission across tasks, sessions and releases.

The kernel keeps the rules.

Durable backlog and memory. Work lifecycle and selection. Evidence, execution and independent verification. Governor, authority boundaries and learning records.

Shared across domains

The domain defines the work.

What a job is. What evidence matters. Which tools can act. The policy, verification rules and conditions that make a job complete.

Reliability: a verified repair.
Content: a checked publication, with impact measured separately.

Cloud guides. Governor supervises.
Factory executes. Evidence decides.

The architectural direction separates a cloud control plane for domains, work, intent, health and receipts from execution on a local or cloud host. Private cloud capabilities exist; a complete hosted, self-service SaaS is still in development.

Cloud architecture and current limits

The building blocks

One core.
Your kind of work.

The core keeps the rules. A project pack brings your goals, tools and checks. The agents do the work.

Each part adds a function. All share one history.

Core

The part every project shares.

The core coordinates the loop: read evidence, choose permitted work, keep its history and require a check. The kernel is the central machinery inside that core.

Why it matters

So each new project inherits the rules instead of inventing another scheduler.

Select a checkout repair only when the evidence and permissions allow it.

Shared private kernel, reused by Reliability and the manually configured Content Factory pilot.

Read every block and its examples

Core: the shared rules

The core coordinates the loop: read evidence, choose permitted work, keep its history and require a check. The kernel is the central machinery inside that core. So each new project inherits the rules instead of inventing another scheduler.

Product: Select a checkout repair only when the evidence and permissions allow it.

Marketing: Apply the same work-and-evidence rules to a campaign, with the project’s own approvals.

Shared private kernel, reused by Reliability and the manually configured Content Factory pilot.

Project pack: your project’s rules

A pack defines the project’s sources, goals, tools, checks, release surfaces and permissions. The core reads this configuration; it does not guess how your business works. So different projects can share a core without sharing credentials, budgets or authority.

Product: Runtime errors, a repository, checkout tests and web or mobile release checks.

Marketing: Campaign signals, approved creative, spending limits and a mature measurement window.

Configured per domain. Content has a real pilot; the campaign example here and a ready-made pack marketplace are not shipped integrations.

Agents & tools: do the work

A supported agent receives a bounded task and the tools the project allows. Providers and subscriptions can change without becoming the owner of the work’s history. So the system can use your tools while keeping scope, cost and permissions explicit.

Product: Investigate the reproduced failure and propose a focused repair.

Marketing: Prepare an approved creative variant or repair the landing page within scope.

Depends on configured adapters, an available host and your agent quota.

Independent checks: prove the result

The agent’s report is not the verdict. A separate review checks the exact revision; release verification checks what users actually receive. A required check must be controllable or allow the job to wait safely. So a finished task cannot quietly turn into an unproven success claim.

Product: The original checkout failure passes on the released version, with enough real traffic.

Marketing: The approved creative is delivered and its destination works. Conversion improvement needs separate evidence.

Checks and review authorities must be configured. A demo receipt is not live proof.

Shared history: remember what is true

Sources and durable receipts record what happened. Ready work, waits and views are reconstructed from that canonical history, rather than the memory of one agent session. So another agent or a restarted process does not have to invent the state again.

Product: Keep the same cause, review, release evidence and traffic wait together.

Marketing: Keep the approval, delivery receipt and measurement window tied to the same campaign.

Durable history is a foundation. ARVO’s demonstrated parking recovery informs Orchestro’s broader target.

Specialists: question and improve

The Governor handles operating limits and stalls. The Critic questions old conclusions. The Lab runs tests with an end date. Each adds a function to the same loop. So learning and supervision improve the system without creating shadow backlogs.

Product: Reopen an old classification when new errors contradict it; end an inconclusive test.

Marketing: Recheck yesterday’s winning creative against a changed audience; retire an expired experiment.

Roles in the operating model. Complete continuous learning and unattended recovery remain roadmap.

An explanation of the architecture, not an installation screen. How project connections work · Current availability

The path behind the progress.

A waiting item leaves the active slot. An urgent regression changes the priority. The whole portfolio keeps moving.

  1. 01

    Notice

    Errors, feedback, QA and post-release monitoring supply the evidence.

  2. 02

    Understand

    Group reports by cause. Explain noise. Give unknowns a bounded evidence plan.

  3. 03

    Choose

    Select actionable work. A serious production regression takes priority.

  4. 04

    Investigate & repair

    Reproduce the failure and fix its cause within the project’s permissions.

  5. 05

    Review independently

    A real finding needs reproduction, a fix and another review. No self-certification.

  6. 06

    Release

    Record merge, deployment and distribution separately for each release channel.

  7. 07

    Watch

    Check the release immediately and over time. Recover or roll back within authority if it regresses.

  8. 08

    Verify

    Require released behavior and the agreed evidence window before declaring a cure.

  9. 09

    Choose again

    Resume the next useful job. Retain evidence and watch for recurrence.

Park. Wake. Resume.

Review unavailable, traffic still maturing or a decision pending: save the work and its wake condition. Return it to selection when it becomes actionable.

Keep questioning the answer.

The Critic can challenge a duplicate, an expected behavior classification or a claimed cure when new evidence contradicts it.

Why these rules exist: what running a Reliability Factory on ARVO taught us. The technical lifecycle below uses the product example.

Every signal has a next step.

First decide what the evidence means. Only actionable work enters the repair path; the other routes keep their context and conditions.

Expected or platform noise

Explain and observe

Keep the evidence. Do not invent a product fix.

Duplicate or known cause

Join its history

Attach the signal to the canonical cause. If cured, check recurrence first.

Unknown

Plan the evidence

Declare baseline, sample target, due date and success, failure or inconclusive criteria. Park until due.

Product bug

Investigate and repair

An actionable cause can enter the repair loop within existing authority.

Feature gap

Resolve the product decision

Get a direction decision where needed before starting implementation.

Operations or cost

Check the operational gate

Respect spending, infrastructure and recovery limits.

Owner or policy

Ask the right person

Record a real permission decision. Keep independent work moving.

200 events do not mean 200 fixes.

150 occurrences of one timeout + 30 platform-noise events + 15 occurrences of another regression + 5 unknowns become two identified causes, explained noise and five open questions. Keep every event linked to its evidence.

Release starts the watch.

Use the observation windows configured for the project. A serious regression can interrupt ordinary work at any point; waiting for a sample frees the active slot.

  1. Immediately

    Check release health

  2. Around 1 hour

    Evaluate early traffic

  3. Several hours

    Check sustained behavior

  4. 24 hours

    Reconfirm the outcome

Verification chooses the way forward.

Each outcome returns to selection. A cure stays open to future evidence; the Critic can challenge any conclusion.

Cured
Record independent proof, keep watching, select the next job.
Same cause returns
Reopen the original boundary and keep its repair history.
New cause
Open a distinct boundary; retain the evidence linking the symptoms.
Failed
Return to remediation, review and the relevant release path.
Inconclusive
Create a bounded follow-up within the extension limit; otherwise record insufficient evidence.
Expected or not a product issue
Change the diagnosis when evidence supports it. The Critic may challenge it later.

From an active task to a verified result.

Inspect a checkout repair in Web, Chat + MCP or CLI. Follow its evidence through a technical wait, release and independent verification. All data is illustrative.

Example
Mobile checkout One job. Explore it in three surfaces.
Ready to investigate

Two reports. One piece of work.

QA and user feedback point to the same checkout failure. Keep the reports together and investigate the cause.

QA finding
Checkout cannot be tapped on mobile Safari.
User feedback
“I fill in my address, but I can’t pay.”
Governor No decision needed

Work is within the project’s approved scope. Investigation can start.

1 of 5

Interactive product example. No live customer data. Availability

Try the difficult moments.

Four product-recovery examples. Choose an event to see how the operating model handles it. These controls only change the demonstration.

What if a release fails?

A failed release stays failed, even when code is merged. Use a bounded retry, an authorized rollback or a technical repair. If a new release harms production, rollback is a first-class recovery option. Then select the next useful job.

Code merged YesRelease succeeded NoRecovery verified No

The affected release stays failed. Independent work can continue.

When the retry limit is exhausted, repair or escalate explicitly. A recovery action never certifies its own success.
What if the sources disagree?

If the work history, repository and production disagree, stop changes to that affected item. Record the authority conflict, repair the evidence and keep other independent work moving.

Repository MergedWork history Claims releasedProduction Old version

Changes to this job are blocked until the sources agree. The rest of the portfolio remains available.

This example corrects a false release claim; it does not invent a deployment.
What if a solved problem returns differently?

Investigate whether the new failure shares the original cause. Reopen the same causal boundary when it does; create a new one when it does not. Similar symptoms alone are not proof of a shared cause.

The checkout was cured under P-B. A new packaging error now appears on the same route. A similar location alone does not identify its cause.

Investigate first. Keep the previous cure and the new evidence visible.

What if there is never enough evidence?

An inconclusive test can have at most one bounded follow-up in these examples. At the limit, record not measurable or insufficient evidence, or request an explicit owner decision about further investment. Remove temporary instrumentation.

Example reliability contract: baseline recorded on 2 October, at least 50 samples, decision by 9 October. Adopt below 1% failure; reject above 3%. The middle band stays inconclusive.

Choose the evidence and evaluate the contract. Elapsed time alone never proves success.

Bounded follow-ups used: 0 / 1. Changing the sample does not reset this limit.

Verification belongs
to the result.

The coding agent’s report is not the final authority. Check the released system against the original failure and agreed acceptance criteria.

Web

Merged is not deployed.Check the actual server, migration or edge-function release and the affected user path.

Mobile

Built is not distributed.Record the app version, distribution channel and device evidence. A store submission is not a verified user result.

Value

Recovery is not business impact.Measure conversion, cost or engagement with defined cohorts and mature data. Keep unproven outcomes explicit.

orchestro / execution replaySimulation · anonymized examples
DEMO-08 · Reliability

Keep checkout working after retries.

Same kernel. Different domain.
  1. 1Observe
  2. 2Investigate
  3. 3Prepare
  4. 4Review
  5. 5Revise
  6. 6Release
  7. 7Verify
  8. 8Next work
1 / 8 · Observe

A mission already has work waiting.

The selector can move “Repair a connection left open after a retry”. Waiting work keeps its context.

shop.example/checkoutPreview
Error trace + regression reproducer

What can we actually support?

Checkout times out after a failed retry.

The reproducer finds a connection left open.

A successful retry must release the connection too.

Recorded in this simulation

No waiting job was discarded

Try an exception

Content examples draw on real editorial work and product-page audits, with names, artifacts and outcomes generalized. Playback compresses the sequence; it does not represent execution time. No command runs and nothing is published.

Read every example without playback

Founder publication: Explain why open work is not always ready work

  1. A mission already has work waiting.

    The selector can move “Explain why open work is not always ready work”. Waiting work keeps its context.

  2. Read the evidence before writing.

    The worker inspects the source and records what it supports. Missing proof is work to do, not permission to invent a claim.

  3. Produce a concrete draft.

    The worker changes content/blog/ready-work.md. The scope stays attached to the work.

  4. The reviewer finds a real problem.

    The draft says “our system” without naming the artifact that supports the claim.

  5. Revise the artifact, then check again.

    Name the work ledger, link the source and keep the project-level limit.

  6. Publish the reviewed revision.

    The configured adapter acts inside the existing authority. A release receipt records the change; the outcome still needs its own check.

  7. Check what actually reached the user.

    Reviewed paragraph, source links and canonical match.

  8. A verified change creates the next job.

    Audience measurement waits for the declared 90-day window.

Product content: Correct an unsupported offline claim

  1. A mission already has work waiting.

    The selector can move “Correct an unsupported offline claim”. Waiting work keeps its context.

  2. Read the evidence before writing.

    The worker inspects the source and records what it supports. Missing proof is work to do, not permission to invent a claim.

  3. Produce a concrete draft.

    The worker changes pages/mobile-companion.md. The scope stays attached to the work.

  4. The reviewer finds a real problem.

    The proposed copy still promises offline behavior the inspected flow does not support.

  5. Revise the artifact, then check again.

    State the connection requirement in the page and FAQ. Preserve its URL and title.

  6. Publish the reviewed revision.

    The configured adapter acts inside the existing authority. A release receipt records the change; the outcome still needs its own check.

  7. Check what actually reached the user.

    Body, visible FAQ and structured answer agree; protected pages unchanged.

  8. A verified change creates the next job.

    Record the baseline; evaluate traffic and CTA impact in a comparable window.

Software reliability: Repair a connection left open after a retry

  1. A mission already has work waiting.

    The selector can move “Repair a connection left open after a retry”. Waiting work keeps its context.

  2. Read the evidence before writing.

    The worker inspects the source and records what it supports. Missing proof is work to do, not permission to invent a claim.

  3. Prepare the patch and its test.

    The worker changes src/checkout/retry.ts. The scope stays attached to the work.

  4. The reviewer finds a real problem.

    The first patch handles success but leaves the failed retry path open.

  5. Revise the artifact, then check again.

    Move cleanup to the shared exit path and add the failing retry to the regression test.

  6. Release the reviewed patch.

    The configured adapter acts inside the existing authority. A release receipt records the change; the outcome still needs its own check.

  7. Check what actually reached the user.

    The released retry passes the original reproduction and the declared recheck window.

  8. A verified change creates the next job.

    Retain the technical verification; business impact remains unmeasured.

What’s real. What’s next.

Implementation and pilot snapshot · 4 October 2026. Private access, with configuration and permissions defined per domain.

Proven in our systems
The Orchestro kernel, autonomous Reliability Factory and Content Factory pilot. Durable state, evidence, independent verification and governed execution. Content uses the same kernel with a manually configured domain.
Required to operate
Readable sources, domain rules, bounded adapters, explicit authority, a running host and agent quota. Pilot availability does not mean unattended operation is guaranteed for every project.
In development
Public self-service onboarding, the phone cockpit, hosted MCP/chat access and further domains. Domain Manifest, Domain Builder / Compiler V1 and capability reconciliation are implemented in the private core. The cloud is not a finished hosted SaaS.
Read the availability reference

Your goals. Your permissions. Your evidence.

Start with a project
you know.

Request early access