STATEMETHOD

Stage 2 / Controlled Pilot

Prove the smallest useful workflow under controlled conditions.

A Controlled Pilot turns an evidence-backed workflow design into working software. It uses representative cases and a bounded number of real integrations to test whether the workflow delivers enough value, whether review is practical and whether failures can be detected and recovered. It is not an open-ended production build.

Starting fee
From €14,000

The buyer problem

A convincing output is not the complete operating path.

A prototype can produce a convincing output without proving the complete operating path. It may not use real permissions, handle duplicate actions, expose uncertainty, preserve traceability or support a reviewer. The pilot tests those conditions before production hardening is considered.

Prerequisites

A pilot starts from evidence, not a demo.

A pilot needs:

  • One bounded workflow and named owner
  • A measurable baseline and business objective
  • Representative cases, including known difficult cases
  • Defined output and review responsibility
  • Access to the required sandbox, staging tools or APIs
  • Known security, privacy and data-handling constraints
  • Measured acceptance criteria
  • A documented reason to proceed from prior evidence

The State Method Diagnostic is a common route to these inputs, but equivalent prior evidence can be reviewed.

Typical scope

Build only what the unresolved assumptions require.

A pilot may include one bounded workflow, one or a small number of real integrations, a sandbox or staging environment, model and deterministic processing, defined tool permissions, a review step, evaluation cases, trace logs, timeout and retry behavior, failure routing, recovery behavior and basic operating documentation.

Scope is agreed around the assumptions the pilot must test. Additional workflows and integrations are separate.

Real-tool integration

The pilot should use enough of the real operating path to expose integration and authority issues. That may mean a real API with restricted credentials, a staging data store or an internal review interface. Access is limited to the actions required by the pilot. Production authority is not implied.

Evaluation

Representative cases are defined before the pilot conclusion. The evaluation should include normal cases, ambiguous inputs, known-bad inputs, edge cases and relevant permission or recovery scenarios. Results are compared with the manual baseline and agreed thresholds. Model quality is assessed inside the complete workflow, including abstention and review effort.

Permissions and approvals

The permission design names what the model may propose, which checks software must enforce, which tools the workflow may access and what a person must approve. Approval records the responsible person and the exact action being accepted, rejected or corrected.

Failure testing

The pilot tests the failures that matter to the workflow: invalid inputs, unavailable tools, timeouts, duplicate requests, rejected approvals, uncertain output and interrupted processing. Each relevant case needs an observable signal and a defined outcome such as safe retry, manual handling or stop.

Deliverables

Working evidence inside the agreed boundary.

Depending on the agreed scope:

  • Working pilot for one bounded workflow
  • One or a small number of real integrations
  • Representative evaluation set and results
  • Implemented permissions and approval points
  • Logs and traceability for tested decisions
  • Timeout, retry and idempotency behavior where required
  • Failure routes and recovery behavior
  • Basic operating documentation
  • Acceptance review and go/change/stop recommendation

Pilot acceptance criteria

Define the decision thresholds before the conclusion.

Criteria are defined before the decision and may cover:

  • Value compared with the current baseline
  • Model behavior on representative and difficult cases
  • Correct operation of deterministic checks
  • Practical effort for human review
  • Prevention of unauthorized actions
  • Traceability of tested decisions
  • Visibility of declared failure states
  • Reliable recovery in the tested boundary

Acceptance demonstrates only the agreed pilot scope. It does not guarantee performance on unseen conditions or production readiness.

Qualification boundary

Good fit and explicit exclusions.

Good fit

  • A Diagnostic or equivalent evidence supports a bounded build
  • Real tools are required to test the remaining assumptions
  • The client can provide representative cases and an operating owner
  • Acceptance can be measured in a sandbox or staging boundary
  • The production decision depends on evidence the pilot can collect

A Controlled Pilot does not automatically include

  • Unrestricted production deployment
  • Broad organizational change
  • Every desired integration
  • Unlimited workflow expansion
  • High-availability infrastructure
  • Full compliance certification
  • Removal of human approval
  • Guaranteed production readiness, ROI or accuracy

Go, change or stop

A technically working pilot is not enough.

The final review compares results with the predeclared acceptance criteria. The decision can be to scope production hardening, change and retest part of the workflow, gather better evidence, retain manual handling or stop. A technically working pilot is not enough if the operational value or review burden is unacceptable.

FAQ

Before the next decision.

Can State Method start from our prototype?

Yes. The prototype, evaluation data, permission model and operating evidence are reviewed first. Reusable work is retained where it supports the pilot boundary.

How long does a pilot take?

No standard duration is stated. Timing depends on integrations, evaluation cases, approval flow and the assumptions being tested. It is defined in the pilot scope.

Does the pilot use production data?

Not by default. The proposal defines the approved data and environment. Representative, sanitized or staged data may be sufficient. Production access requires a separate, explicit boundary.

Can actions run automatically?

Only actions explicitly permitted by the pilot design may run without a person. Consequential actions keep the required approval gate. All tested action authority remains bounded.

What happens when the pilot passes?

Passing supports a production-scoping decision for the proven boundary. It does not automatically authorize deployment.

Continue with the evidence

Review the next relevant boundary.

Controlled Pilot / evidence in use

Test the unresolved assumptions with working software.

If the workflow already has a boundary, baseline and representative cases, share the evidence available and the decision the pilot must support.