How to Define Success Criteria for Your First AI… | ORKA

How to Define Success Criteria for Your First AI… | ORKA

Treat a first AI process as successful only when it solves a defined business task with acceptable quality, measurable operational impact and supervision whose cost does not erase the value. The success criterion does not start with an AI ROI formula or a count of generated responses. It starts with the process baseline, the decision supported by the output and the point at which the pilot should stop or be redesigned.

Teams often start a pilot by asking how much time AI will save. That is a valid business question, but it is too early to use as the only criterion. In a first pilot, the organisation is still testing several basic assumptions:

The number of queries, automated documents or rapidly produced drafts can look positive without improving the process. AI may create more drafts while an expert spends more time checking, correcting and locating sources. It may also speed up one step while total cycle time remains unchanged because of approval waits, missing data or manual transfer into the ERP.

That is why AI success measurement must cover the full workflow, not only the behaviour of the tool.

A baseline is not a broad statement such as "the process is slow". It is a brief, verifiable description of current work. For the selected process, record:

You do not need to measure every detail from day one. You need enough information to compare the situation before and after the pilot. If the organisation already tracks processing time or corrections, use those records. If it does not, agree on a simple sampling method across a representative set of cases.

It is important to separate a process measure from user perception. An employee may find a tool useful while process records show more exceptions. The reverse is also possible: early learning can create resistance even when manual work falls after the process stabilises. Both signals matter, but they answer different questions.

The quality of an AI output is not a universal accuracy score. The measure depends on the task and the consequence of an error.

For classifying incoming invoices, quality may mean assigning the document to the correct category and approval flow. For preparing a customer response, it may mean factual grounding, alignment with internal rules and an appropriate tone. For summarising technical documentation, it may mean coverage of key points with a clear connection to the source.

Before the pilot, answer four questions:

It is often useful to split quality into three levels:

This split is more useful than an average score. It also reveals the nature of failure. A system that often needs small edits may suit draft preparation with review. A system that occasionally produces a serious incorrect output may not suit the task without additional constraints, regardless of favourable averages.

A risk management framework can help define accountability, context and oversight. A relevant starting point is the NIST AI Risk Management Framework 1.0 . The framework does not replace a process decision, but it helps structure questions about measurement and risk.

A high-quality output without operational impact is not necessarily a candidate for wider use. The operational measure should describe what changes in real work. Select one primary measure and a small number of supporting indicators.

The primary measure may be time from case receipt to a ready result, expert active time per case, the number of cases completed without return, or time needed to find relevant information. Supporting measures can track manual transfers, escalations, approval bottlenecks and cases outside the defined scope.

Avoid mixing tool time with process time. If AI creates a draft quickly but the user then waits for access to source data or approval from another department, the business process is not equally fast. Review the full journey from trigger to the next usable decision.

For processes connected to production, accounting or ERP, pay particular attention to points where data moves between systems. A pilot that appears successful in an isolated test can lose value during real entry, master-data alignment or exception handling. ERP and process screening can help map these dependencies before a larger change.

AI ROI is not only the difference between estimated time before and after the tool. It also includes the work needed for safe, repeatable operation:

Supervision is not necessarily a sign of a weak solution. For many tasks, human review remains a required part of the process. The issue arises when supervision is invisible in the success measure. If a pilot reduces drafting time but increases review time or escalations, the team needs to see that change clearly before deciding to expand.

Measure supervision cost in the actual work context. Do not assume every user reviews every case at the same speed. Different documents, exceptions and experience levels create different workloads. Also record which controls exist because of business risk and which arise from unreliable output. That distinction supports a better process redesign decision.

AI pilot criteria should contain more than the question "does it work?". Before starting, write down the decision each possible finding will support.

Continue to limited use when output meets the agreed quality level, the operational measure shows a relevant change and supervision remains workable within existing roles and controls.

Redesign the pilot when the business task is sound but it needs clearer instructions, better inputs, a narrower scope, a different approval flow or a more precise integration. Redesign should have its own hypothesis and new measurement, not simply more testing without a decision.

Stop the pilot when a critical error cannot be detected early enough, when required supervision removes the operational value, when data is not available in an acceptable form or when the process scope is too unstable for a reliable assessment. Stopping is a useful decision: it prevents the spread of an unclear solution and releases the team to examine a more suitable use case.

The stop criterion should be specific to the selected process. For AI-assisted preparation of an internal response, a relevant threshold may be an error that would change a business obligation without expert review. For document routing, it may be an inability to reliably direct cases with particular attachment types. Do not use general thresholds borrowed from another process.

Imagine a pilot where AI prepares a draft response for an employee using approved internal procedures. The scope is not "automate support". It is preparing drafts for one limited group of enquiries.

The baseline includes time spent locating a procedure, composing a response, returns caused by incomplete answers and escalations to an expert. Quality includes alignment with the approved source, coverage of the question and clear identification of situations that need escalation. The operational measure tracks time to a draft ready for review and change in expert workload.

Supervision cost includes reviewing every draft during the pilot, maintaining the set of approved sources and resolving enquiries outside scope. The stop criterion applies when the source cannot be stated or when a draft creates an incorrect direction for action that review cannot readily identify. This framework gives the team a basis for a decision without inventing an expected return.

A first pilot does not establish the value of AI for the entire organisation. It establishes the value of one clearly described process under specific conditions. A change in scope, data source, users or level of autonomy can change the result.

Measurement also has a cost. An overly broad indicator set slows work and makes interpretation harder. Too small a set hides risk. A practical starting point is one quality measure, one operational measure, recorded supervision cost and a clear stop criterion. After the first cycle, the team can add details that prove important.

Processes involving sensitive decisions, weak sources or frequent exceptions often require a narrower scope and stronger review. In such cases, success is not full automation. Success may be a reliable draft, better case preparation or faster retrieval of a verified source while employee accountability remains in place.

Before approving a pilot, prepare a one-page document containing the process name, scope, baseline, owner, data source, quality measure, operational measure, supervision cost, review method and continuation or stop criterion. If these elements cannot be described clearly, the scope is probably too broad or the process is not ready for a pilot.

For processes that touch ERP, documents and established work roles, an AI upgrade for business processes makes sense only after this assessment. A team that wants to test scope and measures for a specific process can Talk to the ORKA team .

Recommended articles