Skip to main content
Back to blog
Blog

Governing AI Pilots: Criteria for Stopping, Adapting and Scaling

An AI pilot needs more than usage figures. Make use cases, quality, human checkpoints and scaling rules testable before the pilot begins.

Andrea Giugliano
Governing AI Pilots: Criteria for Stopping, Adapting and Scaling

An AI pilot should not measure enthusiasm or usage. It should enable a decision.

Many AI pilots are evaluated using the wrong metrics or questions: How many people used the tool? How many licences were allocated? How many users work with AI? More important is whether clearly bounded use cases create value and produce better results under realistic conditions, and whether the organisation can assure the quality, risks and accountability associated with those results.

Licence activity can show that access exists. It demonstrates neither time savings nor better quality. Nor does it reveal how much correction work is required, which tasks do not suit the system or whether the solution would be viable in day-to-day operations.

When a pilot is organised only as a tool test, the outcome is usually a collection of impressions. When it is designed as an evaluation and governance system, it supports a reasoned decision: stop, adapt, continue on a limited basis or scale.

Usage is a signal, not an outcome

Generative AI can accelerate work. Its effects are neither uniform nor automatically transferable to entire workflows.

A randomised field experiment involving 7,137 knowledge workers across 66 companies examined an AI assistant integrated into email, meetings and document work. During the second half of the six-month experiment, active users spent less time on email. At the same time, the researchers found no corresponding change in the quantity or composition of tasks. Individual time savings do not automatically change coordination, decisions or the operating model. The study is an NBER Working Paper; several authors were employed by Microsoft at the time of the research. That limitation belongs in the interpretation. Dillon et al., NBER Working Paper 33795

Task fit varies as well. An experimental study published in Organization Science in 2026 describes a jagged technological frontier. For tasks within the capabilities of the system used in the experiment, AI assistance improved productivity and quality. For a task outside that frontier, it could reduce performance. Tasks that appear similar can therefore produce opposite results. Dell’Acqua et al., Organization Science

The implication for a pilot is that averages across all tasks conceal precisely the differences that matter for a scaling decision. A better approach is to assess the different use cases against different criteria and determine where AI adds value and where its use is not justified.

A pilot needs clear evaluation questions

Suitable questions include:

  • Can the system accelerate a clearly defined task without falling below the agreed quality threshold?

  • Which tasks and roles are suitable for support, and which remain excluded?

  • How much expert review and correction is required?

  • Which data, permissions and control points are necessary for safe operation?

  • Are the benefits and organisational conditions strong enough for limited scaling?

The question must be narrow enough to produce an answer. “How can we use AI in our company?” is too broad. “Can an assistant support the first draft of a standardised internal summary if an accountable expert reviews every statement?” is decidable.

The question also needs explicit non-goals. A limited text-assistance pilot does not automatically test suitability for employment decisions, customer communication, sensitive data or autonomous process steps.

Break the workflow into concrete tasks

Most knowledge-work processes are not a single activity. Research, structuring, drafting, review, decision and approval impose different requirements.

A pilot should therefore describe the tasks being tested individually rather than relying on a broad process label:

Field

Concrete question

Task

Which clearly bounded work step will the system support?

Input

Which information and data classes are required?

Output

What result should be produced?

Quality threshold

Which errors or deviations are unacceptable?

Human checkpoint

Who reviews what, and who may reject the result?

Exclusion

Which tasks, data or decisions are outside the pilot?

The NIST AI Risk Management Framework similarly recommends documenting specific tasks, knowledge limits, expected benefits and costs, human oversight and comparative benchmarks. It is a voluntary US framework, not a source of European law. It is nevertheless useful as a structure for testing, evaluation, verification and validation. NIST AI RMF Core

Define the baseline before the first test

Without a baseline, the impact of the change, with or without AI, cannot be measured. A baseline should therefore be collected before the pilot begins.

The comparison does not need to be complicated, but it must fit the task. A limited pilot may need only a few traceable signals:

  • processing time for comparable tasks

  • expert quality based on a checklist defined in advance

  • the number and severity of required corrections

  • clarification loops, rework and handover errors

  • unauthorised data use or other risk signals

  • perceived relief as supporting evidence, not as the only evidence

The distinction between gross and net time matters. A draft produced ten minutes faster is not a time saving if it creates twenty minutes of additional review.

Quality must not be defined only after the pilot either. Decide in advance which criteria are professionally decisive. These may include completeness, accuracy, traceability, source use, tone or compliance with a required format.

Human review needs accountability, time and authority

A box labelled “human in the loop” is not an effective control.

The reviewer must:

  • understand the task and the plausible failure modes

  • have access to the evidence needed for comparison

  • receive sufficient time to review the output

  • be able to correct, reject or escalate the result

  • know who is accountable for the final decision when an error occurs

This is particularly important when outputs affect people, involve sensitive information or prepare material business decisions.

The European Commission’s current Q&A on AI literacy links appropriate measures to factors including technical knowledge, experience, education and training, and the context in which the system is used. Additional training and human-oversight requirements remain relevant for the deployment of certain high-risk systems. The obligations that apply in a specific case need a separate assessment. A blog article is not a substitute for that review. European Commission: AI Literacy – Questions & Answers

In practice, enablement belongs inside the pilot. A generic introduction is not enough. Users and reviewers need to understand the limits of their specific use case and practise with realistic examples.

Governance must be decidable before use

Governance does not mean sending every prompt to a central committee. It means establishing the necessary boundaries and accountabilities before use begins.

At a minimum, clarify the following:

  1. Permitted data: Which data classes may be processed, and which are explicitly prohibited?

  2. Permissions: Which systems and information can the solution access directly or indirectly?

  3. Provider terms: How are inputs, outputs and logs processed, stored or used for improvement?

  4. Rights: Which copyright, usage-right and confidentiality questions arise?

  5. Accountability: Who owns the task, system use, professional quality and risk?

  6. Escalation: What happens after an error, a data-protection incident or unexpected behaviour?

  7. Employee participation: Which employment-law or co-determination checks are required?

The depth of review should reflect the use case and its risk. An internal drafting aid using non-sensitive data requires different controls from a system that prepares decisions about people.

For the wider organisational context, see our focus area on digitalisation and AI use cases. Where software allocates tasks, sets priorities or evaluates performance, “When AI Directs Work” extends the discussion to decision rights and work design.

Measure by use case, not by individual

A pilot should evaluate tasks and systems. It should not create a hidden employee performance ranking.

Evaluation should therefore be organised by use case, role and context. Person-level prompt logs or individual productivity rankings are usually neither necessary for the decision nor sound organisational practice. Where personal data cannot be avoided, the specific legal, data-protection and employee-participation implications require review.

A compact pilot scorecard might look like this:

Dimension

Baseline

Pilot signal

Decision rule

Time

Current effort for comparable tasks

Net time including review and correction

Positive only if rework does not consume the initial gain

Quality

Professional criteria defined in advance

Errors, completeness and returned work

No scaling below the minimum quality threshold

Risk

Known failure scenarios and permitted data

Incidents, near misses and policy breaches

A critical incident triggers a stop and review

Capability

Existing role knowledge

Safe use and correct escalation

Continue only with sufficient role-specific capability

Viability

Current costs and process ownership

Licence, operating, support and control effort

Benefits must justify the full operating cost

Usage frequency may explain why evidence is limited. It is not a substitute for these dimensions.

Set stop, adapt and scale rules

A pilot without a stop rule develops a life of its own. A pilot without a scaling rule becomes an endless loop.

Describe four possible decisions before the pilot starts:

  • Stop: The benefit is not demonstrated, the quality threshold is missed or the risk is unacceptable.

  • Adapt: Change the use case, data access, control point, training or technical configuration, then test the change deliberately.

  • Continue on a limited basis: The evidence is positive but insufficient for broader use.

  • Scale: Benefits, quality, risk, capability and operational accountability support expansion within the defined scope.

Scaling can apply only to the tasks and conditions that were actually tested. A successful pilot for internal drafting does not authorise external claims, sensitive data or decisions about people.

A change of model or provider may also require reassessment. The capability boundary of generative systems changes. A one-off approval is therefore not permanent evidence.

A good pilot makes limits visible

A robust result does not necessarily have to show that AI should be scaled.

A pilot is also valuable when it demonstrates:

  • that the selected use case is a poor fit;

  • that data or permissions remain unresolved;

  • that correction effort outweighs the time saving;

  • that the organisation lacks a viable review or support structure;

  • that a smaller, more clearly bounded use case would be more sensible.

The important outcome is not a positive result. It is a traceable decision.

Conclusion

Before the pilot begins, the task, baseline, quality threshold, data boundaries, human review and decision rules must be established. Only then can an organisation distinguish between a tool being used and a use case being genuinely viable.

If these questions are addressed only after the pilot, the organisation can document activity. If they are addressed beforehand, it can make a responsible decision.

Would you like to structure an AI pilot so that benefits, risks and scaling conditions can be tested? A focused workshop can define the decision question, use cases, governance gates and evaluation plan. Book a non-binding introductory call.

Sources and limitations

  1. Dillon et al.: Shifting Work Patterns with Generative AI. Randomised experiment across 66 companies and 7,137 knowledge workers; an NBER Working Paper, not a general proof of impact. The authors disclose the relevant vendor affiliations.

  2. Dell’Acqua et al.: Navigating the Jagged Technological Frontier. Peer-reviewed experimental study with management consultants; the task, model and occupational context limit generalisation.

  3. NIST AI Risk Management Framework Core. A voluntary US framework for structuring risk, testing and evaluation work; it is not a source of European law.

  4. European Commission: AI Literacy – Questions & Answers. Official guidance on the current position; the legal assessment of a specific system remains context-dependent.