AI projects have a scoping problem. Teams either scope too loosely, “use AI to improve customer experience”, or too tightly: “build a transformer model with 12 attention layers for intent classification.” The loose scope produces projects that drift for months without delivering value. The tight scope locks the team into architectural decisions before they understand the problem.
A good scope sits between these extremes. It names the problem, constrains the solution space, defines success in measurable terms, and identifies the decisions that need to be deferred until more information is available. This post gives you a template for producing that scope.
The Five Sections of an AI Project Scope
Section 1: Problem Definition
State the problem in terms a non-technical stakeholder can evaluate. Not “build a model that classifies support tickets” but “reduce average ticket resolution time from 48 hours to 12 hours by automatically routing tickets to the correct team.”
The problem definition should include three elements: the current state (what happens today), the desired state (what should happen after the project), and the measurement (how you will know the desired state has been reached). If you cannot state all three, the problem is not well-defined enough to scope an AI project.
Common failure mode: defining the problem as a technology choice. “We need a RAG system for our documentation” is a solution, not a problem. “Our support engineers spend 30 minutes per issue searching for relevant documentation” is a problem. The solution might be RAG. It might be better search. It might be better documentation. The problem definition should not constrain the solution.
Section 2: Data Assessment
Before choosing a model, assess the data. This section answers four questions:
What data exists? List the data sources relevant to the problem. Include structured data (databases, spreadsheets), unstructured data (documents, emails, chat logs), and metadata (timestamps, user IDs, categories). Be specific about volume and format.
What is the data quality? Sample the data and assess completeness, accuracy, consistency, and freshness. A ten-minute data quality assessment, randomly sample 100 records and check for obvious issues, is worth more than a week of model development. If the data is poor, no model will produce good results.
What is labelled? For supervised learning tasks, labelled data is the binding constraint. How many labelled examples exist? Who labelled them? What is the inter-annotator agreement? If labelling has not started, add the labelling effort to the project timeline. Labelling is not a quick step. It is often the longest phase of the project.
What are the access restrictions? Can the team access the data directly, or does it require approvals? Is the data subject to privacy regulations? Can it be used for model training? Access restrictions discovered mid-project cause delays that no amount of engineering skill can overcome.
Section 3: Success Criteria
Define success at three levels: minimum viable, target, and aspirational.
Minimum viable: the threshold below which the project is not worth deploying. If the model cannot route tickets with at least 70 percent accuracy, a rule-based system would be simpler and cheaper. The minimum viable threshold should be above the performance of the naive alternative.
Target: the performance level that justifies the investment. At 85 percent routing accuracy, the team estimates a 60 percent reduction in resolution time. This is the level the project should plan around.
Aspirational: the performance level that would expand the project’s impact. At 95 percent accuracy, the system could handle ticket resolution automatically for common issue types. This level drives stretch goals but should not be in the base plan.
Define each level in business terms, not model terms. Not “F1 score of 0.85” but “85 percent of tickets routed to the correct team on first submission.” Map business metrics to model metrics, but lead with business metrics in the scope document.
Section 4: Build vs. Buy Decision
For each component of the solution, state whether you plan to build, buy, or use an existing service. The components to consider:
- Data ingestion and preparation
- Model training or fine-tuning
- Model serving and inference
- Evaluation and monitoring
- User interface and integration
For each component, the decision should be driven by whether the component is a differentiator. If ticket routing accuracy is the differentiator, invest in custom model development. If the user interface is not a differentiator, use an existing framework. If model serving is a commodity, use a managed service.
Do not build infrastructure that does not differentiate. Every hour spent on undifferentiated infrastructure is an hour not spent on the problem that justifies the project.
Section 5: Timeline and Checkpoints
Break the project into phases with checkpoints where the team evaluates progress against the success criteria. A reasonable structure:
Phase 1 (2-4 weeks): Data assessment and baseline. Establish the current performance of the naive approach. If the naive approach already meets the target success criteria, the project scope needs revision.
Phase 2 (4-8 weeks): Model development and evaluation. Build the model, train it, evaluate it against the success criteria. This phase ends with a go/no-go decision: does the model meet minimum viable criteria?
Phase 3 (4-6 weeks): Integration and deployment. Build the production pipeline, integrate with existing systems, and deploy. This phase should not start until Phase 2 confirms the model meets minimum viable criteria.
The checkpoints are not status meetings. They are decision points. At each checkpoint, the team presents evidence against the success criteria, and the stakeholders decide whether to continue, pivot, or stop. A project that stops at a checkpoint is not a failure. It is a project that saved the remaining budget by learning early.
The Scope Anti-Patterns
Scope by analogy. “Company X built a similar system, so we should be able to.” Company X had different data, different infrastructure, and different requirements. Their success or failure tells you nothing about yours.
Scope by technology. “We should use GPT-4 for this.” The technology choice should follow the problem definition and data assessment, not precede them. If you have already chosen the technology, you have biased the scope.
Scope by optimism. “We think we can do this in six weeks.” The estimate assumes everything goes right. It does not account for data quality issues, integration complexity, or the inevitable scope additions that stakeholders request after seeing the first demo. Multiply optimistic estimates by 1.5x for a realistic plan.
Next Step
Fill in the five sections for your current AI project this week. If you cannot complete Section 2 (Data Assessment) because you have not examined the data, that is your answer: the project is not scoped yet. Examine the data first, then complete the template.