Simor
The AI project scoping template: right-size before you build

The AI project scoping template: right-size before you build

Simor Consulting | 04 Oct, 2026 | 04 Mins read

AI projects have a scoping problem. Teams either scope too loosely, “use AI to improve customer experience”, or too tightly: “build a transformer model with 12 attention layers for intent classification.” The loose scope produces projects that drift for months without delivering value. The tight scope locks the team into architectural decisions before they understand the problem.

A good scope sits between these extremes. It names the problem, constrains the solution space, defines success in measurable terms, and identifies the decisions that need to be deferred until more information is available. This post gives you a template for producing that scope.

The Five Sections of an AI Project Scope

Section 1: Problem Definition

State the problem in terms a non-technical stakeholder can evaluate. Not “build a model that classifies support tickets” but “reduce average ticket resolution time from 48 hours to 12 hours by automatically routing tickets to the correct team.”

The problem definition should include three elements: the current state (what happens today), the desired state (what should happen after the project), and the measurement (how you will know the desired state has been reached). If you cannot state all three, the problem is not well-defined enough to scope an AI project.

Common failure mode: defining the problem as a technology choice. “We need a RAG system for our documentation” is a solution, not a problem. “Our support engineers spend 30 minutes per issue searching for relevant documentation” is a problem. The solution might be RAG. It might be better search. It might be better documentation. The problem definition should not constrain the solution.

Section 2: Data Assessment

Before choosing a model, assess the data. This section answers four questions:

What data exists? List the data sources relevant to the problem. Include structured data (databases, spreadsheets), unstructured data (documents, emails, chat logs), and metadata (timestamps, user IDs, categories). Be specific about volume and format.

What is the data quality? Sample the data and assess completeness, accuracy, consistency, and freshness. A ten-minute data quality assessment, randomly sample 100 records and check for obvious issues, is worth more than a week of model development. If the data is poor, no model will produce good results.

What is labelled? For supervised learning tasks, labelled data is the binding constraint. How many labelled examples exist? Who labelled them? What is the inter-annotator agreement? If labelling has not started, add the labelling effort to the project timeline. Labelling is not a quick step. It is often the longest phase of the project.

What are the access restrictions? Can the team access the data directly, or does it require approvals? Is the data subject to privacy regulations? Can it be used for model training? Access restrictions discovered mid-project cause delays that no amount of engineering skill can overcome.

Section 3: Success Criteria

Define success at three levels: minimum viable, target, and aspirational.

Minimum viable: the threshold below which the project is not worth deploying. If the model cannot route tickets with at least 70 percent accuracy, a rule-based system would be simpler and cheaper. The minimum viable threshold should be above the performance of the naive alternative.

Target: the performance level that justifies the investment. At 85 percent routing accuracy, the team estimates a 60 percent reduction in resolution time. This is the level the project should plan around.

Aspirational: the performance level that would expand the project’s impact. At 95 percent accuracy, the system could handle ticket resolution automatically for common issue types. This level drives stretch goals but should not be in the base plan.

Define each level in business terms, not model terms. Not “F1 score of 0.85” but “85 percent of tickets routed to the correct team on first submission.” Map business metrics to model metrics, but lead with business metrics in the scope document.

Section 4: Build vs. Buy Decision

For each component of the solution, state whether you plan to build, buy, or use an existing service. The components to consider:

  • Data ingestion and preparation
  • Model training or fine-tuning
  • Model serving and inference
  • Evaluation and monitoring
  • User interface and integration

For each component, the decision should be driven by whether the component is a differentiator. If ticket routing accuracy is the differentiator, invest in custom model development. If the user interface is not a differentiator, use an existing framework. If model serving is a commodity, use a managed service.

Do not build infrastructure that does not differentiate. Every hour spent on undifferentiated infrastructure is an hour not spent on the problem that justifies the project.

Section 5: Timeline and Checkpoints

Break the project into phases with checkpoints where the team evaluates progress against the success criteria. A reasonable structure:

Phase 1 (2-4 weeks): Data assessment and baseline. Establish the current performance of the naive approach. If the naive approach already meets the target success criteria, the project scope needs revision.

Phase 2 (4-8 weeks): Model development and evaluation. Build the model, train it, evaluate it against the success criteria. This phase ends with a go/no-go decision: does the model meet minimum viable criteria?

Phase 3 (4-6 weeks): Integration and deployment. Build the production pipeline, integrate with existing systems, and deploy. This phase should not start until Phase 2 confirms the model meets minimum viable criteria.

The checkpoints are not status meetings. They are decision points. At each checkpoint, the team presents evidence against the success criteria, and the stakeholders decide whether to continue, pivot, or stop. A project that stops at a checkpoint is not a failure. It is a project that saved the remaining budget by learning early.

The Scope Anti-Patterns

Scope by analogy. “Company X built a similar system, so we should be able to.” Company X had different data, different infrastructure, and different requirements. Their success or failure tells you nothing about yours.

Scope by technology. “We should use GPT-4 for this.” The technology choice should follow the problem definition and data assessment, not precede them. If you have already chosen the technology, you have biased the scope.

Scope by optimism. “We think we can do this in six weeks.” The estimate assumes everything goes right. It does not account for data quality issues, integration complexity, or the inevitable scope additions that stakeholders request after seeing the first demo. Multiply optimistic estimates by 1.5x for a realistic plan.

Next Step

Fill in the five sections for your current AI project this week. If you cannot complete Section 2 (Data Assessment) because you have not examined the data, that is your answer: the project is not scoped yet. Examine the data first, then complete the template.

Shipping a production AI system?

Find where your AI spend leaks and where quality slips. Take the AI Production Scorecard for a fast baseline across the seven layers, or book a free AI cost review and we will turn it into a plan.

Similar Articles

5 AI Workflows Professional Services Firms Can Deploy This Quarter
5 AI Workflows Professional Services Firms Can Deploy This Quarter
10 Jul, 2026 | 12 Mins read

Professional services firms sell judgment, billed by the hour or by the matter. That makes them both the biggest winners and the most cautious adopters of AI. The upside is real: every firm carries ho

Legacy Data Pipeline Modernisation Without Rewriting Everything
Legacy Data Pipeline Modernisation Without Rewriting Everything
10 Jul, 2026 | 10 Mins read

The pipeline runs every night at 2 a.m. Nobody fully understands it. The original author left in 2019. It is part SAS, part shell, part stored procedures, and part a spreadsheet someone emails in. It

Lightweight MLOps for Mid-Market Teams: Ship Models Without a Platform Engineering Org
Lightweight MLOps for Mid-Market Teams: Ship Models Without a Platform Engineering Org
10 Jul, 2026 | 11 Mins read

A head of ML at a 120-person company told us recently that his team had spent nine months trying to stand up a "proper MLOps platform." They had evaluated three orchestration tools, designed a feature

AI in the Software Development Lifecycle: From Code Review to Deployment
AI in the Software Development Lifecycle: From Code Review to Deployment
27 Jul, 2026 | 22 Mins read

Code completion gets the attention, but it is the narrowest part of what AI can do in a development workflow. Walk into any team that has shipped software for a few years and they will tell you: writi

Building AI Dashboards: Visualising AI System Performance for Executives
Building AI Dashboards: Visualising AI System Performance for Executives
20 Sep, 2026 | 14 Mins read

An executive looking at an AI dashboard does not want to see loss curves. They want to know if the AI system is doing its job, whether it is trustworthy, and what happens when it is not. Translating A

Anatomy of an AI Incident: Post-Mortem of a Model Provider Outage
Anatomy of an AI Incident: Post-Mortem of a Model Provider Outage
19 Jun, 2026 | 09 Mins read

On a Tuesday at 2:14 PM, a major model provider began returning elevated error rates for a specific model endpoint. By 2:31 PM, a customer support platform that depended on that endpoint was producing

AI Rollback Patterns: When to Roll Back a Prompt, a Model, or the Whole Release
AI Rollback Patterns: When to Roll Back a Prompt, a Model, or the Whole Release
27 Jun, 2026 | 11 Mins read

Software rollbacks are well-understood. You deploy a new version, detect an issue, and roll back to the previous version. The rollback is atomic: the entire application reverts to the previous state.

The 7-step vector database selection checklist
The 7-step vector database selection checklist
26 Apr, 2026 | 06 Mins read

Most vector database selection failures come down to one mistake: picking the technology before mapping the workload. Teams benchmark embedding search speed on a curated dataset, pick the fastest opti

Build vs buy: a decision tree for AI infrastructure
Build vs buy: a decision tree for AI infrastructure
03 May, 2026 | 06 Mins read

Every AI infrastructure team eventually faces the same argument. One faction wants to build a custom solution because the commercial options do not handle their specific requirements. The other factio

How to design a prompt ops pipeline from scratch
How to design a prompt ops pipeline from scratch
10 May, 2026 | 06 Mins read

Prompt management in most AI teams starts the same way. One engineer writes a prompt, it works well enough, and the prompt gets committed to a config file. Three months later, there are forty prompts

The data quality scorecard: metrics that actually matter
The data quality scorecard: metrics that actually matter
17 May, 2026 | 06 Mins read

Most data quality initiatives fail not because teams lack tools, but because they measure the wrong things. Teams track hundreds of data quality metrics, generate dashboards full of green indicators,

A cost optimisation framework for LLM inference
A cost optimisation framework for LLM inference
24 May, 2026 | 06 Mins read

LLM inference costs follow a pattern that catches teams off guard. The first prototype costs almost nothing: a few hundred dollars a month during development. The pilot scales to a few thousand. Produ

Migration playbook: batch to streaming in 5 phases
Migration playbook: batch to streaming in 5 phases
31 May, 2026 | 06 Mins read

The case for streaming is straightforward: data that arrives in minutes instead of hours enables decisions that were previously impossible. Fraud detection catches transactions before they clear. Pers

How to audit your AI pipeline for bias: step by step
How to audit your AI pipeline for bias: step by step
07 Jun, 2026 | 06 Mins read

Bias in AI systems is not a theoretical risk. It is a measurable property that can be detected, quantified, and mitigated at every stage of the pipeline. The teams that treat bias as an audit problem

The 30-day AI readiness assessment
The 30-day AI readiness assessment
14 Jun, 2026 | 07 Mins read

Organisations that skip readiness assessment before investing in AI tend to discover their gaps expensively. A financial services firm spent four months building a customer churn prediction model only

Your first 90 days as a Head of AI Engineering
Your first 90 days as a Head of AI Engineering
28 Jun, 2026 | 07 Mins read

The first Head of AI Engineering at a company inherits one of three situations. Situation one: there is no AI team, no AI infrastructure, and the mandate is to build from scratch. Situation two: there

The RAG evaluation framework you'll actually use
The RAG evaluation framework you'll actually use
08 Jul, 2026 | 06 Mins read

Most RAG systems are evaluated with vibes. An engineer runs ten queries, eyeballs the results, and declares the system "working." Three months later, a customer reports that the system confidently ret

How to write an AI incident response plan
How to write an AI incident response plan
12 Jul, 2026 | 07 Mins read

AI systems fail differently than traditional software. A traditional software bug produces incorrect output deterministically. The same input always produces the same wrong output, and a fix eliminate

Capacity planning for vector databases
Capacity planning for vector databases
19 Jul, 2026 | 07 Mins read

Vector database capacity planning fails in predictable ways. Teams estimate storage based on vector count alone and discover at 60% capacity that memory consumption is growing faster than disk because

The procurement checklist for AI vendors
The procurement checklist for AI vendors
26 Jul, 2026 | 07 Mins read

AI vendor procurement is where organisations make binding commitments that are expensive to unwind. A three-year contract with a model provider locks you into their pricing, their rate limits, their m

Setting up a model registry: the minimal viable approach
Setting up a model registry: the minimal viable approach
02 Aug, 2026 | 06 Mins read

A model registry is the version control system for your trained models. Without one, teams track model versions by filename, store artifacts in ad-hoc cloud storage locations, and discover which model

Data contract template and negotiation guide
Data contract template and negotiation guide
09 Aug, 2026 | 07 Mins read

Data pipelines break because data producers and data consumers have different assumptions. The producer assumes the consumer can handle null values in a column. The consumer assumes the column is neve

How to run an AI architecture review
How to run an AI architecture review
12 Aug, 2026 | 07 Mins read

An architecture review for an AI system catches design flaws at the cheapest possible stage: before implementation. A data pipeline that cannot handle the expected volume, a model serving architecture

The observability maturity model for AI systems
The observability maturity model for AI systems
16 Aug, 2026 | 07 Mins read

Most AI systems in production operate with observability that was designed for traditional software. Teams monitor CPU, memory, network, and error rates. These metrics tell you whether the server is r

Building an internal AI platform team: org chart and responsibilities
Building an internal AI platform team: org chart and responsibilities
23 Aug, 2026 | 07 Mins read

The decision to centralise AI infrastructure into a platform team usually comes after a period of decentralised pain. Three product teams independently built model serving pipelines. None of them shar

LLM cost calculator: estimating spend before you deploy
LLM cost calculator: estimating spend before you deploy
30 Aug, 2026 | 05 Mins read

Teams approve LLM projects based on per-query cost estimates, then get blindsided by the actual invoice. The gap between estimate and reality is not a rounding error. It is a structural problem: the e

Designing a data mesh operating model: roles, responsibilities, and boundaries
Designing a data mesh operating model: roles, responsibilities, and boundaries
06 Sep, 2026 | 04 Mins read

Most data mesh initiatives fail not because the architecture is wrong, but because nobody can answer the question: who owns this data product? When ownership is ambiguous, quality drops, SLAs go unmet

The LLM cost optimisation playbook: 12 techniques that actually save money
The LLM cost optimisation playbook: 12 techniques that actually save money
13 Sep, 2026 | 04 Mins read

LLM costs are easy to start and hard to control. A team ships a feature that calls GPT-4, the feature works, users like it, and the invoice climbs 15 percent month over month. The cost is not a proble

How to run a pre-mortem on your AI project
How to run a pre-mortem on your AI project
23 Sep, 2026 | 04 Mins read

Post-mortems are useful. Pre-mortems are cheaper. A post-mortem tells you why a project failed after the money is gone. A pre-mortem tells you why a project might fail while you can still change cours

Data pipeline testing strategy: unit, integration, and contract tests
Data pipeline testing strategy: unit, integration, and contract tests
27 Sep, 2026 | 05 Mins read

Data pipelines break in production more often than they should, and the breakage is expensive. A pipeline that silently produces wrong data for three days before anyone notices has corrupted downstrea

AI Enablement Programs: Building Organisational Capability, Not Just Technology
AI Enablement Programs: Building Organisational Capability, Not Just Technology
19 Mar, 2026 | 11 Mins read

A technology company built an impressive AI platform. They had GPU clusters, fine-tuning pipelines, evaluation frameworks, and a growing model registry. They opened access to any team that wanted to u

Building an AI Centre of Excellence: Structure, Mandate, and Success Metrics
Building an AI Centre of Excellence: Structure, Mandate, and Success Metrics
05 Jul, 2026 | 11 Mins read

Most organisations have attempted some form of AI initiative. Some succeeded and delivered measurable business value. Many failed and produced results that were technically interesting but did not mov

Prompt Engineering as Infrastructure: Version Control, Testing, and Deployment
Prompt Engineering as Infrastructure: Version Control, Testing, and Deployment
22 May, 2026 | 11 Mins read

Prompts are not prompts in the casual sense of suggestions or starting points. They are software. They take inputs, produce outputs, have failure modes that manifest in specific conditions, and requir

Why Small Businesses Need AI Now: A 2026 Practitioner's Guide
Why Small Businesses Need AI Now: A 2026 Practitioner's Guide
10 Jul, 2026 | 11 Mins read

If you run a small business, you have heard the AI pitch a hundred times. Most of it is aimed at enterprises with data teams, seven-figure budgets, and a CIO to translate. That framing is now out of d