AI Production Scorecard · Method and evidence
A worked production scorecard
Follow one fictional system from questionnaire answers to review priorities. See the calculation, the findings, and what the numbers cannot tell you.
Worked example using fictional inputs. This is not a client assessment. Scores summarize self-reported controls; they do not certify security, compliance, or production readiness.
The fictional situation
A fictional Head of Engineering at a 600-person B2B software company. The production, customer-facing, multi-tenant support assistant handles customer PII and uses three tools or APIs, not an MCP server. The priority is reliability following recurring poor answers in production. No frequency, losses, client outcomes, or measured improvement is claimed.
This is a complete example with no missing answers. It is not evidence of a typical customer or a completed architecture review.
Overall result
45/100
Band: piloting. “Piloting” is the rubric’s control-maturity label, not the deployment stage: this fictional system is in production.
These are rubric scores, not percentages of safety, industry percentiles, or statistically calibrated estimates. They summarize self-reported controls, not verified enforcement.
Ranked review priorities
- Evals — 13/100
- Tool governance — 25/100
- Budget governance — 38/100
The rubric suggests these priorities from the three lowest scores; they are not proven incident causes. Exposure and evidence can change their order during a review. Evals and Tool governance score below 35. A strong average cannot cancel a consequential gap.
Seven layer findings and good-state targets
The findings below use only the selected answers. The good-state descriptions are proposed practice standards, not measured benchmarks. They do not change scoring, establish an audit result, or mandate a vendor or centralized architecture.
Model control — 75/100
Questionnaire finding. Most traffic uses a central model layer. Model swaps are controlled and tested, but the answers do not establish complete traffic coverage or fast, reversible changes across all providers and regions.
What good looks like. Model access follows an auditable policy. Changes are versioned, evaluated, and reversible. Provider-held continuation state retains affinity unless portability is established; routing alone does not make failover safe.
Prompt operations — 50/100
Questionnaire finding. A registry preserves prompt history, but release checks are manual spot checks.
What good looks like. Version prompts and settings together. Review changes, evaluate representative cases before release, and retain a tested rollback path.
Guardrails — 50/100
Questionnaire finding. Important flows have basic checks. Higher-risk tools have pre-use validation, but comprehensive enforcement is not established.
What good looks like. Validate inputs, outputs, and proposed actions at trusted boundaries. Enforce authorization outside model instructions. Guardrails are neither a sandbox nor proof of hostile-process containment.
Budget governance — 38/100
Questionnaire finding. Provider or cloud alerts and detection within hours do not establish automatic containment.
What good looks like. Enforce operation, tenant, and tool limits with bounded retries and loops. Alert an accountable owner and stop or degrade work through tested containment paths.
Tool governance — 25/100
Questionnaire finding. Secrets are centrally stored but broadly accessible. Tool permissions are documented informally.
What good looks like. Use scoped credentials, rotation, explicit permissions, and approvals for consequential actions. Verify enforcement with permission tests and audit records. A secret store alone does not establish least privilege.
Observability — 63/100
Questionnaire finding. Some requests are traceable manually. Errors, latency, tool health, and model metrics are visible together, but complete reconstruction is not established.
What good looks like. Correlate model and tool steps with versions, decisions, and incident handling. Redact sensitive data and define access and retention. Replay must not repeat external side effects. Local retention policy does not establish provider retention.
Evals — 13/100
Questionnaire finding. Pre-release checks are manual spot checks. User complaints are the main post-release signal.
What good looks like. Gate changes with representative component and system evaluations. Review production failures and add regression cases. Monitor post-release quality and drift. Passing a finite suite does not prove universal correctness.
Every selected answer
Context — not scored
- What stage is this AI system in?
- Customer-facing production (option index 4 ; context only )
- Who uses the system today?
- Multi-tenant production system (option index 4 ; context only )
- Does the system use tools, APIs, or MCP servers?
- 3 to 5 tools or APIs (option index 2 ; context only )
- How sensitive is the data handled by this system?
- PII, financial, or contractual data (option index 3 ; context only )
- What do you most want to improve right now?
- Reliability (option index 1 ; context only )
Control answers — scored
- How is model access handled today?
- Most traffic goes through a central model layer (option index 3 ; 75 points )
- How quickly can you swap a model, provider, or region when something changes?
- Model swaps are controlled and tested before rollout (option index 3 ; 75 points )
- How are prompts and model settings managed?
- We use a prompt registry with version history (option index 3 ; 75 points )
- How are prompt changes tested before release?
- We do manual spot checks only (option index 1 ; 25 points )
- What input and output guardrails are in place?
- Basic checks on some important flows (option index 2 ; 50 points )
- How do you control risky or non-compliant tool actions?
- Higher-risk tools have validation before use (option index 2 ; 50 points )
- How do you control spend and runaway behavior?
- We rely on provider or cloud alerts (option index 1 ; 25 points )
- How quickly would you know if an agent loop or tool bug started burning money?
- Alerts would catch it within hours (option index 2 ; 50 points )
- How are tool and MCP credentials managed?
- Secrets are stored centrally, but access is broad (option index 1 ; 25 points )
- How well do you know what each tool is allowed to do?
- Tool access is documented informally (option index 1 ; 25 points )
- If a user reports a bad result, how easily can you trace it?
- We can trace some requests manually (option index 2 ; 50 points )
- What do you monitor in production?
- Errors, latency, tool health, and model metrics in one place (option index 3 ; 75 points )
- How do you evaluate quality before changes go live?
- We do manual spot checks (option index 1 ; 25 points )
- How do you catch regressions after release?
- User complaints are the main signal (option index 0 ; 0 points )
See the methodology and worked calculation, then read the fictional written review summary.