AI Production Scorecard · Method and evidence
How the seven-layer scorecard works
A self-assessment for a review conversation—not a certification, client proof, or prediction of incident probability.
Inputs, weights, and rounding
Five context questions describe deployment stage, users, tools, data, and the improvement priority. They inform the conversation but do not affect the score. Fourteen scored questions cover seven layers, with two answers per layer.
Each selected option contributes its zero-based index multiplied by 25: indexes 0, 1, 2, 3, and 4 contribute 0, 25, 50, 75, and 100 points. Each layer is the rounded mean of its two answers. The overall is the rounded mean of the seven already-rounded layer scores. Half points round up. All layers have equal weight.
Complete every question before interpreting a score. An omitted answer is not evidence that a control is absent; a partial input is not a valid completed assessment.
Score bands
- 0–34: fragile (Fragile)
- 35–54: piloting (Piloting)
- 55–69: shipping (Shipping)
- 70–84: controlled (Controlled)
- 85–100: hardened (Hardened)
Band names describe rubric maturity, not actual deployment stage. The three lowest-scoring layers suggest review priorities. Ties retain rubric order. The numbers are not safety percentages, industry percentiles, or statistical calibration.
The worked calculation
The complete fictional example uses these point pairs. Its context answers contribute no points.
- Model control: (75 + 75) / 2 = 75 → 75
- Prompt operations: (75 + 25) / 2 = 50 → 50
- Guardrails: (50 + 50) / 2 = 50 → 50
- Budget governance: (25 + 50) / 2 = 37.5 → 38
- Tool governance: (25 + 25) / 2 = 25 → 25
- Observability: (50 + 75) / 2 = 62.5 → 63
- Evals: (25 + 0) / 2 = 12.5 → 13
Overall: (75 + 50 + 50 + 38 + 25 + 63 + 13) / 7 = 314 / 7 ≈ 44.86 → 45, piloting. Rounding only at the very end would be a different calculation.
What each layer covers
The good-state descriptions are proposed practice standards, not measured benchmarks. They do not mandate a particular vendor or centralized architecture. Artifacts below are examples a reviewer might inspect; none was supplied or inspected for this fictional company.
Model control
Coverage and purpose. Model access, routing, and change control make provider and model changes auditable and reversible.
What good looks like. Model access follows an auditable policy. Changes are versioned, evaluated, and reversible. Provider-held continuation state retains affinity unless portability is established; routing alone does not make failover safe.
Common failure and possible evidence. Routing traffic through a gateway without testing a model change or its rollback leaves change control unproven. A reviewer could inspect model-change records, evaluations, and rollback evidence.
Prompt operations
Coverage and purpose. Prompt and configuration lifecycle controls help teams reproduce behavior and release changes deliberately.
What good looks like. Version prompts and settings together. Review changes, evaluate representative cases before release, and retain a tested rollback path.
Common failure and possible evidence. A prompt registry without release checks records history but does not test behavior. A reviewer could inspect prompt versions, settings, and release checks.
Guardrails
Coverage and purpose. Input, output, and action validation place controls at trusted boundaries before an unsafe operation can proceed.
What good looks like. Validate inputs, outputs, and proposed actions at trusted boundaries. Enforce authorization outside model instructions. Guardrails are neither a sandbox nor proof of hostile-process containment.
Common failure and possible evidence. Instructions that ask the model to respect a policy do not establish enforcement. A reviewer could inspect action-policy decisions and tests of rejected actions.
Budget governance
Coverage and purpose. Cost and execution limits bound resource use, including retries and agent loops, so an owner can contain runaway work.
What good looks like. Enforce operation, tenant, and tool limits with bounded retries and loops. Alert an accountable owner and stop or degrade work through tested containment paths.
Common failure and possible evidence. An alert alone can report overspend without stopping it. A reviewer could inspect limit configuration, alert records, and stop-path tests.
Tool governance
Coverage and purpose. Credentials, permissions, and action approvals constrain what tools can do and who can authorize their use.
What good looks like. Use scoped credentials, rotation, explicit permissions, and approvals for consequential actions. Verify enforcement with permission tests and audit records. A secret store alone does not establish least privilege.
Common failure and possible evidence. Central secrets with broad access leave permission boundaries unclear. A reviewer could inspect permission matrices, credential scopes, denied-operation tests, and audit records.
Observability
Coverage and purpose. Correlated model and tool telemetry helps reconstruct decisions and investigate failures while respecting sensitive data.
What good looks like. Correlate model and tool steps with versions, decisions, and incident handling. Redact sensitive data and define access and retention. Replay must not repeat external side effects. Local retention policy does not establish provider retention.
Common failure and possible evidence. Separate dashboards may show errors without reconstructing a request. A reviewer could inspect redacted correlated traces, access controls, and retention settings.
Evals
Coverage and purpose. Component and system evaluations test representative behavior before release and turn production failures into regression cases.
What good looks like. Gate changes with representative component and system evaluations. Review production failures and add regression cases. Monitor post-release quality and drift. Passing a finite suite does not prove universal correctness.
Common failure and possible evidence. Manual spot checks and complaints leave release behavior largely untested. A reviewer could inspect eval results, regression cases, and post-release quality monitoring.
Interpreting the result
A strong average cannot cancel a consequential gap. A narrowly scoped permission failure can deserve attention even when other layers score well. Questionnaire responses do not establish exposure, actual harm, or control enforcement.
Layers overlap. Action guardrails check proposed operations; tool permissions constrain the authority available to perform them. Both may affect the same action. Their scores are not independent risk probabilities and must not be multiplied or combined into an incident likelihood.
A human review adds architecture, evidence, uncertainty, and context to the priority order. It can confirm, reject, or rescope the questionnaire’s suggested priorities. See the fictional five-section written summary for the format, not a claim of completed inspection.