AI Production Scorecard · Method and evidence
An example written review summary
A fictional format example in five sections—not a completed review. This illustrates useful next steps without claiming inspected evidence or delivered outcomes.
Worked example using fictional inputs. This is not a client assessment. Scores summarize self-reported controls; they do not certify security, compliance, or production readiness.
1. Situation and evidence limits
A fictional Head of Engineering at a 600-person B2B software company. The production, customer-facing, multi-tenant support assistant handles customer PII and uses three tools or APIs, not an MCP server. The priority is reliability following recurring poor answers in production. No frequency, losses, client outcomes, or measured improvement is claimed.
This example uses questionnaire answers only. No architecture artifacts were inspected. Reported controls are not independently verified. The full worked scorecard records every selected answer and the methodology explains its limits.
2. Findings by layer
The overall rubric score is 45/100, band piloting. “Piloting” describes control maturity, not deployment stage. These are not safety percentages or industry percentiles.
Severity is a provisional review priority, not inferred automatically from the number. Evals and Tool governance merit first review; actual harm and exposure remain unverified. An average cannot cancel a consequential gap.
Model control — 75/100
Questionnaire finding. Most traffic uses a central model layer. Model swaps are controlled and tested, but the answers do not establish complete traffic coverage or fast, reversible changes across all providers and regions.
What good looks like. Model access follows an auditable policy. Changes are versioned, evaluated, and reversible. Provider-held continuation state retains affinity unless portability is established; routing alone does not make failover safe.
Prompt operations — 50/100
Questionnaire finding. A registry preserves prompt history, but release checks are manual spot checks.
What good looks like. Version prompts and settings together. Review changes, evaluate representative cases before release, and retain a tested rollback path.
Guardrails — 50/100
Questionnaire finding. Important flows have basic checks. Higher-risk tools have pre-use validation, but comprehensive enforcement is not established.
What good looks like. Validate inputs, outputs, and proposed actions at trusted boundaries. Enforce authorization outside model instructions. Guardrails are neither a sandbox nor proof of hostile-process containment.
Budget governance — 38/100
Questionnaire finding. Provider or cloud alerts and detection within hours do not establish automatic containment.
What good looks like. Enforce operation, tenant, and tool limits with bounded retries and loops. Alert an accountable owner and stop or degrade work through tested containment paths.
Tool governance — 25/100
Questionnaire finding. Secrets are centrally stored but broadly accessible. Tool permissions are documented informally.
What good looks like. Use scoped credentials, rotation, explicit permissions, and approvals for consequential actions. Verify enforcement with permission tests and audit records. A secret store alone does not establish least privilege.
Observability — 63/100
Questionnaire finding. Some requests are traceable manually. Errors, latency, tool health, and model metrics are visible together, but complete reconstruction is not established.
What good looks like. Correlate model and tool steps with versions, decisions, and incident handling. Redact sensitive data and define access and retention. Replay must not repeat external side effects. Local retention policy does not establish provider retention.
Evals — 13/100
Questionnaire finding. Pre-release checks are manual spot checks. User complaints are the main post-release signal.
What good looks like. Gate changes with representative component and system evaluations. Review production failures and add regression cases. Monitor post-release quality and drift. Passing a finite suite does not prove universal correctness.
3. Good-state targets for the three rubric priorities
- Evals. Gate changes with representative component and system evaluations. Review production failures and add regression cases. Monitor post-release quality and drift. Passing a finite suite does not prove universal correctness.
- Tool governance. Use scoped credentials, rotation, explicit permissions, and approvals for consequential actions. Verify enforcement with permission tests and audit records. A secret store alone does not establish least privilege.
- Budget governance. Enforce operation, tenant, and tool limits with bounded retries and loops. Alert an accountable owner and stop or degrade work through tested containment paths.
These targets are practice standards, not measured benchmarks. Exposure and inspection evidence may change the priority order.
4. Ranked next steps
- Turn one reported failure into a release evaluation. Verify it fails before the correction and passes after it. This would make that failure testable at release; it would not prove universal correctness. Which reported failure can the team reproduce, and who owns its regression case?
- Inspect and narrow tool permissions. Verify denied operations remain denied. This would constrain tool authority rather than rely on documented intent. Which operations are consequential, and who can approve their use?
- Establish a bounded budget. Test its stop path in a non-production environment. This would provide a containment path rather than alerts alone. What operation and tenant limits apply, and who owns safe degradation?
This is example advice, not authority to perform changes. No tests or changes are claimed to have occurred. No hours, savings, or incident-reduction estimates are supported by these answers.
5. Conditional fit statement
If the review confirms these gaps and an accountable owner can support remediation, we would propose a separately scoped readiness audit. If your team can resolve them internally, use the checks above and return with remaining questions.
This is not a guaranteed proposal or a paid-work commitment. A real summary must state no fit plainly when that is the conclusion.