LLM Project Checklist

An ordered checklist for a new project, or for reviewing an existing LLM deployment. It is the companion artifact to Sequence Beats Model Choice, which explains the reasoning behind the two orderings the checklist depends on: controls are designed in Phase 4 and built in Phase 6, and the evaluation suite in Phase 5 gates the build. Vendor names appear as examples and traps, not as requirements.

How to use

Contents

Discovery

Planning

Implementation

Release

Operate

Reference


Phase 0: Intake

Gate: the charter records the requester, sponsor, budget owner, governing constraint, data classes, and candidate metric; the engagement is classified; and the premise has been sanity-checked against a non-LLM solution.


Phase 1: Discovery

Elicitation

The four question categories [Per use case]

Use-case inventory

Capability list [Per use case]

Translate preferences into constraints [Per use case]

Record requirements [Per use case]

Data and systems

Gate: requirements trace to the business case; assumptions are labeled; the before metric exists; and the Phase 0 open questions are answered or moved into the assumption register.


Phase 2: Feasibility and sizing [Per use case]

Technical feasibility (the four AI properties)

Sizing

SLA

Verdict and value

Gate: cost and latency modeled at production volume; the verdict, the load-bearing boundary condition, and the SLA are documented (the full boundary-condition list is finalized in Phase 3).


Phase 3: Architecture and pattern design [Per use case]

Decomposition

Pattern selection (five factors, first that rules out)

Reference architecture (the wiring, not a re-pick of the pattern)

Multi-agent (if orchestrator-workers)

Tools (if the system calls tools)

Model, context, prompts

Gate: every architecture choice names the constraint it resolves and the alternative it rejected.


Phase 4: Safety, controls, and fairness

Design the guarded path here, before it is built. This is the alignment boundary plus every control that will sit on the request path in Phase 6.

Alignment boundary

Guardrails

Risk assessment [Per use case]

Fairness and transparency [Per use case]

Human review [Per use case]

Compliance

Gate: each control has an owner, a failure direction, and an evidence artifact.


Phase 5: Evaluation [Per use case]

Complete the eval suite here, after the safety design, so it tests the control behaviors as well as task quality. This phase gates the build.

Eval suite

Gate every change

Gate: the eval suite exists and passes; every change runs through it. No production code before this gate.


Phase 6: Integration, reliability, and enterprise readiness

Build the guarded path decided in Phase 4, wire it into the enterprise stack, and add the reliability controls where the route and boundary are now known.

Delivery-team environment (before any code) [If you use an AI coding assistant]

Entry point and route (compliance first)

Identity, authorization, data

Observability

Reliability

Gate: identity, authorization, data handling, observability, and reliability are each owned and evidenced.


Phase 7: Documentation and handoff

The document serves three readers: the inheriting engineer, the auditor, and the returning architect. Serving one without the others leaves it incomplete.

Gate: completeness test passes; no load-bearing decision is undocumented.


Phase 8: Pre-launch readiness

Everything that must exist before launch: the team’s readiness to run AI-assisted work, the evidence that AI-assisted work is trustworthy, and the monitoring that decides what happens when a signal moves. All of it ships with the launch, not after it.

Team readiness at launch [If AI-assisted development is used]

Monitoring readiness

Gate: the governance table, SLA wiring, owners, and verification checklist all exist before launch.


Phase 9: Launch

Gate: the system is live with controls, logging, and owners in place.


Phase 10: Operate

Run the feedback loop, run experiments, and complete the outcome document.

Monitor

Experiment (if optimizing a live system)

Outcome [Per use case]

Gate: monitoring has triggers and owners; the outcome document is either complete or has a named owner and completion milestone.


Review Mode: auditing an existing project

Run the same phases as an audit. For each, ask “does this exist, is it current, and who owns it?”


Document set (summary)

The risk picture appears three times by design: boundary conditions (#12) list the constraints, the risk assessment (#13) scores likelihood and impact, and the failure-mode table (#25) maps architecture-specific failures. Produce them in that order, or merge them into one register if the project is small. Boundary conditions are finalized in Phase 3, after owner assignment; the load-bearing condition identified in Phase 2 travels in the feasibility memo.

Each row is one file to produce. The numbers are stable identifiers carried over from the original template set, so they are not in strict order; the Phase column is the order to produce them in. Phase 9 adds no document of its own: the launch record goes into the decision log (#11) and the architecture document (#26).

# Document Phase Required when
1 Project charter / one-page brief 0 always
2 Discovery notes 1 always
3 Translation table 1 always
4 Assumption register 1 always
5 Baseline metric record 1 always
6 Feasibility memo 2 always
7 Cost and latency model 2 always
8 ROI / business case 2 always (needed for expansion)
9 SLA definition 2 always
10 Solution architecture document 3 always
11 Decision log 3 always
12 Boundary conditions 3 if feasible with constraints
29 Statement of work / scope document 3 always
13 Risk assessment 4 always
14 Guardrail design and failure directions 4 always
15 Fairness instrumentation and decision-log design 4 always
16 Review-routing rule 4 if human review exists
17 Control register 4 if regulated
18 Eval plan and golden dataset 5 always
19 Model-change gate record 5 always
20 Entry-point-responsibility map 6 if multi-entry-point
21 Integration design document 6 always
22 Data-flow record 6 if residency or PHI
23 Constraint-to-integration matrix 6 if regulated
24 Reliability design 6 always
25 Failure-mode and mitigation table 6 always
30 Team setup and shared configuration 6 if you use an AI coding assistant
32 Spend policy 6 if you use an AI coding assistant
26 Architecture document with rationale 7 always
27 Runbook 7 always
28 Escalation path 7 always
31 Verification checklist 8 always
33 Governance table 8 always
34 Rollback plan 8 always
35 Experiment design and result record 10 if optimizing
36 Outcome document 10 always (may be deferred to a milestone)

Cross-cutting conditionals (keep visible throughout)

Condition What changes
Regulated (HIPAA, GDPR, FedRAMP, privilege) Compliance rules routes first; control register; scheduled checkpoints; business associate or data processing agreement per configuration
PHI or PII Minimum-necessary data; redaction before the call and before logging; log scope; retention controls
Data residency Explicit region pinning at the integration layer; verify logs, caches, monitoring, retention
Retrieval Screen retrieved content; monitor precision and recall; never use for live state
Agentic Tool budgets, stopping criteria, gate placement, coverage reconciliation, no per-step approvals
Multi-agent Shared trace ID; recoverable vs unrecoverable boundaries; coverage check at synthesis
Multi-tenant Separate API keys; per-tenant attribution and isolation
Multi-entry-point / multi-platform Entry-point-responsibility map; explicit region; model-id and feature-lag differences
High-consequence action Deterministic authorization before the action; human gate before irreversible steps
Conversational Multi-turn evals with a transcript golden dataset

Appendix: beyond the source material

A few items in this checklist are standard delivery practice rather than framework guidance. They are included because the checklist runs real projects, but the source material does not require them.

Operational delivery

Artifact hygiene

Everything else traces to the source material. The core is the safety controls (Phase 4), the evaluation suite (Phase 5), and the per-use-case scoping chain in Phases 1 and 2.