Managing AI Projects: A Lightweight Method
When the thing you are building is an AI feature, “done” stops being obvious. A lightweight method: write the success criteria first, build an evaluation set before the feature, treat every prompt and model change as a task, and accept work by its scores rather than by a demo.
7 min read
Project management for AI projects needs four changes to the way you would run any other software project. Write down what success means as numbers before anyone builds anything. Build an evaluation set — a fixed list of inputs with the answers you expect — before the feature, and run it on every change. Treat each prompt edit, model switch or retrieval tweak as its own task with a plan and a recorded result. And accept work when the evaluation says it passes, not when a demo looks good. Everything else — a board, a short risk list, a weekly rhythm — you already know how to do.
This is about projects whose deliverable is an AI feature or automation: a support-ticket classifier, a summariser, an extraction step in a pipeline, an assistant inside your product. Using AI assistants to help run a project is a different subject, covered in what is AI project management and how to use AI agents in project management.
Why the usual plan breaks
In ordinary software, a feature either does what the ticket says or it does not, and a test proves which. An AI feature is judged on a distribution. It gets most cases right, some wrong, and a few badly wrong, and the same input can give a different answer tomorrow. OpenAI’s evaluation best practices (opens in a new tab) put it plainly: models can produce different output from the same input, which makes traditional software testing insufficient on its own.
Three things follow for whoever runs the project. You cannot estimate “make it accurate” the way you estimate a form; you can only estimate the next experiment. Progress is a number moving, not a checklist emptying. And a change that fixes one case can quietly break five others, so nothing counts as finished until the whole set has been run again.
Step 1: write success criteria as numbers
Before the first prompt is written, agree what good looks like. Anthropic’s guide to defining success criteria (opens in a new tab) asks for criteria that are specific, measurable, achievable and relevant, and notes that most uses need several of them at once. “Classifies tickets well” is a wish. “At least 90% of tickets land in the right queue on our 300-ticket set, no ticket containing a refund request is ever routed to spam, and the answer comes back in under two seconds” is something you can accept or reject.
- One quality number the feature is mostly judged on.
- One or two “never” rules for the failures that would really hurt: leaking data, a rude reply to a customer, a wrong amount.
- An operational limit: time per answer, or cost per thousand requests, whichever your owner actually cares about.
The thresholds are a decision, not a detail. Somebody who owns the outcome sets them, and the reason is written down, because in three months someone will ask why 90% and not 95%.
Step 2: build the evaluation set before the feature
The evaluation set is the project’s real specification. Collect real inputs — anonymised tickets, documents, messages — and write the expected answer beside each. Anthropic’s advice on building evaluations (opens in a new tab) is to mirror the real distribution of your task, include edge cases, automate the grading where you can, and prefer many automatically graded cases to a few hand-graded ones.
- Start with 50 to 100 cases you can grade automatically: an exact label, a field that must match, a number within a tolerance.
- Add the awkward ones on purpose: empty input, very long input, two languages at once, a case even your team argues about.
- Keep a small hand-graded slice for what code cannot judge, such as tone. Grade it the same way each time, against a written rubric.
- Freeze a held-out part nobody tunes against, so the headline number is not just the cases you practised on.
Building the set is a task in its own right, usually the first one on the board. It is also the one most often skipped, because a working demo arrives sooner without it.
Step 3: every prompt or model change is a task
On an AI project the code may change very little while the behaviour changes a lot. A reworded instruction, a new model version, a different chunk size for retrieval: each is a change to the product, and each should be as visible as a code change. Give each one a task, a plan that says what will change and what result is expected, and a recorded result once the set has run.
ENH-041 Ticket classifier: add refund examples to the prompt
Problem: 7 of 40 refund tickets on the eval set go to "general".
Plan: Add three refund examples to the system prompt.
Run the full set (300) and the held-out set (60).
Expect refund accuracy up; overall must not drop.
Testing: Partly tested. Full set 91.3% (was 89.0%), refunds 38/40.
Held-out 88.3% (was 88.3%). Not yet run on live traffic.Two habits make this work. Keep the problem and the plan apart, so the original description survives however many times the plan is rewritten. And write the result as numbers against the previous run, including what was not checked. “Better” is not a result.
Step 4: accept by evaluation, not by demo
The demo is where AI projects fool themselves. Someone tries five inputs, all five look right, and the feature is declared done. OpenAI’s guide names this anti-pattern “vibe-based evals” and recommends running evaluations early and on every change. Make that the acceptance rule: a task moves to Completed when the set has been run on the final version and the numbers meet the criteria from step 1. A demo is still useful, for showing people what the feature does. It just does not decide whether it ships.
When a number falls short, you have three honest options: keep working, narrow the scope (route only the ticket types it handles well), or change the threshold. The third is allowed, but it is a new decision made by the person who set the first one, and recorded as such.
Step 5: keep a short risk list
AI features carry risks an ordinary feature does not: wrong answers delivered with confidence, personal data in prompts or logs, inputs designed to make the model misbehave, and a model provider changing behaviour underneath you. The NIST AI Risk Management Framework (opens in a new tab) organises this work into four functions — Govern, Map, Measure and Manage — and NIST has published a companion profile for generative AI. You do not need the whole framework for one feature, but its shape is a good checklist for a small team:
- Govern: who owns the feature, who may change the thresholds, and who can switch it off.
- Map: where it is used, who sees its output, and what happens downstream when it is wrong.
- Measure: the evaluation set, plus the “never” rules as their own checks.
- Manage: what you do when a number drops — roll back the prompt, fall back to a person, or turn the feature off.
Each risk that needs work becomes a task like any other. A risk list that lives in a slide deck is a risk list nobody reads twice.
Running it on a board
None of this needs a special tool, and it does not need sprints or story points. On fenbs it maps onto what is already there. Each experiment is a task — a feature for new capability, an enhancement for a tuning change, a bug for a case the feature gets wrong — moving through To Do, Next Up, In Progress and Completed. The task’s problem box says what is wrong, its plan says what will change and what result is expected, and its Testing status (Tested, Partly tested, Failed or Needs owner check) and notes carry the scores. Filtering for “Completed, not tested” before a release shows anything that shipped without a run.
The thresholds and the “never” rules go on the Decisions page, with who decided and why, linked to the tasks that depend on them. An AI assistant can write a decision down but is never recorded as the one who made it. If an assistant runs the evaluation for you, its comments and status changes are recorded in History by name — “Claude via” the person who connected it — so a score can always be traced to who produced it.
Related
How to check what an assistant hands back: verifying AI-generated work. Planning in fixed iterations without a sprint feature: sprint planning with AI. The three kinds of task: feature, enhancement, bug.