Claude Code Tasks vs the Ralph Loop: How Each Keeps an Agent Going
Claude Code’s task list keeps Claude oriented inside a session. The Ralph loop keeps it working by feeding the same prompt back until the job is done. What each one is, how they differ, and where people and a board come in.
7 min read
Claude Code’s task list and the Ralph loop solve different problems. Claude Code’s task list is Claude’s own checklist for a multi-step job: it keeps Claude oriented, survives context compaction, and shows you what is pending, in progress and done. It does not make Claude keep working. The Ralph loop does exactly that and nothing else: it feeds the same prompt to Claude again and again, so each pass picks up from what the last one left in the files, until a stop condition is met. One is a map, the other an engine. Both leave open the questions a team cares about: who decided this work should happen, and who checked the result.
What the Ralph loop is
The technique was written up by Geoffrey Huntley in July 2025 under the title “Ralph Wiggum as a software engineer (opens in a new tab)”, after the character from The Simpsons. His own definition is short: “Ralph is a technique. In its purest form, Ralph is a Bash loop.”
while :; do cat PROMPT.md | claude-code ; done
The prompt never changes. What changes is the repository: each pass reads the code, the specifications and the plan left by the previous pass, does some work and exits, and the shell starts it again. Huntley’s advice around that loop is where most of the craft lives:
- One item per loop. He repeats it for emphasis: each pass should do one thing, which keeps the context small and the output better.
- Write the specifications first, in conversation with the model, and keep them in a folder every pass reads.
- Keep a prioritised plan file,
fix_plan.mdin his write-up, listing what is still to do. He warns that you will sometimes wake up to a broken codebase, and the plan is how you recover. - Tune it. When the loop does something bad, add an instruction to the prompt that steers it away, and run it again.
The official ralph-loop plugin
Anthropic publishes a plugin called ralph-loop in its official Claude Code marketplace (opens in a new tab), claude-plugins-official, which implements the technique inside a running session instead of a shell loop. It credits Huntley and links his post. It works through a Stop hook: when Claude finishes a turn and tries to stop, the hook blocks the stop and hands the original prompt back as the next instruction.
/plugin install ralph-loop@claude-plugins-official /ralph-loop "Make every test in tests/api pass. Output <promise>DONE</promise> when they all do." --completion-promise "DONE" --max-iterations 20 /cancel-ralph
--max-iterationsstops the loop after that many passes. Without it the limit is unlimited, and the README tells you to always set one.--completion-promiseends the loop when Claude’s last message contains that exact text inside<promise>tags. It is an exact string match, so it cannot tell “finished” from “blocked”.- The loop’s state is kept in
.claude/ralph-loop.local.mdin the project, including the iteration count, and the hook only acts on the session that started it. /cancel-ralphstops an active loop.
The plugin’s README (opens in a new tab) is direct about fit. It suggests the loop for well-defined work with clear success criteria, work that needs iteration such as getting tests to pass, greenfield projects you can walk away from, and anything with automatic verification. It advises against it for work that needs human judgement or design decisions, one-shot operations, unclear success criteria, and production debugging.
How it differs from the built-in task list
Claude Code’s task list is kept with the TaskCreate, TaskList, TaskGet and TaskUpdate tools, and Ctrl+T shows it. The documentation (opens in a new tab) says the items persist across context compactions, and that CLAUDE_CODE_TASK_LIST_ID shares one list across sessions in a named directory under ~/.claude/tasks/. The history and the other ways of tracking work are covered in Claude Code tasks vs to-dos and Claude Code tasks vs Beads vs a shared board. Set side by side with the loop:
- What keeps Claude going — task list: nothing; Claude stops when it judges the turn is done. Ralph: the shell loop or the Stop hook (opens in a new tab), which refuses to let it stop.
- Where the plan lives — task list: Claude’s checklist on your machine. Ralph: files in the repository, such as specifications and a plan file, which each pass reads afresh.
- When it stops — task list: when Claude finishes or you interrupt. Ralph: at the iteration limit, when the promise text appears, or when you cancel it.
- Who says it is done — task list: Claude marks its own items completed. Ralph: Claude writes the promise, and the loop takes its word for it.
They are not alternatives. A pass inside a Ralph loop can use Claude’s own task list to plan its steps, the same way any session does. That checklist is the inside view of one pass; the loop is what starts the next.
The risks of an unattended loop
The appeal of the loop is that nobody has to sit with it. That is also where the risks come from, and each has a plain counter-measure:
- It may never end. A goal that cannot be met keeps the loop running, and every pass uses your plan’s usage or your API budget. Set
--max-iterationsevery time, as the README says. - It ends on a claim, not a check. The completion promise is text Claude chooses to write. Make the prompt tie the promise to something a command proves, such as a test run, and check that command yourself afterwards.
- It drifts. The same prompt read against a changing codebase can wander into work nobody asked for. One item per pass and a written plan file keep it narrow.
- It breaks things between checks. Huntley’s own warning is that you will sometimes come back to a broken codebase. Run it on a branch or a separate worktree, commit often, and keep main out of reach.
- It runs with whatever permissions you gave it. A loop that stops to ask for approval is not unattended, so decide in advance which actions it may take without asking, and keep anything irreversible off that list.
Where human review fits
The loop removes people from the middle of the work, not from the ends. Three points stay human, and they line up with human in the loop for AI agents:
- Before: someone decides this piece of work is worth a loop, and writes the specification and the definition of done. This is most of the value, and the plugin’s own advice on clear completion criteria says as much.
- At the limit: when the loop stops on
--max-iterationsrather than on its promise, a person reads what happened before starting it again. Treat the limit as a checkpoint, not a failure. - After: a person reviews the diff and runs the check the promise was tied to. A promise is where review starts.
Where a board fits
A loop’s plan file is written for the loop. It is detailed, it changes every pass, and only people who open the repository see it. A board sits one level up: which pieces of work should get a loop at all, in what order, and what came of each. In fenbs that is one task per loop run. A person moves it to Next Up when it is ready; the prompt names its ref; the loop comments on it when it stops.
This run is for BUG-042 on the fenbs board. Read it first with fenbs_get_item and follow its Problem and Plan. Do one item per pass. Do not open or change any other task. When the tests named in the Plan pass, comment on BUG-042 with what changed, the commit and the test output, then output <promise>DONE</promise>. If you are blocked, comment why on BUG-042 and stop working.
Two settings keep the board honest while the loop runs unattended. Connect the session with the read and comment scopes only, so the loop can report on a task but cannot move it: a person moves it to Completed after reading the result, and every comment is kept in the history as “Claude via” the person who approved the connection. And if the loop misbehaves, revoke that one connection; the loop loses the board and nothing else changes. The wider rules for an assistant on a board are in how to keep an AI agent from wrecking your board.
Running several loops at once is a separate problem — two loops can start the same piece of work — and running Claude Code tasks in parallel covers claiming and worktrees.
Related
Planning work for Claude Code to pull: a kanban board for Claude Code. What the scopes mean: assistant tokens and scopes. Connecting Claude Code: Claude Code integration.