Feature Flags: Ship Dark, Release Safely
A feature flag lets you deploy code switched off and turn it on later, for some users or all of them, without another deploy. The four kinds of flag, the best practices that keep them safe, and the cleanup habit that stops old flags becoming debt.
7 min read
A feature flag is a switch in your code that decides at runtime whether a piece of functionality is on, for everyone, for some users or for nobody. It separates deploying code from releasing it: you ship the new checkout switched off, turn it on for your own team, then for a small share of customers, then for all of them, and if something breaks you turn it off without a redeploy. The catch is that every flag is a second code path, and flags that outlive their purpose become debt. Give each flag an owner, a type and a removal task the day you create it.
What a feature flag is
In its simplest form a flag is an if around new code, with the condition read from configuration rather than written into the code. Pete Hodgson’s article on feature toggles (opens in a new tab), published on Martin Fowler’s site, is the standard reference, and it treats “feature toggle” and “feature flag” as the same thing. The configuration can be an environment variable, a row in your database or a hosted flag service; what matters is that someone can change the answer without shipping new code.
if (flags.isOn('new-checkout', { userId: user.id })) {
return renderNewCheckout(cart);
}
return renderCheckout(cart);That small if buys three things. You can merge unfinished work to the main branch every day instead of keeping a long-lived branch. You can release to a few people first and watch. And you can switch a bad release off in seconds, which is usually faster and safer than rolling back a deploy.
Types of feature flags
Hodgson sorts flags into four categories by how long they live and how often their answer changes. The type decides how you should treat the flag, so name it when you create one.
- Release flags hide unfinished or unreleased work. They should live days or weeks, and they are the ones that most often get forgotten.
- Experiment flags split users between variants for an A/B test. They change per request and live as long as the experiment needs to reach an answer.
- Ops flags are kill switches and load valves: turn off the expensive recommendations panel when the database is struggling. Some are short-lived; a few are meant to stay for good.
- Permission flags turn features on for certain users, such as paying customers or internal staff. They can live for years and are really part of your product’s access rules.
The first two should always be temporary. The last two may be permanent, and when they are, treat them like any other product setting: documented, tested and owned.
Shipping dark, then releasing in steps
“Shipping dark” means deploying code that is switched off for everyone. The release then becomes a series of small, reversible decisions rather than one big one:
- Deploy with the flag off. Nothing changes for users; you are only checking that the new code does not break the old path.
- Turn it on for your own team. Use the feature for real work for a day or two.
- Turn it on for a small share of users, or one customer who asked for it. Watch errors and support messages.
- Widen in steps, pausing at each one long enough to see a problem if there is one.
- Turn it on for everyone, then schedule the removal. The release is done when the flag is gone, not when it reaches 100%.
Before step 1, test both states. A flag doubles the paths through that code, and the off path is the one that gets forgotten. The QA checklist is the pre-release pass; run its functional checks with the flag on and with it off.
Feature flag best practices
Vendor guidance agrees on most of the basics. Unleash’s feature flag best practices (opens in a new tab) include making flags short-lived, giving them unique names, evaluating them on the server to protect personal data, and folding flag-removal tasks into planning “just as you would with technical debt.” The practices that matter most for a small team:
- One flag, one purpose. Never reuse an old flag’s name for new behavior; create a new one.
- Name flags so they read in a sentence:
new-checkout, notflag7ortest-jen. - Make the off state the safe state. If the flag service is unreachable, the code should fall back to the behavior you are sure of.
- Keep decisions on the server when they depend on who the user is, so rules and customer data are not shipped to the browser.
- Record who owns each flag and when it should be gone. An ownerless flag is permanent by default.
- Keep flag changes in history. When a customer reports a problem, “what changed at 2:14 p.m.?” should include flips, not only deploys.
Flags are debt with a carrying cost
Hodgson’s line is the one to remember: “Savvy teams view their Feature Toggles in their codebase as inventory which comes with a carrying cost and seek to keep that inventory as low as possible.” Every flag left in place after release is a branch nobody tests anymore, a condition every reader has to parse, and a trap for whoever reuses it. It is technical debt of the most predictable kind: you knew the day you created it that it would need to come out.
The best-known cautionary tale is Knight Capital. According to the SEC’s 2013 order against the firm, code deployed for a new trading program in July 2012 repurposed a flag that had once activated an old, unused function called Power Peg. The old code had never been removed, and a technician did not copy the new code to one of eight servers. On August 1, orders carrying the repurposed flag reached that server and woke the dead code, which sent millions of orders in about 45 minutes. The firm lost more than 460 million dollars. The order also faults the deployment process, but the flag lesson is plain: dead code behind an old flag is not harmless, and a flag’s name should never be recycled for new behavior.
Track every flag as a task
Hodgson describes the habit that works: some teams always add a toggle removal task to the backlog whenever a release toggle is first introduced, and others put expiration dates on flags, or even “time bombs” that fail a test if a flag outlives its date. The task is the cheapest of these and it works with any flag tool. Write it at the same time as the flag:
Title: Remove flag new-checkout Type: release flag (temporary) Owner: [name] Created: September 30, 2026 Remove when: on for 100% of users for two weeks with no rollback Plan: 1. Delete the old checkout path and the flag check 2. Delete the flag from the flag service 3. Run the checkout tests; confirm nothing still reads the flag
Delete the flag from the code before you delete it from the service. LaunchDarkly’s documentation on archiving flags (opens in a new tab) makes the same point: remove every code reference first, because an archived flag that is still evaluated serves its fallback value, which may not be the behavior you expect.
Build or buy a flag system
A small team can start with a configuration file or a table and a helper function; that covers release and ops flags. Percentage rollouts, per-user targeting, experiments and an audit log of flips are where hosted services earn their place. If you want to keep the choice open, OpenFeature (opens in a new tab), a Cloud Native Computing Foundation incubating project, defines a vendor-neutral SDK that plugs into different flag providers, so your code calls one interface whichever provider sits behind it.
Flag tasks on a fenbs board
fenbs is not a feature flag service and does not switch anything in your app. It holds the work around each flag. The feature, say FET-310 “New checkout,” gets a linked enhancement, ENH-311 “Remove flag new-checkout,” created the same day and linked with relatesTo, so opening either one shows the other. The removal task’s plan lists the steps above, and its test status records whether someone checked that nothing still reads the flag.
- Filter the board by a category such as “Flags” to see every flag still waiting to come out.
- fenbs has no due dates, so “remove when” is a condition in the plan, reviewed when you plan the week, not a reminder.
- A standing rule such as “every release flag gets a removal task when it is created” belongs on the Decisions and rules page, where every connected AI assistant reads it before touching the board.
- An AI assistant that adds a flag while building FET-310 can file the removal task itself over MCP, with
fenbs_create_item.
Related
Checks to run with the flag on and off: QA checklist. Why old flags count as debt: technical debt. Catching what a flag change broke: regression testing checklist. Telling users what just turned on: release notes template.