DevOps Tools: The Small-Team Stack

A small team needs one tool for each of eight jobs, from source control to incident tracking, and can skip most of what large platform teams run. The stack by job, a one-page template to write yours down, and what to leave out until it hurts.

6 min read

A small team’s DevOps tooling comes down to eight jobs: source control, CI/CD, infrastructure as code, containers, secrets, monitoring and logging, alerting and on-call, and incident tracking. Pick one tool per job, prefer the one already built into something you use, and write the whole stack on one page so anyone can find where a thing lives. Most of what large organizations run, such as Kubernetes clusters, service meshes, internal developer portals and several overlapping monitoring products, can wait until a specific problem asks for it.

If you want the idea behind the tools first, what is DevOps covers the definition and the practices. This page is the shopping list.

DevOps tools by job

1. Source control

Git, on a hosted service: GitHub, GitLab, Bitbucket or Azure Repos. Everything else in the stack hangs off it, so choose it first. What matters on day one is not the host but the rules: protect the main branch, require a pull request and a passing check to merge, and keep branches short-lived. Git branching strategy and code review best practices cover both.

2. CI/CD

The service that tests every change and deploys what passes (the stages are explained in What is a CI/CD pipeline?). Start with the one built into your code host, on the vendor’s machines. The eight common choices, hosted and self-hosted, are compared in CI/CD tools compared.

3. Infrastructure as code

Servers, databases, DNS and cloud permissions described in files and changed through the same pull requests as code. HashiCorp’s Terraform introduction (opens in a new tab) describes it as “an infrastructure as code tool that lets you build, change, and version cloud and on-prem resources safely and efficiently”, with a write, plan, apply workflow and a state file that records what exists.

Know that the ecosystem split in two: the OpenTofu FAQ (opens in a new tab) says OpenTofu is a Terraform fork, created in response to HashiCorp’s switch from an open-source license to the Business Source License, and now held by the Linux Foundation. Cloud-specific options exist too, such as AWS CloudFormation and Azure Bicep.

For a team on one platform-as-a-service host, infrastructure as code can start as small as the host’s own config file checked into the repository. The rule that matters: nobody changes production settings by clicking, or if they must, they write it down the same day.

4. Containers

A container packages the app with everything it needs, so it runs the same on a laptop, in CI and in production. In Docker’s own overview (opens in a new tab), an image is “a read-only template with instructions for creating a Docker container” and a container is “a runnable instance of an image.” A Dockerfile in the repository and a registry to hold the built images is the whole requirement at small scale. Running containers is a separate decision: a managed container service or platform host is enough long before a cluster is.

5. Secrets

API keys, database passwords and deploy tokens belong in your CI/CD tool’s secret store or a cloud secret manager, never in the repository. Two cheap safeguards: scanning that stops a secret before it lands, such as GitHub’s push protection (opens in a new tab), which blocks pushes that contain secrets before they reach the repository and is on by default for users pushing to public repositories; and an expiry on every credential that allows one. Tokens held by AI assistants count too: a security review of AI agent access is a quarterly habit worth borrowing.

6. Monitoring and logging

One place for errors, one for metrics, one for logs, and often a single product covers all three. Instrument with an open standard so you can change the back end later: the OpenTelemetry project (opens in a new tab) describes itself as vendor- and tool-agnostic, for traces, metrics and logs, and is a Cloud Native Computing Foundation project. Measure what users feel before what machines feel.

7. Alerting and on-call

Alerts go to a person who is expected to act. The chapter on monitoring distributed systems (opens in a new tab) in Google’s SRE book names four golden signals, latency, traffic, errors and saturation, and sets the bar for paging: “Every page should be actionable.” A small team needs a rotation, a paging app or your monitoring tool’s own notifications, and a runbook behind every alert. On-call rotation and the runbook template cover the setup.

8. Incident and work tracking

When something breaks, three records matter: the incident itself, the postmortem, and the follow-up work. The incident postmortem template covers the second. The third is where most small teams lose things: action items written in a document nobody reopens. Put each one on the same board as the rest of your features, enhancements and bugs, where it competes for priority in the open.

Write the stack on one page

The most useful DevOps document a small team can own is a list of which tool does which job, who owns it, and where its config lives. Keep it in the repository, next to the code it describes:

docs/stack.md
# Our stack (reviewed: September 30, 2026)

job                 tool                         owner    config lives in
source control      GitHub                       Maria    repo settings, rulesets
ci/cd               GitHub Actions               Maria    .github/workflows/
infrastructure      OpenTofu                     Sam      infra/
containers          Docker + host registry       Sam      Dockerfile
secrets             CI secret store              Sam      repo + environment secrets
monitoring/logs     one hosted product           Priya    instrumentation in src/telemetry
alerting/on-call    monitoring alerts + rotation Priya    docs/on-call.md
work + incidents    one task board               Priya    board; postmortems in docs/incidents/

## Rules
- Nothing reaches production except through the CI/CD deploy job.
- Every alert that pages links to a runbook.
- Every postmortem action item becomes a task on the board.

The owner is not the only person allowed to touch the tool; they are the person who notices when it drifts. Review the page when someone joins or leaves.

What to skip early

  • Kubernetes, until you run enough services that a managed container service or platform host is the thing slowing you down.
  • A service mesh, multi-cloud, and multi-region failover, until a customer contract or a real outage asks for them.
  • A self-hosted CI server, unless you need special hardware or must keep builds on your own network.
  • An internal developer portal. With one repository and a stack page, the portal is the README.
  • A second monitoring product that overlaps the first. Two dashboards that disagree are worse than one.
  • A metrics platform for delivery numbers. A spreadsheet and your deploy log are enough to track DORA metrics for months.
  • Chat bots that deploy or roll back. Make the deploy job easy to run first; wrap it later.

DevOps best practices that matter more than the tools

  • Ship small changes often. Small batches are easier to review, test and roll back than big ones.
  • Keep everything that defines the system in version control: code, pipeline, infrastructure, dashboards where you can.
  • Have exactly one way to deploy, and make it the automated one.
  • Alert on symptoms users feel, not every cause a machine can report.
  • Run blameless postmortems, and track every action item to done.
  • Decouple deploying from releasing with feature flags once risky changes become routine.

Where fenbs fits, and where it does not

fenbs covers the last job only: tracking the work. A board has four lanes, To Do, Next Up, In Progress and Completed, and every task is a feature, an enhancement or a bug, with a priority from 1 to 10 where 1 is the most urgent. A postmortem’s action items can be pasted in at once with Add many, one per line. Each task has a note for the problem, a plan for how it will be fixed, and a test status with notes, and History records who changed what, whether a person or an AI assistant connected over MCP.

Team conventions such as “never deploy on a Friday” can be recorded on the Decisions and rules page as rules, with the person who decided them; every connected AI assistant reads the rules first. What fenbs does not do is just as important for this stack: it has no alerting, paging or on-call schedule, no due dates, no sprints, no settable assignee, and no GitHub or CI/CD integration. Those jobs stay with the tools above.

Related

The whole lifecycle these tools support: software development life cycle. When the incident was caused by an AI agent: AI agent incident response. Letting an assistant read your errors: Sentry MCP. Filing and triaging bugs: bug tracking tools.

Questions people ask.

What are DevOps tools?

The software a team uses to build, ship and run its product: source control, CI/CD, infrastructure as code, containers, secret storage, monitoring and logging, alerting and on-call, and incident and work tracking. The tools support the practices; they are not DevOps on their own.

Which DevOps tools does a small team need first?

Source control with branch protection, CI/CD that tests every change and deploys what passes, a secret store, and one monitoring product with alerts that reach a person. Infrastructure as code and containers usually come next.

Does a small team need Kubernetes?

Usually not at first. A managed container service or a platform host runs a handful of services with far less to operate. Kubernetes starts to pay off when you run many services and need its scheduling and scaling, and someone has time to run it.

Should I use Terraform or OpenTofu?

Both are infrastructure as code tools with the same roots; OpenTofu is a fork created after HashiCorp moved Terraform to the Business Source License, and it is held by the Linux Foundation. Check the license terms against how you use it, and check which one your cloud providers and modules support.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.