Performance and Load Testing: A Starter Guide for Small Teams

Load testing puts realistic traffic on your system to see how fast it stays and when it starts to fail. Load vs stress vs spike vs soak, the three numbers to watch, a small k6 script and a Locust one, how to set a goal before you run anything, and how to test without hurting anyone.

7 min read

Load testing means sending your app the amount of traffic you expect, from a script, and measuring how it responds: how long requests take, how many it handles per second, and how many fail. Performance testing is the wider family; load testing is its most common member, alongside stress tests that push past the expected peak, spike tests that hit it suddenly, and soak tests that hold a normal load for hours. For a small team the recipe is short: write down a goal first (“95% of requests under 500 ms at our expected peak, under 1% errors”), script the main user path in a tool such as k6 or Locust, run it against a staging environment you own, and treat a missed goal as a bug.

Load vs stress vs spike vs soak testing

The ISTQB glossary defines load testing (opens in a new tab) as performance testing that evaluates a system “under varying loads, usually between anticipated conditions of low, typical, and peak usage.” The other types change the shape or length of that load.

  • Load test: expected traffic, typical up to peak. Answers “are we fast enough on a normal busy day?”
  • Stress test: beyond the expected peak, or with fewer resources than usual. Answers “where does it break, and does it fail gracefully?”
  • Spike test: a sudden jump in traffic, such as a launch email or a mention on the news, then a drop. Answers “does it survive the burst and recover?”
  • Soak test (ISTQB calls it endurance testing): normal load held for hours. Catches memory leaks, filling disks and connection pools that slowly run dry.
  • Breakpoint test: load that keeps rising until something gives, to find the capacity limit.

Grafana’s k6 documentation describes the same family in its guide to load test types (opens in a new tab), and adds one more worth copying: start every round with a smoke test at minimal load, to check that the script works before you scale it up. That is a different use of “smoke” from the post-deploy check described in smoke testing vs sanity testing.

What to measure: latency percentiles, throughput, error rate

  • Latency as percentiles, not averages. The p95 is the time 95% of requests beat; the p99 is the slow tail. Google’s SRE book chapter on service level objectives (opens in a new tab) explains why: a simple average can hide tail latency, and a high percentile such as the 99th shows a plausible worst case. A p50 of 80 ms with a p99 of 4 seconds means one request in a hundred is painfully slow.
  • Throughput: requests or completed user journeys per second. It should rise with load until something saturates; when it flattens while latency climbs, you have found the bottleneck.
  • Error rate: the share of requests that fail, including timeouts. A fast error is still an error.
  • Saturation on the server side: CPU, memory, database connections and queue depth during the run. These tell you why the three numbers above moved.

How to set a performance goal

  1. Find your expected peak from your own analytics or logs: the busiest hour you have had, and what you expect at the next launch or seasonal peak. Do not borrow someone else’s numbers.
  2. Turn it into load: concurrent users or requests per second on the paths that matter, usually sign-in, search or browse, and the core action such as checkout.
  3. Write the goal as a pass or fail line per path: “checkout p95 under 800 ms and errors under 1% at 2x last year’s peak hour.” Decide the numbers with whoever owns the product, before the first run, not after you have seen results.
  4. Put the goal in the script as a threshold, so the run passes or fails on its own rather than on someone’s reading of a chart.

A tiny k6 load test

k6 scripts are JavaScript. This one ramps to 20 virtual users, holds for five minutes and ramps down, and fails the run if the goal is missed. The k6 thresholds (opens in a new tab) documentation explains that when a threshold is not met, k6 marks the test failed and exits with a non-zero code, which is what lets a pipeline stop on it.

load.js (k6)
import http from 'k6/http';
import { sleep } from 'k6';

export const options = {
  stages: [
    { duration: '2m', target: 20 }, // ramp up to 20 virtual users
    { duration: '5m', target: 20 }, // hold
    { duration: '1m', target: 0 },  // ramp down
  ],
  thresholds: {
    http_req_failed: ['rate<0.01'],   // under 1% errors
    http_req_duration: ['p(95)<500'], // 95% of requests under 500 ms
  },
};

export default function () {
  http.get(`${__ENV.BASE_URL}/api/products?zip=10001`);
  sleep(1);
}

// run against staging:
//   k6 run -e BASE_URL=https://staging.example.com load.js

The same idea in Locust

If your team writes Python, Locust describes users as classes. The Locust quickstart (opens in a new tab) shows a user class with @task methods that call self.client.get, started with locust and a web UI at port 8089, or without the UI using --headless.

locustfile.py (Locust)
from locust import HttpUser, task, between


class Shopper(HttpUser):
    wait_time = between(1, 3)

    @task(3)
    def browse(self):
        self.client.get("/api/products?zip=10001")

    @task(1)
    def view_cart(self):
        self.client.get("/api/cart")

# run without the web UI:
#   locust --headless --users 20 --spawn-rate 2 -H https://staging.example.com

Load testing tools compared

  • k6 (Grafana): JavaScript scripts, a command-line runner, thresholds built in. A good default for teams that already write JavaScript or TypeScript.
  • Locust: Python classes, a web UI for watching a run, and a headless mode for CI.
  • Apache JMeter: test plans built in a desktop GUI and saved as files. Its getting started guide (opens in a new tab) is blunt that the GUI is for creating and debugging tests, and load tests must run in CLI mode, for example jmeter -n -t my_test.jmx -l log.jtl.
  • Gatling: tests as code, with SDKs for Java, JavaScript, TypeScript, Scala and Kotlin, and a commercial Enterprise Edition alongside the open-source Community Edition.

Pick the one whose script language your team already reads. The tool matters far less than having a written goal and a realistic user path.

Testing safely

  • Never load test a system you do not own or have written permission to test. That includes someone else’s production, a partner’s API and a payment provider. Pointing heavy traffic at other people’s servers can look exactly like an attack.
  • Use a staging environment sized like production. A test against a smaller box tells you about the smaller box.
  • Stub third parties. Replace payment, email and SMS providers with test modes or fakes, or the test sends real messages and may break their terms.
  • Check your host’s rules first. Cloud providers publish testing policies; AWS, for example, separates ordinary network stress tests from simulated denial-of-service tests, which need their own approval.
  • Tell the team before a run, watch it live, and keep a stop button. If you must test production, do it in a quiet window with a small load and a person watching.

Where the results go

fenbs does not run load tests, store results or connect to CI; keep the scripts in your repository and the reports with your pipeline. The board is where the follow-up lives. A missed goal becomes a task of kind bug, such as “Checkout p95 is 1.4 s at peak load, goal 800 ms”, with the run’s numbers and the script’s commit in the note and a test status once the fix is re-measured. The goal itself is a decision someone made, so record it on the Decisions and rules page; a goal that holds from now on is a rule, and every connected AI assistant reads the rules first. fenbs has no due dates or sprints, so if the fix must land before a launch, say so in the task.

Related

Where load testing fits among the rest: types of software testing. Fast checks at the bottom of the pyramid: unit testing. Speed targets for pages rather than servers: the performance section of the QA checklist. Running a script on every change: what a CI/CD pipeline is.

Questions people ask.

What is the difference between load testing and performance testing?

Performance testing is the family of tests that measure speed, capacity and stability. Load testing is one member: it measures how the system behaves under expected traffic, from typical to peak. Stress, spike and soak tests are the others.

What is the difference between load testing and stress testing?

A load test uses the traffic you expect, up to your peak, to check you meet your goals. A stress test goes beyond the expected peak, or removes resources, to find where the system breaks and whether it fails gracefully.

Which numbers matter in a load test?

Latency percentiles such as p95 and p99 rather than averages, throughput in requests or journeys per second, and the error rate including timeouts. Server CPU, memory and database connections explain why those numbers moved.

Which load testing tool should a small team use?

The one whose scripting language the team already reads: k6 for JavaScript, Locust for Python, Gatling for JVM languages or TypeScript, and JMeter if you prefer building plans in a GUI. A written goal and a realistic user path matter more than the tool.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.