Load testing

Know your p99
before your customer does.

Enterprise buyers ask how your service performs under load and what uptime you can commit to. This guide covers the kinds of performance test, how to load test an API properly, which numbers matter, and how to keep the results as evidence.

LoadStressSpikeSoakp95 / p99SLOsCI

Last updated Published by TryTrustableNot legal advice

Short answer

Load testing sends a realistic, controlled volume of traffic to your system and measures how it responds: latency percentiles (p50, p95, p99), throughput and error rate. You compare the results with a target you set in advance, a service level objective such as "p95 under 300 ms with under 0.1% errors at 200 requests per second". Stress, spike and soak tests vary the shape of that traffic to find the breaking point, sudden surges and slow leaks. Run a small load test on every deploy and a larger one before big launches, and keep the results: enterprise buyers and SOC 2 auditors ask for them.

01

What is load testing?

Load testing is a type of performance testing that puts a known amount of traffic on a system and measures how it behaves. The output is a handful of numbers (latency percentiles, throughput, error rate) compared with a target. It answers a practical question: at the traffic we expect, do users get fast, correct responses?

02

Load testing vs stress testing vs spike and soak testing

TestTraffic shapeQuestion it answersWhen to run it
Load testExpected traffic, held steadyDo we meet our SLO at normal and peak load?Every deploy (small), before launches (full)
Stress testRamp beyond expected load until something breaksWhere is the breaking point, and how does it fail?Quarterly, and before capacity planning
Spike testA sudden jump, then back downDo we survive a burst (a launch, a marketing email, a retry storm) and recover?Before known events
Soak testNormal load for hoursDo memory, connections or queues leak over time?Before major releases, after infrastructure changes
Smoke testA handful of requestsDoes the endpoint work at all?On every deploy, before the real test

Names vary between teams and tools; what matters is that each test has a stated question and a pass/fail target.

03

How to load test an API, step by step

  1. Pick the endpoints that matter. Sign-in, the main read path, the main write path, search, and anything a customer integration calls. Do not test only the health check.
  2. Set the target first. Write the SLO before the run: for example, p95 under 300 ms and error rate under 0.1% at 200 requests per second. A target set after seeing the result is not a test.
  3. Use a production-like environment with realistic data volumes. An empty database is fast.
  4. Authenticate like a real client. Use test accounts and tokens so you measure the real code path, not a stream of 401s.
  5. Ramp up, hold, ramp down. Warm caches and connection pools, then hold steady long enough to see the tail.
  6. Watch the server too. CPU, memory, database connections and queue depth explain the client-side numbers.
  7. Record and compare. Keep each run with its target and the version tested, and compare with the last one.

Never load test systems you do not own or have permission to test, and warn your cloud provider or CDN if their terms require it.

04

Which percentiles matter: p95 and p99, not averages

Google's SRE book makes the point plainly: averaging request latencies obscures the long tail, where most requests are fast but some are much slower. A service with a p50 of 110 ms and a p99 of 2.4 seconds feels fast in a demo and slow to the customer whose dashboard makes twenty requests per page load: if 1% of requests are that slow, about 18% of those page loads (1 minus 0.99 to the power 20) include at least one.

MetricWhat it tells youTypical use
p50 (median)What a typical request experiencesTrack trends; do not set the SLO on it alone
p95The experience of the slowest 1 in 20 requestsThe usual latency SLO for user-facing APIs
p99The slowest 1 in 100: queueing, garbage collection, cold caches, slow queriesSLOs for critical paths; capacity warnings
AverageMixes fast and slow requests into one numberAvoid as a target: a long tail can hide behind a good average
Throughput (requests per second)How much work the system completedConfirms the test actually generated the intended load
Error rateShare of failed requests (timeouts, 5xx)Always part of the SLO; a fast error is still an error
05

Setting a performance SLO

An SLO combines a metric, a threshold and a condition: "p95 latency of the orders API under 300 ms, error rate under 0.1%, at 200 requests per second." Start from what users and your contracts need, measure where you are today, and set a target you can meet with some margin. Tighten it later. The SLA, SLO and SLI guide covers error budgets and how SLOs relate to the SLA you sign.

06

Load testing in CI/CD

A full load test is too slow for every pull request, but a short, small test on every deploy catches the regressions that matter: an N+1 query, a missing index, a synchronous call added to a hot path. A common pattern:

  • Every deploy to staging: a 30 to 60 second test of the key endpoints against the SLO.
  • Nightly or weekly: a longer load test at expected peak.
  • Before launches and quarterly: stress, spike and soak tests.

Store every result with the commit it tested, so a slowdown can be traced to a change.

07

Open-source load testing tools

ToolTests written inNotes
k6 (Grafana Labs)JavaScriptOpen source; thresholds let a run pass or fail, which suits CI
Apache JMeterGUI test plans (Java-based)Long-established, wide protocol support, large plugin ecosystem
LocustPythonOpen source; user behaviour as Python code; distributed workers
GatlingJava, other JVM languages, or JavaScript/TypeScriptOpen-source Community Edition plus a commercial edition

A neutral overview, not a ranking. Pick the one whose scripting language your team already uses.

08

How TryTrustable runs load tests

TryTrustable treats performance as an availability control: the test, the target and the result are recorded together, so the record is the evidence an auditor or buyer can inspect.

In TryTrustableWhat it does
API load testsIn-process load tests with concurrent workers, reporting p50, p95 and p99 latency, average, throughput and error rate
SLO evaluationEach run is checked against the target you set (p95 latency, error rate, minimum requests per second) and recorded as met or breached
Frontend auditsGoogle PageSpeed Insights (Lighthouse) for a public URL; see Core Web Vitals
SchedulesManual, hourly, daily or on deploy
Endpoint discoveryDuring trytrustable sync the SDK finds API routes in Express, Fastify, NestJS and Next.js code
Load tests in your CItrytrustable loadtest <url> runs the same test inside your pipeline and reports results to the platform
EvidenceResults, including failures, are stored as evidence for availability controls such as SOC 2 A1.1

Hosted runs are deliberately bounded (up to 50 concurrent workers, up to 30 seconds). They prove an SLO at steady load; for breaking-point stress tests use a dedicated tool.

Read more on the performance testing page, or see what buyers ask for in the enterprise readiness guide.

Questions

The things people ask us

What is the difference between load testing and stress testing?

A load test checks behaviour at expected traffic against a target. A stress test pushes beyond expected traffic until something fails, to find the limit and see how the system fails and recovers.

How long should a load test run?

Long enough to reach steady state and see the tail: minutes for a deploy check, an hour or more at peak for a full test, and several hours for a soak test that looks for leaks.

Why use p95 or p99 instead of average response time?

Averages hide the slow tail. p95 and p99 show what the slowest 5% and 1% of requests experience, which is what users with busy pages and integrations actually hit.

What is a good API response time?

It depends on the endpoint and the user. Set the target from what your users and contracts need, write it as an SLO on p95 or p99 with an error rate, and test against it.

Can load tests count as SOC 2 evidence?

They can support availability criteria such as A1.1, which is about monitoring capacity and planning for demand. Auditors want results over time against a stated target, not a single screenshot.

Can TryTrustable replace k6 or JMeter?

Not entirely. It covers steady-load SLO checks with the results stored as evidence. Hosted runs are bounded (up to 50 concurrent workers, 30 seconds), so for large stress or soak tests use a dedicated tool alongside it.

Book a walkthrough

Turn your load tests into evidence buyers can check.

Thirty minutes: we run a load test against one of your endpoints, set an SLO, and show the result landing as availability evidence.