Continuous Load Testing in CI/CD Pipeline
With every deploy to production, the risk of introducing performance regression is high. One suboptimal SQL query or a forgotten N+1 problem can increase response time by 50%, and users will switch to competitors. We integrate automated load tests directly into your CI/CD pipeline to catch such regressions before they reach master. The result: you are confident that every commit does not break SLA for latency and throughput.
We configure smoke tests with threshold values that turn the pipeline into a gatekeeper: if p95 latency exceeds 500ms or error rate exceeds 1%, the pipeline fails and the developer receives a notification. This is not a replacement for full-scale load testing, but a quick safeguard.
Problems We Solve
- Performance regression after deploy — even a micro-change in code can slow down a critical endpoint. Without automated tests, you learn about the problem only after user complaints or monitoring alerts. We set up baseline comparison: each test runs against the stable branch, and if deviation >20%, the pipeline is blocked.
- Inefficient queries and bottlenecks — our scripts emulate typical user scenarios (listing, creating a post, searching). If a query starts to take longer, we see it on metric charts.
- Lack of performance-first culture — developers often don't think about performance during code review. Continuous Load Testing makes performance visible: each PR is accompanied by a comment with test results (p95, error rate).
Tools and Their Place in CI
k6 is the best choice for CI: JS scripts, built-in statistics, threshold-based pass/fail, native integration with GitHub Actions and GitLab CI. More at k6 documentation.
Artillery — YAML configuration, convenient for describing scenarios without code.
Gatling — Scala/Java, detailed HTML reports, convenient for Java teams.
Basic k6 Script
// tests/performance/api-smoke.js
import http from 'k6/http'
import { check, sleep } from 'k6'
import { Rate, Trend } from 'k6/metrics'
// Custom metrics
const errorRate = new Rate('errors')
const postCreateDuration = new Trend('post_create_duration')
export const options = {
// Load profile for CI: fast, non-destructive
stages: [
{ duration: '30s', target: 10 }, // warm-up
{ duration: '1m', target: 10 }, // sustained load
{ duration: '10s', target: 0 }, // cool-down
],
// Pipeline will break if thresholds are not met
thresholds: {
http_req_duration: [
'p(95)<500', // p95 < 500ms
'p(99)<1000', // p99 < 1000ms
],
errors: ['rate<0.01'], // errors < 1%
http_req_failed: ['rate<0.01'], // HTTP errors < 1%
post_create_duration: ['p(95)<800'],
}
}
const BASE_URL = __ENV.BASE_URL || 'http://localhost:3000'
const AUTH_TOKEN = __ENV.AUTH_TOKEN
export function setup() {
// Once: get token or prepare data
const res = http.post(`${BASE_URL}/api/auth/login`, JSON.stringify({
email: '[email protected]',
password: 'testpassword'
}), { headers: { 'Content-Type': 'application/json' } })
return { token: res.json('token') }
}
export default function(data) {
const headers = {
'Content-Type': 'application/json',
'Authorization': `Bearer ${data.token || AUTH_TOKEN}`
}
// Scenario 1: list posts (70% traffic)
const postsList = http.get(`${BASE_URL}/api/posts?limit=20`, { headers })
check(postsList, {
'posts list: status 200': (r) => r.status === 200,
'posts list: has items': (r) => r.json('data').length > 0
})
errorRate.add(postsList.status !== 200)
sleep(Math.random() * 0.5) // random pause 0-500ms
// Scenario 2: create post (20% traffic)
if (Math.random() < 0.2) {
const start = Date.now()
const createPost = http.post(`${BASE_URL}/api/posts`, JSON.stringify({
title: `Test post ${Date.now()}`,
content: 'Load test content'
}), { headers })
postCreateDuration.add(Date.now() - start)
check(createPost, {
'create post: status 201': (r) => r.status === 201,
})
errorRate.add(createPost.status !== 201)
}
sleep(0.3)
}
How to Set Performance Thresholds in k6?
Thresholds are the key mechanism for automatic pipeline failure. In the script options, we define acceptable boundaries: for example, http_req_duration: ['p(95)<500', 'p(99)<1000']. This means 95% of requests must be faster than 500ms, and 99% faster than 1 second. If a threshold is violated, the test finishes with an error, and CI stops deployment.
We also use custom metrics for specific business operations (e.g., post creation time) and set separate thresholds on them. That way, you know for sure that critical functionality is not degrading.
GitHub Actions Integration
# .github/workflows/performance.yml
name: Performance Tests
on:
push:
branches: [main, staging]
pull_request:
branches: [main]
jobs:
performance:
runs-on: ubuntu-latest
services:
postgres:
image: postgres:15
env:
POSTGRES_DB: testdb
POSTGRES_PASSWORD: testpass
ports: ['5432:5432']
options: >-
--health-cmd pg_isready
--health-interval 10s
steps:
- uses: actions/checkout@v4
- name: Start application
run: |
docker compose -f docker-compose.test.yml up -d api
npx wait-on http://localhost:3000/health --timeout 60000
- name: Run k6 smoke test
uses: grafana/[email protected]
with:
filename: tests/performance/api-smoke.js
flags: --out json=results.json
env:
BASE_URL: http://localhost:3000
K6_PROMETHEUS_RW_SERVER_URL: ${{ secrets.PROMETHEUS_URL }}
- name: Parse results
if: always()
run: |
# Show summary in PR comment
jq -r '.metrics | {
p95: .http_req_duration["p(95)"],
p99: .http_req_duration["p(99)"],
errors: .http_req_failed.rate
}' results.json
- name: Comment PR with results
if: github.event_name == 'pull_request'
uses: actions/github-script@v7
with:
script: |
const fs = require('fs')
const results = JSON.parse(fs.readFileSync('results.json'))
const p95 = results.metrics.http_req_duration['p(95)'].toFixed(0)
const errorRate = (results.metrics.http_req_failed.rate * 100).toFixed(2)
github.rest.issues.createComment({
issue_number: context.issue.number,
owner: context.repo.owner,
repo: context.repo.repo,
body: `## Performance Test Results\n\n| Metric | Value | Threshold |\n|--------|-------|-----------|\n| p95 latency | ${p95}ms | <500ms |\n| Error rate | ${errorRate}% | <1% |`
})
Baseline Comparison Between Deployments
#!/bin/bash
# scripts/compare-performance.sh
CURRENT_BRANCH=$(git branch --show-current)
BASELINE_BRANCH="main"
# Test current code
k6 run --out json=current.json tests/performance/api-smoke.js
# Switch to baseline
git stash
git checkout $BASELINE_BRANCH
docker compose up -d --build api
sleep 10
k6 run --out json=baseline.json tests/performance/api-smoke.js
# Comparison
node - <<'EOF'
const current = require('./current.json')
const baseline = require('./baseline.json')
const metrics = ['http_req_duration']
for (const m of metrics) {
const cp95 = current.metrics[m]['p(95)']
const bp95 = baseline.metrics[m]['p(95)']
const delta = ((cp95 - bp95) / bp95 * 100).toFixed(1)
if (cp95 > bp95 * 1.2) { // regression > 20%
console.error(`REGRESSION: ${m} p95 degraded by ${delta}%`)
process.exit(1)
}
console.log(`${m} p95: ${cp95}ms vs ${bp95}ms baseline (${delta}%)`)
}
EOF
# Return to current branch
git checkout $CURRENT_BRANCH
git stash pop
Artillery for Scenario Description
# tests/performance/user-journey.yml
config:
target: "{{ $processEnvironment.BASE_URL }}"
phases:
- duration: 60
arrivalRate: 5
rampTo: 20
name: "Ramp up"
- duration: 120
arrivalRate: 20
name: "Sustained load"
ensure:
thresholds:
- http.response_time.p95: 500
- http.request_rate: 15
scenarios:
- name: "Browse and purchase"
weight: 70
flow:
- get:
url: "/api/products"
expect:
- statusCode: 200
- post:
url: "/api/cart"
json:
productId: "{{ $randomInt(1, 100) }}"
quantity: 1
- name: "Search only"
weight: 30
flow:
- get:
url: "/api/search?q={{ $randomString(5) }}"
Load Testing Tool Comparison
| Feature | k6 | Artillery | Gatling |
|---|---|---|---|
| Scripting language | JavaScript | YAML | Scala/Java |
| Native CI integration | Yes (GitHub Actions, GitLab CI) | Yes (via NPM) | Yes (Maven/Gradle) |
| Built-in thresholds | Yes | Yes (via ensure) | Yes |
| Report generation | JSON, Prometheus, HTML | JSON, HTML | HTML (detailed) |
| Performance | High (Go) | Medium (Node.js) | High (JVM) |
| License | Open Source (AGPL) | Open Source (MPL) | Open Source (ALv2) |
Why Use k6 for CI?
k6 has built-in CI/CD support: it works as a CLI tool, requires no GUI, and results can be output to JSON for further processing. Unlike Gatling (requires Scala/Java and HTML report generation) or Artillery (YAML configs, but fewer metrics), k6 provides flexible metrics and simple integration with GitHub Actions via ready-made actions. For teams already using JavaScript, the entry threshold is minimal.
Process of Work
- Analysis — identify critical endpoints and user scenarios.
- Design — develop load scenarios, configure load profiles (ramp-up, sustained).
- Implementation — write scripts (k6, Artillery, or Gatling), integrate into CI.
- Testing — run smoke tests on staging, adjust thresholds.
- Deployment — include tests in pipeline, configure automatic PR comments.
Timelines and What's Included
Setting up a basic set of k6 smoke tests with thresholds and CI integration takes from 1 to 2 business days. If baseline comparison and custom metrics are needed, we extend to 3-5 days. The scope includes:
- Load test scripts (2-3 scenarios)
- Thresholds configuration
- Integration with GitHub Actions or GitLab CI (YAML pipeline)
- PR comment with test results
- Documentation for running and maintenance
- Team consultation on interpreting results
We guarantee that tests will not interfere with the main development process: they run in parallel and take no more than 5 minutes.
Why Choose Us
- Over 5 years of experience in load testing and CI/CD
- 50+ successful projects — from startups to enterprise
- We use only proven tools (k6, Grafana, Prometheus)
- We provide a guarantee on correct test operation for a month after implementation
Contact us to discuss your project: get a consultation on integrating Continuous Load Testing into your pipeline.







