Spitfire
Version comparisonat the same time

Is the new release slower? Find out before it goes live.

Run the same test, with the same load, at the same time, against the environment running the old release and the one running the new release. Every runner splits its load evenly between the two, so network and timing differences do not skew the result. At the end Spitfire shows the difference step by step and says it plainly: better, no difference, or worse.

spitfire.local/comparisons/…
Spitfire: a finished comparison run. Verdict: worse; checkout p95 from 219 ms to 268 ms (+22%, 95% confidence interval ±2.4%), search from 150 ms to 130 ms (−13%, ±1.7%). Below, both releases' request rate and p95 charts.
A finished comparison run: the verdict card, the largest differences with their confidence intervals, and both releases' charts. Real screenshot.

When to use it

Before a release

The new release runs next to the old one under the same load before it goes live.

An infrastructure change

A new database version or a new server type: the old and the new infrastructure meet the same test.

A library or framework upgrade

See where performance goes when the framework, the ORM or the runtime version changes.

A configuration change

Connection pool size or a cache setting: measure the change instead of guessing.

How it works

  1. STEP 1

    Pick the environments

    From the test's Environments tab pick the one running the old release as A and the one running the new release as B, and write the version labels.

  2. STEP 2

    Equivalence check

    Spitfire compares the two environments and warns about differences; you go on by confirming or fixing them.

  3. STEP 3

    One run, both at once

    The same test runs against both environments at the same time with the same load; the live screen shows both arms side by side.

  4. STEP 4

    Verdict

    At the end, a step-by-step difference table, a confidence interval for every difference, and the verdict: better, no difference, worse, or inconclusive while the measurement is not clear yet.

spitfire.local/comparisons/…
Spitfire: the comparison screen while the run is going; the v2.3 · test and v2.4 · dev series fill side by side on the request rate and p95 charts.
While it runs: both releases' series fill side by side on every chart. Real screenshot.
spitfire.local/comparisons/…
Spitfire: the per-step difference table (p50, p95, p99, errors, requests/s, a confidence interval for every difference and the step verdict) and a separate comparison for every step of the load.
When it ends: the per-step difference table and a separate comparison for every step of the load. Real screenshot.

Like for like

A comparison means something only when both sides are measured under the same conditions.

Evenly split load

Every runner splits its virtual users evenly between the two arms, so a slowdown on the runner (CPU, network, GC) hits both alike. Load steps change on both arms at the same moment, and the location split is the same.

Pre-check

Before it starts, a light health request goes to both environments: response time, network latency from the runners, TLS and HTTP version are compared, and a clear difference gets a warning.

Shared infrastructure warning

When both environments point at the same database or cache (same host:port), or you marked them so, it warns: the two arms compete for one resource and the difference can hide.

A/A calibration

A short calibration run while both environments run the same release measures how they differ on their own; later comparisons show that difference apart from the result.

Version endpoint

Give an environment a version address (say GET /version) and both arms' versions are read and recorded at the start; if both report the same version, it warns.

Only one environment?

In sequential mode the two releases take turns on the same environment: A, B, A, B (2 rounds by default, each the same length). You switch the release; between rounds Spitfire waits and tells you with a run.switch_needed webhook; you confirm in the UI, the CLI or with an API call, and it sees the switch itself when a version endpoint is set. Taking turns spreads time-bound effects like cache warm-up and daily traffic over both releases; the verdict card says 'sequential comparison' and the uncertainty band is wider.

A release gate in CI

Give spitfire cloud run --compare in the pipeline. If the new release is worse, the step exits with code 98 and the deployment stops.

GitHub Actions

- name: Spitfire version comparison
  run: spitfire cloud run "Checkout flow" --compare A=test,B=dev --label-a v2.3 --label-b $GITHUB_SHA
  env:
    SPITFIRE_URL: https://spitfire.example.com
    SPITFIRE_TOKEN: ${{ secrets.SPITFIRE_TOKEN }}

GitLab CI

version-comparison:
  script:
    - spitfire cloud run "Checkout flow" --compare A=test,B=dev --label-a v2.3 --label-b $CI_COMMIT_SHORT_SHA
  # SPITFIRE_URL, SPITFIRE_TOKEN: Settings → CI/CD → Variables (token masked)
Exit codeMeaning
0Better or no difference. "Inconclusive" exits 0 too; 98 with --fail-on-inconclusive.
98The comparison says worse.
97The comparison is invalid: an arm dropped, or with --strict the equivalence check failed.
99A threshold failed (as today, k6-compatible).

How to read the result

Which plans

PlanVersion comparisonFrom CI and scheduled
Free1 a month—
Growth yearly✓✓
Growth monthly✓—
Growth 1 month (one-time)✓—
Scale yearly✓✓
Scale monthly✓—
Scale 1 month (one-time)✓—
Enterprise yearly✓✓

In a comparison run the virtual user, requests/s and runner limits apply to both arms together: each arm uses half.

Frequently asked

How do I know whether the new release is slower than the old one?
With a comparison run. Pick the environment running the old release as A and the one running the new release as B. Spitfire runs the same test against both at the same time, and every runner splits its load evenly between them. The result shows the difference and the verdict for every step: better, no difference, or worse.
My two environments are not identical. Can I still trust the result?
Before it starts, Spitfire compares the two environments and warns about the differences it sees. To measure how the environments differ on their own, run a calibration with the same release on both; later comparisons show that difference apart from the result.
Why not run twice and compare the two runs?
Between two runs at different times the network, the caches and the load on shared infrastructure can change. Running both at once spreads those outside effects evenly over the two sides; what is left is the difference the release makes.
I have one test environment. Can I still use it?
Yes. In sequential mode the two releases take turns on the same environment (A, B, A, B); each time you switch the release, Spitfire moves on to the next round.
Can CI stop the deployment when the new release is slower?
Yes. Give spitfire cloud run the --compare flag; if the new release is worse the step exits with 98 and the deployment stops.