Spitfire

Version: 0.32.0This documentation is for Spitfire 0.32.0.

MQTT, RabbitMQ, NATS and AWS SQS

This page covers four messaging protocols: MQTT for IoT and device traffic, RabbitMQ (AMQP 0-9-1) for queue-based services, NATS for events between services and JetStream, and AWS SQS and SNS for Amazon queues. For Kafka see its own page: Kafka.

MQTT

What it does

Every VU is its own MQTT client, like a device: it connects on its first call, keeps its subscriptions for the run and reconnects on the next call if the connection is lost. There are three actions:

  • Publish: sends a message to a topic. With QoS 1/2 the time runs until the broker acknowledges.
  • Wait for a message (subscribe): the VU subscribes to the topic once; every iteration waits for the next message.
  • Request/reply: subscribes to the response topic, publishes the message and waits for the reply; the time is the round trip.

When to use it

To measure the capacity and latency of a broker (Mosquitto, EMQX, HiveMQ, the RabbitMQ MQTT plugin…) when thousands of devices connect at once and send telemetry.

Creating the connection

  1. Connections → Add connection, Type: MQTT.
  2. Brokers (comma-separated): tcp://broker:1883, for TLS ssl://broker:8883, over WebSocket ws://broker:8083/mqtt. Required. Don't put a password in the address.
  3. Client ID: a template; default spitfire-{{__VU}}. Each VU needs a unique client ID; when two clients connect with the same ID, the broker drops one of them.
  4. Username and Password (a secret).
  5. Keep-alive (s): default 30.
  6. Clean session: on by default. When off, the broker keeps the session and persistent subscriptions.
  7. MQTT version: 3.1.1 (the default) or 5. MQTT 5 works with tcp:// and ssl:// broker addresses only (no WebSocket).
  8. Expand TLS for TLS.
  9. Save, Test.

Adding a step

  1. Add step → Protocol: MQTT → pick the Connection.
  2. Pick an Action.
  3. Type the Topic. Wildcards (+, #) are only valid for Wait for a message (subscribe); on publish the save fails with wildcards (+ #) are only valid for subscribe.
  4. Write the Payload; {{variables}} are allowed.
  5. Pick QoS (0, 1 or 2) and retain if needed.
  6. For Request/reply, write the Response topic.
  7. Wait (default 10s): for Wait for a message and Request/reply, the step fails when it runs out.
json
[
  {"id": "telemetry", "name": "Telemetry", "protocol": "mqtt", "connection": "iot-broker",
   "mqtt": {"action": "publish", "topic": "devices/{{__VU}}/telemetry",
     "payload": "{\"temp\": {{$randInt 18 30}}}", "qos": 1}},
  {"id": "command", "name": "Wait for command", "protocol": "mqtt", "connection": "iot-broker",
   "mqtt": {"action": "subscribe", "topic": "devices/{{__VU}}/commands", "qos": 1, "wait": "10s"}}
]

MQTT 5

When the connection's MQTT version is 5, publish and request steps also have:

  • User properties: key/value pairs sent with the message (templates work: tenant = {{tenant}}). An incoming message's user properties reach checks and extractors as headers.
  • Content type and Message expiry: when it runs out, the broker no longer delivers the message.
  • Request/reply: the request carries the response topic and fresh correlation data on every call; only the reply with that correlation data is accepted. A late reply to an earlier request does not answer this one (on 3.1.1 the first message on the response topic is taken).

Metrics it produces

req_duration (publish: sending with QoS 0, until the broker acknowledges with QoS 1/2; subscribe: waiting for the next message; request/reply: the round trip), req_failed, reqs, data_sent, data_received.

Tip

To connect like a new device on every iteration, tick Close connections between iterations in the Options tab; it applies to MQTT clients too. Handy for measuring the cost of connecting.

RabbitMQ (AMQP)

What it does

Publishes a message to an exchange, takes a message from a queue, or does request/reply (RPC). The VUs of a runner share one connection (each with its own channels); optionally every VU opens its own connection. Every publish waits for its publisher confirm, so the measured time is the broker's.

When to use it

To see how the workload that flows through RabbitMQ (orders, notifications, job queues) affects the broker and its consumers.

Creating the connection

  1. Connections → Add connection, Type: RabbitMQ (AMQP).
  2. URL: amqp://rabbit:5672/, or with a vhost amqp://rabbit:5672/orders. For the TLS port usually amqps://rabbit:5671/. Required; don't put the password in the URL.
  3. Username and Password (a secret).
  4. Separate connection per VU: when ticked, each VU opens its own TCP connection (to imitate many clients). Default: one connection per runner, a channel per VU.
  5. TLS section for TLS.
  6. Save, Test.

Adding a step

  1. Add step → Protocol: RabbitMQ (AMQP) → pick the Connection.
  2. Action:
    • Publish: write the Exchange (empty means the (default) exchange) and the Routing key. A message no queue receives (unroutable) is a failure.
    • Consume from queue: write the Queue. The queue must exist; each iteration takes one message and acknowledges it (ack).
    • RPC (request/reply): the request is sent with direct reply-to and the matching reply is awaited.
  3. Write the Message; add Content-Type and Headers if needed.
  4. Tick Persistent message to keep the message across a broker restart.
  5. Wait (default 10s): for consume and RPC.
json
[
  {"id": "publish", "name": "Order event", "protocol": "amqp", "connection": "rabbit",
   "amqp": {"action": "publish", "exchange": "orders", "routingKey": "orders.created",
     "body": "{\"id\": \"{{$uuid}}\"}", "contentType": "application/json", "persistent": true}},
  {"id": "take", "name": "Take from queue", "protocol": "amqp", "connection": "rabbit",
   "amqp": {"action": "consume", "queue": "orders.created", "wait": "10s"}},
  {"id": "rpc", "name": "Ask price", "protocol": "amqp", "connection": "rabbit",
   "amqp": {"action": "rpc", "routingKey": "pricing.rpc", "body": "{\"sku\": \"A-1\"}", "wait": "5s"}}
]
Warning

Consume from queue takes the message from a real queue and acknowledges it; that message no longer reaches your real consumer. Don't use it on production queues; bind a separate queue for testing.

Metrics it produces

req_duration (publish: until the publisher confirm; consume: waiting for the message; RPC: the round trip), req_failed, reqs, data_sent, data_received.

NATS

What it does

Core NATS and JetStream steps. The VUs of a runner share a few connections (NATS carries a lot over one connection); a connection per VU is an option. There are five actions:

  • Publish: core publish. No acknowledgement; the time is writing the message to the client's buffer.
  • Wait for a message (subscribe): the VU subscribes to the subject once (wildcards: orders.*, orders.>); every iteration waits for the next message. With a Queue group, messages are shared among the VUs in the group.
  • Request/reply: the message goes with a reply inbox and the reply is awaited; the time is the round trip. With nobody listening the error is no_responders.
  • Publish to JetStream: the time is to the stream's acknowledgement (PubAck). With no stream on the subject the error is no_stream.
  • Consume from JetStream (pull): one message is fetched from a durable pull consumer and acknowledged. The consumer is shared by the runner's VUs, so every message is handled once.

When to use it

To measure the capacity and end-to-end latency of a NATS server, cluster or leaf node for event streams between microservices, request/reply services and JetStream queues.

Creating the connection

  1. Connections → Add connection, Type: NATS.
  2. Servers (comma-separated): nats://nats:4222; for TLS tls://nats:4222 or the TLS section. Required. Don't put a password in the address.
  3. Fill in only one way to authenticate: Username + Password, Token, NKey seed or .creds file (JWT + seed; paste the file's content). All are stored encrypted.
  4. JetStream domain: only with leaf nodes / domains.
  5. Shared connections (per runner): 4 by default. Connection per VU opens one for every VU (to try the server's connection limits).
  6. Save, Test.

Adding a step

  1. Add step → Protocol: NATS → pick the Connection.
  2. Pick the Action and write the Subject. Wildcards only work when waiting for messages and consuming from JetStream; publishing to one is refused.
  3. For publish and request, write the Message; {{variables}} work. Headers can be added.
  4. Consume from JetStream needs the Stream. The Consumer is spitfire by default; when it does not exist, a durable consumer is created that starts from new messages, acknowledges explicitly and filters on the subject. An existing consumer is used as it is; its settings are not changed.
  5. Wait (10s by default): waiting for a message, a request and a JetStream consume fail when it runs out.
json
[
  {"id": "place", "name": "Order", "protocol": "nats", "connection": "bus",
   "nats": {"action": "jsPublish", "subject": "orders.{{__VU}}", "payload": "{\"vu\": {{__VU}}}"}},
  {"id": "handle", "name": "Handle the order", "protocol": "nats", "connection": "bus",
   "nats": {"action": "jsConsume", "subject": "orders.>", "stream": "ORDERS", "consumer": "spitfire", "wait": "5s"}},
  {"id": "price", "name": "Ask the price", "protocol": "nats", "connection": "bus",
   "nats": {"action": "request", "subject": "price.get", "payload": "{{sku}}", "wait": "2s"}}
]

Careful: on work-queue streams an acknowledged message is deleted. Don't use a real service's consumer name; Spitfire would take and acknowledge that service's messages.

End-to-end latency

Publishes are stamped with a Spitfire-Ts header (turn it off with Don't stamp). Wait for a message and Consume from JetStream steps measure nats_e2e_latency for messages published in the same run by the same runner: both ends read one clock, so there is no clock skew. Messages published on another runner are not measured (as with Kafka).

Metrics it produces

req_duration (by action: writing to the buffer, PubAck, round trip or waiting for the message), nats_e2e_latency, req_failed, reqs, data_sent, data_received. A threshold example: p(95)<50 on nats_e2e_latency.

AWS SQS and SNS

What it does

Sends messages to Amazon SQS queues, receives and deletes messages from a queue, and publishes to SNS topics. It also works with compatible services such as LocalStack and ElasticMQ. There are three actions:

  • Send (SendMessage): the time is to SQS's answer.
  • Receive and delete (ReceiveMessage): a message is awaited with long polling (in polls of up to 20s, until the wait runs out) and deleted when it arrives. The time is until the message arrives. With Don't delete the message it returns to the queue after the visibility timeout.
  • Publish to SNS (Publish): the time is to SNS's answer.

Creating the connection

  1. Connections → Add connection, Type: AWS (SQS, SNS).
  2. Region: e.g. eu-central-1. Required.
  3. Access key ID and Secret access key (stored encrypted); Session token for temporary credentials.
  4. Or Use the machine's credentials: the controller and every runner use their own environment, profile or IAM role (EC2, ECS, EKS); no key is stored. Only an installation admin can turn this on, because everyone who uses the connection acts with the installation's AWS identity.
  5. Endpoint: only for LocalStack (http://localstack:4566), ElasticMQ or a VPC endpoint.
  6. Save, Test (tries to list queues; when listing is not allowed, accepted credentials still count as success).

A least-privilege IAM policy: sqs:SendMessage, sqs:ReceiveMessage, sqs:DeleteMessage, sqs:GetQueueUrl on the queue; sns:Publish on the topic.

Adding a step

  1. Add step → Protocol: AWS SQS / SNS → pick the Connection.
  2. Pick the Action. In Queue write the queue's name or URL (a name is turned into its URL once); for SNS the SNS topic ARN.
  3. For sending and publishing, the Message and optional Message attributes. On FIFO queues and topics, Message group ID and Deduplication ID (e.g. {{$uuid}}; leave it empty with content-based deduplication). Delay is 0–15 minutes.
  4. For receiving, the Wait (10s by default).
json
[
  {"id": "queue", "name": "Queue the order", "protocol": "sqs", "connection": "aws-prod",
   "sqs": {"action": "send", "queue": "orders", "body": "{\"orderId\": {{__ITER}}}"}},
  {"id": "take", "name": "Handle the order", "protocol": "sqs", "connection": "aws-prod",
   "sqs": {"action": "receive", "queue": "orders", "wait": "10s"}}
]

Careful: the receive step deletes the message. Don't receive from a queue a real service works on; you would consume its messages. Use a queue of its own for the test. SQS charges per request: a load test shows up on your bill.

End-to-end latency

Sent messages get a spitfire-ts attribute (turn it off with Don't stamp). Receive and delete measures sqs_e2e_latency for messages sent in the same run by the same runner. For SNS → SQS, with raw message delivery on for the subscribed queue the attributes are kept and the latency is measured too.

Metrics it produces

req_duration, sqs_e2e_latency, req_failed, reqs, data_sent, data_received. Error types: not_found (no such queue), auth (credentials or permission), throttled (AWS throttling), timeout.

Common problems

Symptom: MQTT Test fails with connection refused or not authorized. Cause: Wrong broker address/port, or the username/password is refused. Fix: Check the scheme (tcp://, ssl://, ws://) and the port; type the password again.

Symptom: MQTT VUs keep disconnecting and reconnecting. Cause: Client IDs collide (e.g. a fixed Client ID); the broker drops the older connection with the same ID. Fix: Use {{__VU}} in the Client ID (default spitfire-{{__VU}}).

Symptom: Wait for a message times out every time. Cause: No message arrives on the topic, or the topic name/wildcard is wrong. Fix: Check the topic; add a publish step to the same test; raise the Wait.

Symptom: RabbitMQ ACCESS_REFUSED - Login was refused. Cause: Wrong username/password, or the user has no access to the vhost. Fix: Type the password again; give the user permissions on the vhost in the URL.

Symptom: The publish step fails with publish: unroutable: 312 NO_ROUTE. Cause: No queue bound to the exchange matches the routing key. Fix: Check the exchange name and the routing key; create the queue binding.

Symptom: Consume from queue fails with NOT_FOUND - no queue. Cause: The queue does not exist; Spitfire does not create queues. Fix: Create the queue in RabbitMQ beforehand.

Symptom: A NATS step fails with no_responders. Cause: Nothing listens on the request subject. Fix: Check the subject and that the service is up; if the service is in a queue group, check its name.

Symptom: Publish to JetStream fails with no_stream. Cause: No stream covers the subject. Fix: Create the stream (nats stream add) or match the subject to the stream's subjects.

Symptom: Consume from JetStream times out every time. Cause: The consumer starts from new messages and nothing is published in the test, or the filter subject does not match. Fix: Add a Publish to JetStream step to the test, or create the consumer beforehand (--deliver all) and write its name in Consumer.

Symptom: An SQS step fails with not_found. Cause: The queue name is wrong, or the queue is in another region or account. Fix: Check the connection's Region; for a queue in another account write its URL, not its name.

Symptom: An SQS or SNS step fails with auth. Cause: The key is wrong or expired, or the IAM policy does not allow the action. Fix: Add the action named in the error (e.g. sqs:DeleteMessage) to the policy.