Here we show how to stress test a Cassandra cluster using the cassandra-stress tool.
What this tool does is run inserts and queries against a table that it generates by itself or an existing table or tables.
(This article is part of our Cassandra Guide. Use the right-hand menu to navigate.)
The Basic Flow
The test works like this:
- Insert random data. You specify the length of text fields or random numeric values by selecting a fixed value or a statistical distribution such as a normal, uniform, or other distribution. The normal distribution are values drawn from some mean and standard deviation, i.e. the familiar bell curve. The uniform distribution is random numbers drawn from a range like 1,2,3,….
- Run select statements using the values generated.
- Calculate the time in milliseconds to run each operation.
- Calculate the mean time, standard deviations number of garbage collections etc. for each iteration. This gives you an average and the bell curve so you can see how widely your operations are disbursed. It gives you graphs over time so you can see whether performance degrades over time. Lots of operations far from the mean indicate a high level of variance. This could point to items you need to tune, such as indexes, partitions, add more memory, etc.
Test Setup
You need a Cassandra cluster. If you do not have one yet follow these instructions.
We need to create a keyspace, table, and index and to create a stress test configuration file in YAML format.
Now deactivate Python 2.7 virtual environment, if you are using that, then run cqlsh. Paste in the following Cassandra SQL.
CREATE KEYSPACE Library
WITH REPLICATION = { 'class' : 'SimpleStrategy', 'replication_factor' : 3 };
CREATE TABLE Library.book (
ISBN text,
copy int,
title text,
PRIMARY KEY (ISBN, copy)
);
create index on library.book (title);
Now, create a text file named stress-books.yaml and paste the following into it. This is a YAML format, so don’t mess up the indentation.
Below we explain the fields.
keyspace: library table: book columnspec: - name: text size: uniform(5..10) population: uniform(1..10) - name: copy cluster: uniform(20..500) - name: title size: uniform(5..10) insert: partitions: fixed(1) select: fixed(1)/500 batchtype: UNLOGGED queries: books: cql: select * from book where title = ? fields: samerow
The field names are:
| keyspace | Name of existing keyspace. You could also put SQL here to create one if it does not exist. |
| table | Name of existing table. You could also put SQL here to create one if it does not exist. |
| insert, queries | These are the functions we will call. The insert does the insert and the queries/books does the query we put there. It runs queries against the values inserted in the batch it just ran. |
| columnspec: – name: text size: uniform(5..10) | There is one columnspec for each value we want to populate. The size is the length of the field. uniform(5..10) means to generate a text string from 5 to10 characters. If it was a numeric field it would create random numbers in that range. |
| insert: partitions: fixed(1) select: fixed(1)/500 batchtype: UNLOGGED | This means to insert a fixed number of rows in each partition in each batch. We explained batch operations here. |