invariantity looks at the structure in your data before choosing how to compress it. Different data compresses differently, so we measure by domain and publish the results below, including the worst one.
| No. | Data domain | Measured result | Remark |
|---|---|---|---|
| 01 | Structured data (JSON) | 16.8× | Full 912 MB GitHub Archive event dump, not a sample |
| 02 | Backups and version history | 71.4× | 20-version incremental backup set, cross-file dedup |
| 03 | vs. zstd, xz, bzip2, brotli (max level) | 1 loss (of 38) | 35 wins, 2 ties elsewhere; up to 2.20× tighter on the best case (Apache logs). The one loss is 0.76% on a 21 KB object file. |
| 04 | Documents (PDF) | 3.8–6.1× | Varies by how much of the PDF is already-compressed image data |
| 05 | Log files (Apache, HDFS, Linux, OpenSSH, Spark, Zookeeper, HealthApp) | 12.2–56.2× | 7 real production log formats from the LogHub benchmark set |
| 06 | Already-compressed binary media (JPEG, PNG) | ≈1.0× | The floor. Measured on all 24 Kodak PNGs + 4 JPEGs. On 10 of 24 PNGs, a specialized codec still finds a small residual gain (0.04–0.85%) that this build does not chase. |
All values are measured, not projected, logged run by run. Row 06 is the worst case and it is printed on purpose, including where it currently falls slightly short.
| File | Original | invariantity | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| bib | 111,261 | 23,197 | 4.80× | 4.05× (bzip2) |
| book1 | 768,771 | 208,220 | 3.69× | 3.31× (bzip2) |
| book2 | 610,856 | 136,991 | 4.46× | 3.88× (bzip2) |
| geo | 102,400 | 48,220 | 2.12× | 1.94× (brotli) |
| news | 377,109 | 101,746 | 3.71× | 3.34× (brotli) |
| obj1 | 21,504 | 9,472 | 2.27× | 2.30× (brotli, loss) |
| obj2 | 246,814 | 61,472 | 4.01× | 4.02× (xz, tie) |
| paper1 | 53,161 | 13,365 | 3.98× | 3.44× (brotli) |
| paper2 | 82,199 | 20,638 | 3.98× | 3.31× (brotli) |
| paper3 | 46,526 | 12,434 | 3.74× | 3.18× (brotli) |
| paper4 | 13,286 | 3,718 | 3.57× | 3.10× (brotli) |
| paper5 | 11,954 | 3,720 | 3.21× | 2.93× (brotli) |
| paper6 | 38,105 | 9,796 | 3.89× | 3.42× (brotli) |
| pic | 513,216 | 33,909 | 15.13× | 12.87× (xz) |
| progc | 39,611 | 10,038 | 3.95× | 3.41× (brotli) |
| progl | 71,646 | 12,171 | 5.89× | 5.12× (brotli) |
| progp | 49,379 | 8,665 | 5.70× | 5.00× (brotli) |
| trans | 93,695 | 13,672 | 6.85× | 6.08× (brotli) |
| File | Original | invariantity | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| alice29.txt | 152,089 | 36,262 | 4.19× | 3.52× (bzip2) |
| asyoulik.txt | 125,179 | 33,809 | 3.70× | 3.16× (bzip2) |
| cp.html | 24,603 | 6,313 | 3.90× | 3.57× (brotli) |
| fields.c | 11,150 | 2,295 | 4.86× | 4.10× (brotli) |
| grammar.lsp | 3,721 | 1,026 | 3.63× | 3.31× (brotli) |
| kennedy.xls (spreadsheet) | 1,029,744 | 26,706 | 38.56× | 19.85× (xz) |
| lcet10.txt | 426,754 | 92,164 | 4.63× | 3.96× (bzip2) |
| plrabn12.txt | 481,861 | 131,363 | 3.67× | 3.31× (bzip2) |
| ptt5 (fax scan) | 513,216 | 33,909 | 15.13× | 12.87× (xz) |
| sum | 38,240 | 9,516 | 4.02× | 4.02× (xz, tie) |
| xargs.1 | 4,227 | 1,324 | 3.19× | 2.89× (brotli) |
| File | Original | invariantity | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| reymont (Silesia corpus, Polish novel) | 6,627,202 | 1,085,999 | 6.10× | 5.32× (bzip2) |
| attention.pdf (public arXiv paper) | 2,215,244 | 584,284 | 3.79× | 2.15× (brotli) |
| File | Original | invariantity | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| Apache_2k.log | 171,239 | 3,046 | 56.22× | 25.54× (brotli) |
| OpenSSH_2k.log | 225,216 | 5,222 | 43.13× | 23.10× (xz) |
| Linux_2k.log | 216,485 | 6,728 | 32.18× | 21.62× (xz) |
| Spark_2k.log | 196,268 | 6,688 | 29.35× | 21.63× (xz) |
| HealthApp_2k.log | 187,456 | 7,388 | 25.37× | 15.49× (brotli) |
| Zookeeper_2k.log | 279,891 | 11,450 | 24.45× | 18.87× (bzip2) |
| HDFS_2k.log | 287,848 | 23,505 | 12.25× | 7.02× (brotli) |
| File | Original | invariantity | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| kodim01/05/12/18.jpg (4 JPEGs, avg) | 161,864 avg | 161,880 avg | 1.00× | ≤1.00× (brotli, no meaningful difference) |
| 14 of 24 PNGs (avg, no meaningful gap) | 621,477 avg | 621,493 avg | 1.00× | ≤1.00× (16-byte store header only) |
| kodim02.png | 617,995 | 618,011 | 1.00× | 1.00× (xz, 0.22% smaller) |
| kodim04.png | 637,432 | 637,448 | 1.00× | 1.00× (xz, 0.28% smaller) |
| kodim08.png | 788,470 | 788,486 | 1.00× | 1.01× (bzip2, 0.85% smaller) |
| kodim09.png | 582,899 | 582,915 | 1.00× | 1.00× (xz, 0.22% smaller) |
| kodim10.png | 593,463 | 593,479 | 1.00× | 1.00× (xz, 0.05% smaller) |
| kodim15.png | 612,582 | 612,598 | 1.00× | 1.00× (xz, 0.04% smaller) |
| kodim18.png | 780,947 | 780,963 | 1.00× | 1.01× (xz, 0.73% smaller) |
| kodim19.png | 671,476 | 671,492 | 1.00× | 1.01× (xz, 0.80% smaller) |
| kodim22.png | 701,970 | 701,986 | 1.00× | 1.00× (xz, 0.36% smaller) |
| kodim24.png | 706,397 | 706,413 | 1.00× | 1.00× (xz, 0.17% smaller) |
Once invariantity recognizes data as already compressed, it stores it as-is rather than spending time trying to shrink it further. On the 4 JPEGs and 14 of the 24 PNGs that costs nothing beyond a 16-byte header. On the other 10 PNGs, a general-purpose codec still finds a small residual gain (0.04–0.85%) that this build does not chase — we are looking at whether a cheap second-pass check is worth adding.
Petabytes of user files and artifacts, most of them structured or repetitive. Across 38 real files tested at max settings against standalone zstd, xz, bzip2, and brotli, invariantity won 35 and tied 2, with a single 0.76% loss on one file — so switching is rarely a downside on the files those tools already handle well.
Version history is the most redundant data there is, which is why it is our best measured result at 71.4x on a 20-version backup set. Longer retention in the same footprint, and restores move less data over the wire.
JSON events are highly structured, measured at 16.8x on a full 912 MB dump, not a favorable sample. Application logs are even more repetitive: up to 56.2x on real Apache logs. Compressing them well cuts both the storage bill and the egress bill, at every hop where the data sits or moves.
On already-compressed binary media such as JPEG and PNG, invariantity measures close to 1.0x across the standard Kodak reference set (24 PNGs, 4 JPEGs). That is the honest floor, and it is on this page because a benchmark table without a worst case is an advertisement, not a measurement — including the part where, on 10 of those 24 PNGs, a general-purpose codec still finds a small residual gain (up to 0.85%) that our current build does not chase. PDFs turned out not to belong in this floor: two real documents we tested measured 3.8x and 6.1x, since a PDF's own internal structure varies far more than an already-DCT-compressed image does. If most of your data is genuinely already-compressed pixels, we will tell you the gains are modest before you spend a day integrating anything.
zstd and xz are excellent general-purpose compressors and we do not pretend otherwise. invariantity differs in one way we can state without disclosing the method: it looks at the structure in your data before choosing how to compress it, then still checks its choice against zstd, xz, bzip2, and brotli at their own maximum settings. Across 38 real files tested this way, it won 35 outright, tied 2, and lost 1 (a 0.76% gap on a 21 KB object file) — up to 2.20x tighter than the best single alternative on the best case. A full methodology will be published with the evaluation build so you can reproduce the comparison on your own data.
Not today. The core is proprietary. We plan a public evaluation build so you can verify these numbers on your own data before committing to anything, which we think matters more than reading our source.
On genuinely already-compressed binary media, such as JPEG and PNG, you get close to 1.0x, the floor in the table above (on 10 of 24 PNG reference images, a specialized codec still finds a small 0.04–0.85% residual gain our current build doesn't chase). PDFs are a partial exception, since they often carry their own compressible structure: two real PDFs we tested measured 3.8x and 6.1x. If your data is mostly already-compressed pixels rather than documents, the size gains alone probably do not justify a migration, and we will say so when you tell us what you store.
We are pre-launch. Private pilots come first, prioritized by data type, since the gains depend on what you store. Leave your details below and tell us what your data looks like; that is genuinely the fastest route in.
We reply with what the measured numbers say about your data type, including when the answer is that the gains would be modest. No drip campaign follows.