Filename: 370-compression-dict-transport.md
Title: Even smaller consensus downloads with Compression Dictionary Transport
Author: Nick Mathewson
Created: 22 September 2026
Status: Open
Introduction
Current versions of Tor reduce the amount of bandwidth used for consensus downloads by enabling clients to request a diff from a previous consensus; these diffs are then compressed.
This use of diffs introduces complexity into client and relay implementations, and additionally delivers suboptimal performance. Notably, both diffs and compression involve performing a series of approximate longest-common-subsequence hunts. In this proposal, we describe a standards-based mechanism for reducing consensus download bandwidth even further, and offloading the complexity into our compression library.
The building blocks
Let's talk about the pieces we can use to replace the use of ed diffs.
Compression and dictionaries
LZ-based lossless compression schemes work, roughly, by recording places where the document copies from itself. When a piece of data would be duplicated, the compressed datastream instead says "insert a copy of the N bytes that appeared M bytes ago."
Some modern LZ-based compression schemes (including zstd and brotli) support the use of an "external compression dictionary": if both parties share a single "dictionary" document, they can additionally say "insert a copy of the N bytes starting at position M in the dictionary."
Significantly, using dictionaries in this way can replace the use of compressed diffs, and often gives better compression. (I'll discuss results later.)
RFC9842 and Compression Dictionary Transport
RFC9842 describes a set of extensions to HTTP that parties can use to negotiate compression dictionaries.
In brief,
when serving a document that can be used as a dictionary,
the server adds a Use-As-Dictionary header to its response,
telling the client
which documents it can use the document as dictionary for.
The client later provides an Available-Dictionary header
in subsequent requests, telling the server that it has
a dictionary with a given SHA-256 digest.
Consensus documents, signature variance, and dictionary types
Unfortunately, we allow some variance in ways in which a given consensus can be represented. Only the signed portion of a given consensus is invariant: it is possible for caches to have different sets of signatures, depending on whether the authorities exchanged signatures successfully or not.
Additionally, we do not currently (AFAICT!) specify an order in which the signatures must appear.
Moreover, a client will accept any consensus document that is signed by a threshold of the authorities that the client recognizes. The client will ignore signatures from unrecognized authorities, and will ignore missing signatures, so long as not too many are missing.
Taken together, this means that we cannot rely on the SHA-256 digest of the signed consensus as a unique identifier for the dictionary.
Fortunately, RFC9842 specifies a "type" field
in the Use-As-Dictionary header that explains
how to use the resulting dictionary.
The only type it specifies is "raw",
meaning that we use the document as-is,
as "an unformatted blob of bytes suitable for any compression scheme".
So if we specify a new "type", we can make RFC9842 work for Tor's needs.
Application to Tor
Here we use "cache" to mean any party serving consensus documents. We use "client" to mean any party fetching consensus documents.
Unless otherwise stated, all requirements from RFC9842 are in effect.
When serving consensus documents, caches SHOULD
return a Use-As-Dictionary header.
If they do, the "id" field for the header SHOULD be
the consensus flavor, followed by a dash,
followed by the valid-after date of the consensus
in YYYY-MM-DDTHH:MM:SS format.
Clients MUST NOT rely on this format,
and MUST return this ID verbatim.
Caches MUST also include the "type" field
with the value tor-signed-consensus.
When receiving a Use-As-Dictionary header,
clients SHOULD ignore it
(and therefore not use the consensus as a dictionary)
unless it is associated with a request for a consensus,
and unless the "type" is tor-signed-consensus.
Clients MUST NOT use a consensus document as a dictionary unless they have found that it is well formed and correctly signed.
To derive the dictionary from the signed consensus,
the tor-signed-consensus type means that
clients (and relays) will strip all unsigned material from the consensus.
This means that the dictionary consists of all of the consensus up to and including the first space after the first
directory-signaturetoken.
Clients SHOULD use these dictionaries when
downloading subsequent consensus documents,
using the Available-Dictionary and Dictionary-ID headers.
(The SHA256 hash is a hash of the dictionary,
not of the signed consensus.)
As an extension to RFC9842, clients SHOULD use a dictionary when fetching a consensus from any cache, even if that cache was not the one that served the dictionary.
Because all caches fetch the same set of consensus documents, clients can reasonably expect that all caches will have the same consensuses.
Clients and caches SHOULD advertise support for the "dcz" content encoding whenever a dictionary may be used, per RFC9842 section 6.
Additionally, clients and caches MAY support brotli compression, via the "br" and "dcb" encodings.
Given the performance results, we may want to switch those, or support only one.
Parties MUST NOT support dictionary compression for anonymized HTTP connections, such as those used for hidden service descriptors.
Parties MUST NOT use compression settings that require the decompressor to use more than (approximately) 16 MiB of RAM.
Appendices
Appendix A: Compression improvements
Below are the size of compressed diffs (or dictionary-based compressions) between pairs of consensuses taken from the CollecTor archives in June 2026. The "time difference" shows the difference between the pairs of consensuses' Valid-After times. Other values are in bytes.
For reference a typical consensus from this period has an xz-6 compressed size of around 730 KiB.
For xz, I'm using "-6" compression, since that's what we use in C tor to limit the memory needed for decompression.
With zstd, I've used "wlog=24", to allow a ~16 MiB requirement for decompression. (By default it only allows less memory.) I've set the zstd compression quality to "12", since that is the highest compression quality that is still faster (in my experiments) than xz-6. (The zstd quality parameter can go up to 19. In my experiments, zstd-12 is about as fast for compression as ed+xz-6.)
Brotli's default window size is 24 bits, so it also obeys the same ~16 MiB decompression requirement. I've chosen "9" as the highest brotli compression level that is still faster than xz-6. (The brotli compression parameters can go up to 11. In my experiments, br-9 is about twice as fast for compression as ed+xz-6.)
| Time difference | ed+zstd | ed+xz | dict+zstd | dict+brotli |
|---|---|---|---|---|
| 1 hour | 10509 | 9816 | 9244 | 8945 |
| 6 hours | 31715 | 27124 | 27513 | 26152 |
| 12 hours | 49606 | 42448 | 42971 | 40795 |
| 1 day | 69577 | 58864 | 56065 | 54322 |
| 2 days | 107450 | 93088 | 87386 | 85807 |
| 1 week | 216307 | 197384 | 180754 | 181109 |
| 2 weeks | 320610 | 298948 | 271731 | 271285 |
Note that dict+zstd outperforms ed+xz in most cases, by around 5-10%. In the worse cases, it is about 1% worse. Dict+brotli outperforms ed+xz in all cases, by around 5-10%. (And remember, the compression is significantly faster at this level.)
Not pictured here: If we pre-compute dict-compressed values, and are willing to accept higher compression costs, we can make the compressed outputs smaller while retaining the same required memory usage for the decompressor. In the most expensive case, dict+zstd-19 is about 5x slower than ed+xz-6, and gives compression performance between 10-12% better; whereas dict+brotli-11 is about 10x slower than ed+xz-6 but gives compression performance around 15% better.
We could look at the distribution of actual time distances in offered consensuses for diffs, and use those to determine what we want to pre-compute at a higher compression level, and what we want to generate on the fly at a lower compression algorithm.
Appendix C: Walking onions
With walking onions, clients will only download a small "Parameters Document", but relays will still have to download and serve a larger "ENDIVE". We can use compression dictionaries for both, if we choose to do so.