Replies: 3 comments 1 reply
|
Hi! It's unlikely we'll switch away from x/crypto/chacha20poly1305 specifically for age. You should probably bring this upstream. With my upstream hat on, a pure-Go implementation with simd/archsimd is very appealing if it lets us delete the existing assembly and get a performance speedup! |
|
Thanks for the quick feedback, @FiloSottile! We actually opened the Gerrit CL earlier today for It implements the zero-allocation We also staged every individual building block (the adversarial Looking forward to your review on Gerrit whenever you have cycles! |
|
Great question! The difference comes from the scope and granularity of each implementation:
In short: the Gerrit CL is the conservative drop-in for |
Uh oh!
There was an error while loading. Please reload this page.
Hi @FiloSottile and
agemaintainers,While benchmarking streaming AEAD pipelines on large payloads (64 KiB chunks), we explored whether the current sequential
cipher.AEAD.Sealinvocation ininternal/streamcould be accelerated in pure Go by interleaving keystream generation with polynomial reduction.We wanted to share some empirical microarchitectural data and a standalone adversarial testing harness we built, in case this is useful for
age.Microarchitectural Context (Intel Core i9-14900K, Single Thread)
In
internal/stream, sequential chunk sealing delegates tox/cryptoassembly routines that execute ChaCha20 rounds and Poly1305 carry chains in alternating bursts.By interleaving 8-block ChaCha20 (
simd/archsimd/ AVX2 unrolling) with 64-bit Poly1305 arithmetic (math/bits.Mul64->MULX/ADCX), instructions from both domains retire simultaneously across separate execution ports in the same clock cycle:Measured Hardware Performance (64 KiB Chunk via in-process
perf_event_open)ageBaseline (x/crypto)c2fused)Standalone Deliverables Available
agetorture(Harness & Adversarial Suite):A zero-dependency verification suite covering 6 strata (Micro 64B to MultiStream 4MiB), fragmented adversarial I/O (
OneByteReader, non-aligned boundary reads), and degraded entropy keys (internal/streamFused Engine:A drop-in, zero-allocation acceleration for the chunk seal/open loop (bit-exact compatibility with
agev1 stream format and terminal segment flags).perf_eventsProfiler:A lightweight Go wrapper over
SYS_PERF_EVENT_OPENto measure IPC and cache misses directly withingo test -v.If any of these components (even just the adversarial test harness) are of interest to the project, we'd be happy to open a clean PR or share a standalone repository for review.
Cheers,
Hazyhaar
All reactions