Repository navigation
Prepare v0.4 with verified current evidence and installable modules - #17
Merged
Merged
Conversation
fetch-traces.sh could not download MSR Cambridge and printed instructions for fetching it by hand from SNIA IOTTA trace 388. That source hands files out only through a browser form (cookies, name, affiliation, email), and in September 2026 iotta.snia.org did not respond at all, from two separate networks. MSR volumes were therefore never part of a scripted evidence run. The script now takes six volumes (hm_0, prn_0, proj_0, src1_2, usr_0, web_0, 210 MB) from the mirror the cacheMon project keeps at cache-datasets.s3.amazonaws.com/cache_dataset_txt/2008_msr, which holds SNIA's original msr-cambridge1.tar and msr-cambridge2.tar. The SNIA Trace Data Files Download License v2.0 permits use and redistribution without restriction, so this is a lawful copy. Only the listed volumes are fetched, by byte range out of the 5.3 GB of uncompressed tar. Each entry pins the volume's tar header offset, its size, and the MD5 given for it in the archive's MD5.txt. Before downloading, the script reads the member name from the tar header at that offset; after downloading, it checks the MD5. Either mismatch fails the script and leaves no file behind. Those checksums come from inside the mirrored archive, so they detect a corrupted or repacked download, not deliberate tampering; the mirror has not been compared with a copy from SNIA, because SNIA could not be reached. docs/benchmarking.md says so. Verified: - the script fetches all six volumes and every MD5 matches (19.7 s); - a wrong MD5 and a wrong offset each fail with exit 1, naming the problem (for the offset, the member actually found there), and leave no file; - bash -n passes; shellcheck is not installed here and was not run; - golangci-lint and go vet on bench: clean (the Go change is a comment).
Every number the evidence suite reports from a real trace depends on the loader turning the file into the right request sequence and on the replay counting hits the way other simulators do. The format fixtures prove only that a loader reads the rows it was tested on. Nothing compared the whole pipeline with an independent implementation, and the v0.4.0 plan requires that comparison before the trace tables are regenerated. make verify-ref (scripts/verify-ref.sh) does it for all twelve traces the suite reads: Twitter cluster052, LIRS loop and 2_pools, ARC P3 and OLTP, Meta kvcache 202206 and the six MSR volumes. - awk in the script expands each raw file into one key per request, following the loader's documented rules: ARC block runs, Meta op_count repeats over GET rows, MSR 512-byte block ranges over reads. It does not call the Go loaders, so a loader bug cannot cancel itself out. - libCacheSim's cachesim, built into .tools/ at the pinned commit 1d7415569978330ea95c9cff06a260630406f7e3, replays that sequence through LRU with object sizes ignored, at 0.25x to 4x the suite's capacity. - TestLRUMatchesReference loads the same files through the Go loaders, replays this repository's LRU at the same capacities, and requires the same request count and a miss ratio within 0.5 points. It also requires the reference to include the capacity the suite actually uses. The gate fails rather than passing over nothing: libCacheSim that will not build, a missing trace, an unset AS_CACHE_TRACES, or a skipped Go test each exit 1 with a message. The test skips when no reference is supplied, so make test and make evidence are unaffected. Nothing runs in CI: libCacheSim needs a C toolchain with glib and argp, and the traces are not committed. Verified: - make verify-ref: 60 points, largest miss-ratio difference 0.005 points (the rounding of cachesim's four-decimal output), request counts equal at every point; - the test fails with a miss ratio off by 1.4 points, with a request count off by one, and with a reference that omits the suite's capacity, and passes on the correct row; it skips when the reference is unset; - the script exits 1 with a trace missing and with AS_CACHE_TRACES unset; - golangci-lint on bench: 0 issues. shellcheck is not installed here. The first run of the script stopped after the third trace with status 141 and no message: once awk reached the request limit, gzip took SIGPIPE and pipefail ended the script. Decompression now tolerates exactly that status, and an ERR trap reports the line on which any other failure stopped.
TestTraceEvidence replayed the adaptive cache on a 2ms wall-clock epoch, once per trace. How many epochs a replay saw depended on how fast the machine ran it, so the result moved between runs, and the table in docs/evidence.md (measured at 50ms, a setting that was never in the test) could not be reproduced by make evidence. On cd8502f and on 1e2e599, before the B1 and B2 changes, the 2ms test gave ARC P3 3.3-3.7% against the documented 11.4%. At 50ms, five runs across both commits gave 6.2-9.6%. The test now: - ends epochs on request counts (EpochRequests), set as a number of epochs over the whole trace (10, 20 and 50), reporting all three rather than the best of them; - keeps the configuration production would use: all nine arms and a shadow sample rate of 0.05; - replays every subject that is not reproducible five times and reports the median with its range. That covers the adaptive cache, whose sampler hash is seeded per cache, and the Random and W-TinyLFU arms. The deterministic arms are replayed once; - prints a per-trace table and a summary row per trace, and, when AS_CACHE_EVIDENCE_OUT names a file, writes every run as JSON with the commit, whether the tree was modified, the Go version, the platform and the settings. The assertion is unchanged in substance: the adaptive median at each epoch length must beat the median of the worst fixed policy. It moved to a new file, bench/trace_evidence_test.go, because trace_test.go keeps the trace list and the loader checks. make evidence's timeout rises from 20m to 45m for the extra replays. Verified: TestTraceEvidence passes on all twelve traces (Twitter, LIRS loop and 2_pools, ARC P3 and OLTP, Meta kvcache, six MSR volumes) in 330 s, and writes the JSON. go vet and golangci-lint on bench: clean. The full make evidence run and the docs rewrite follow separately.
sshaplygin
marked this pull request as ready for review
September 27, 2026 22:05
sshaplygin
marked this pull request as draft
September 28, 2026 20:41
sshaplygin
marked this pull request as ready for review
September 28, 2026 21:20
…ecise reference inputs
The previous dataset was measured at 00bdcb1. Commits after it changed the measurement code itself - bench/evidence_test.go, bench/competitor_test.go and bench/competitors.go, through the V1 and V2 repairs - so the retained numbers no longer came from the code in this branch. The dataset validator still passed, because it checks the dataset's internal consistency, not whether measurement sources changed after the recorded commit. Recorded with scripts/record_evidence.py --out against the committed snapshot 1cb65fe, in an isolated checkout with developer files excluded: the libCacheSim reference gate, the Meta byte experiment, and three consecutive full evidence runs. Wall time 31m43s. Nothing was edited by hand; README.md is the generator's output. What moved, as the non-deterministic subjects (adaptive with sampling, Random, W-TinyLFU) move between batches: - Adaptive medians trail the best fixed median on 11 of 12 traces at every tested epoch setting (10 of 12 in the previous dataset, where msr_usr_0 at 50 epochs was +0.05 pp; it is now -0.16 pp). ARC P3 remains the one trace with positive deltas at all three settings: +0.35, +0.82, +0.32 pp. - The single adaptive median below the worst fixed median is still msr_prn_0 at 20 epochs: -0.0207 pp against W-TinyLFU 0.75% [0.67-0.93] (previously -0.0510 pp). - W-TinyLFU on LIRS loop: 49.26% [42.39-59.77], up from 44.24% [40.52-45.80]. Its upper range now reaches into the higher mode recorded earlier, which widens the adaptive deficits on that trace to -8.58, -6.39 and -5.17 pp. - reference.tsv and reference.log are byte-identical: the calibration is deterministic. No document quotes any of the changed figures; README, docs and the site link the generated tables, so nothing outside bench/results/current changes. Verified: record_evidence.py --verify on the recorded directory and again in place: "Verified 16 artifacts and generated report at 1cb65fe".
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The previous evidence gate rejected unfavorable random outcomes, and the release check could package uncommitted code. This prepares the v0.4.0 candidate with a complete current dataset, explicit measurement limits, and eight modules checked as external dependencies.
bench/results/current/. Replace old results and remove the historical site explorer and timeline. Remove unsupported maximum-deficit and single-run timing claims from the documentation and site.benchclient.DefaultArmsto the four arms published in v0.3.1. Add pinned Python linting and formatting and make prepublication tidy skips explicit.Validation: three consecutive full
make evidenceruns passed on clean source3331d5f; all 60 reference points passed. The strict manifest verifier checks all 16 artifacts and recomputes pooled results from the raw batches. Finalmake allpassed on29a831f, including vet/lint/race-short checks across twelve modules, 19 script tests, 10 release-check fixtures and eight external consumer builds. Repeated-race -short -count=3tests passed in the root, bench and benchclient modules, together with basic and cold/warm/gradual migration smoke runs. GitHub CI passed all 25 checks on29a831f. Independent QA, critic and reviewer all approved the final candidate.This PR prepares a candidate. No v0.4.0 tags or GitHub release have been published. Human review and merge, tag publication and
make release-check-publishedremain required. The offline ObserveOnly sweep is complete; a real-service trial remains pending a service and environment.