Skip to content

Correct the Welch and Mann-Whitney tests, and test every group - #49

Merged
Burdantes merged 1 commit into
mainfrom
fix/stat-udfs
Sep 30, 2026
Merged

Burdantes merged 1 commit into
mainfrom
fix/stat-udfs

Conversation

@Burdantes

Copy link
Copy Markdown
Collaborator

Why

Step 02 decides whether a group degraded using the persistent hermes.welchs_t_test and hermes.mann_whitney_u_test. Checked against scipy, they have three defects:

  1. Welch p-value: the incomplete-beta continued fraction (betacf) uses the wrong index terms in both recurrence steps, compared with Numerical Recipes. p is overestimated by up to ~0.2 when |t| is roughly 1 to 1.8, at every n.
  2. Mann-Whitney p-value: the continuity correction moves Z away from the mean (U = min(U1,U2) ≤ meanU, and 0.5 is subtracted), so p is underestimated for small groups. For example, n=5 gives 0.144 against scipy's 0.210.
  3. Both return p_value = 1e-10 with zeroed fields, without testing, when a sample exceeds 20,000 values. No reason is recorded, and both run 100,000 × 100,000 samples in ~2.6 s.

These values are now published through events_enriched.performance.tests.

What changes

  • Fixed code in udfs/welch_t_test.sql and udfs/mann_whitney_u.sql, inlined into Step 02 as TEMP functions via @requires-udf, as compute_wasserstein_p_value already is. Each image carries the exact detector it runs, staging can't share a live routine with production, and rollback is the previous image.
  • The shared persistent routines are untouched; only legacy queries use them. Their defects are documented in the UDF README.
  • tests/test_stat_udfs.py runs both JS bodies in Node on 11 deterministic cases (n=5 to 60,000, unequal sizes, heavy ties, three over the old cap) against frozen scipy results. The unfixed copies fail; the fixed ones match (Welch p to 1e-8, Mann-Whitney p to 1e-6, U exactly).

Staging: controlled comparison (hermes_staging, 2026-08-04, Step 02 run twice back to back on identical inputs)

control (main) treatment (this branch)
rows / groups 149,777 / 146,076 identical
non-test columns identical (src_lat/src_lon within 1.7e-13: float AVG order)
groups skipped by the 20k cap 112 0
RTT verdicts 3,713 3,526 (−187, +0)
download verdicts 2,764 2,684 (−81, +1)
upload / loss verdicts 753 / 4,684 751 / 4,675
runtime 3.3 min 3.1 min

All 187 lost RTT verdicts come from the Mann-Whitney fix (its p rose above 0.05). They are concentrated in small groups: median baseline n is 10, against 46 for all RTT verdicts. Mann-Whitney p rose in 64,752 groups and never fell. Welch p fell in 43,020 groups (the overestimate removed). The 20k fix changed no verdicts on this date.

⚠️ Methodology discontinuity

Deploying this reduces RTT verdicts by ~5% and download verdicts by ~3% from the first date processed by the new image. Earlier partitions keep the old verdicts unless Steps 02+ are backfilled. Any time series spanning the deploy date, and any figures computed from earlier partitions, mix the two.

Deploy

Image only: merge, git pull in ~/hermes-build-gate, docker build. The next nightly runs the fixed tests. Rollback: the previous image tag.

🤖 Generated with Claude Code

Step 02 called the persistent `hermes.welchs_t_test` and
`hermes.mann_whitney_u_test`, which have three defects, all confirmed against
scipy:

1. Welch p-value: the incomplete-beta continued fraction (betacf) used wrong
   index terms in both the even and odd steps (vs Numerical Recipes), so p was
   overestimated by up to ~0.2 when |t| is roughly 1 to 1.8, at every n.
2. Mann-Whitney p-value: the continuity correction was applied away from the
   mean (U = min(U1,U2) <= meanU, and 0.5 was subtracted), so Z was inflated
   and p underestimated for small samples (n=5: 0.144 vs scipy 0.210).
3. Both returned p_value = 1e-10 with zeroed fields, without testing, when a
   sample exceeded 20,000 values (131 groups/day). No reason is recorded; both
   run 100,000 x 100,000 samples in ~2.6 s in BigQuery.

The fixed code is inlined into Step 02 as TEMP functions (`welch_t_test`,
`mann_whitney_u`) via @requires-udf, as compute_wasserstein_p_value already
is, so each image carries the exact detector it runs, staging cannot share a
live routine with production, and rollback is the previous image. The shared
persistent routines are untouched (only legacy queries use them) and their
defects are documented in the UDF README.

tests/test_stat_udfs.py runs both JS bodies in Node on 11 deterministic cases
(n=5 to 60,000, unequal sizes, heavy ties, three over the old cap) against
frozen scipy results: the unfixed copies fail, the fixed ones match (Welch p
to 1e-8, Mann-Whitney p to 1e-6, U exactly).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Burdantes added a commit to m-lab/website that referenced this pull request Sep 30, 2026
m-lab/hermes#49 corrects the Welch t and Mann-Whitney U implementations and
removes a 20,000-sample short-circuit. Detection verdicts shift at the
deploy boundary (about 5% fewer RTT and 3% fewer download verdicts, measured
on 2026-08-04 with detection run both ways on identical inputs), so the
boundary is documented alongside the 2026-08-01 grouping change. Also records
that the reverse-path flags start on the same date.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Burdantes
Burdantes merged commit 4d73983 into main Sep 30, 2026
5 checks passed
sermpezis added a commit to m-lab/website that referenced this pull request Oct 5, 2026
* Add HERMES documentation section (draft for review)

HERMES has a blog post and a paper but no home on the site, so there is
nowhere to point someone who wants to query the data. This adds that section.

It sits under /tests/ rather than /data/, in a new "Analysis Systems" group:
HERMES is not a test M-Lab runs, but it is a system M-Lab operates, and the
site's existing pattern is that a system's pages live with the system while
/data/ links across to them. NDT works the same way.

The framing throughout keeps HERMES distinct from its traceroute-enrichment
stage. Enrichment produces the annotated paths HERMES uses as evidence; it is
one stage of the pipeline, reachable only through the methodology page, and
every enrichment page opens by saying so. Without that boundary the enrichment
tends to become the description of the whole product, which understates the
detection and localization work that is most of the published data.

Pages added:

  /tests/hermes/                                landing
  /tests/hermes/quickstart/                     access, cost, first queries
  /tests/hermes/schema/                         both tables, field by field
  /tests/hermes/examples/                       worked queries
  /tests/hermes/methodology/                    pipeline, stage by stage
  /tests/hermes/methodology/path-enrichment/    the enrichment component
  /learn/traceroute/                            general-audience explainer

The traceroute explainer is in /learn/ rather than under HERMES because it
serves scamper1, reverse traceroute, and IPRS readers equally, and /learn/
already says it is being restructured to hold exactly this kind of resource.

Field names, types and structure are read from create_events_enriched.sql and
04_mapping_union.sql in m-lab/hermes, and from the pipeline DDLs. Enrichment
data sources are as confirmed by the HERMES team.

Still a draft, and marked as such on every page. Open before publication:
  - access path is unresolved (mlab-collaboration grants vs. measurement-lab)
  - no example query has been run against the live table
  - row grain and coverage window need confirming
  - types need checking against INFORMATION_SCHEMA.COLUMNS
  - events_enriched does not expose the statistical test outputs

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167k8uZh7uL3Ei33tgpBfAV

* Update _pages/learn-traceroute.md

Co-authored-by: Pavlos Sermpezis <sermpezis.pavlos@gmail.com>

* Update _pages/datadocs.md

Co-authored-by: Pavlos Sermpezis <sermpezis.pavlos@gmail.com>

* Update _pages/tests/hermes/path-enrichment.md

Co-authored-by: Pavlos Sermpezis <sermpezis.pavlos@gmail.com>

* Address review on the HERMES docs and point them at the complete view

Review feedback from PR #891, plus what checking the drafts against the live
tables and the deployed pipeline turned up.

Review items: accept the drafted wording (partition_date note, hop-RTT
caveat, test-output note, scan estimate), drop duplicate links from the data
pages in favour of /tests/hermes/#start-here, plain-text section headings,
reframe traceroute enrichment as a general pipeline that HERMES uses, link
the paper where its figures are quoted, and add a traceroute.md pointer.

Placeholders resolved against m-lab/hermes and the live tables:
- a row is one NDT measurement with a traceroute in any analysed group, not
  only degraded ones; a group is ASN + metro + site + IP version;
- 7-day baseline; daily 15:00 UTC run; coverage 2025-02-21 onward with the
  listed gaps; three grouping periods (backfilled <= 2025-07-31 metro,
  2025-08-01..2026-07-31 city, >= 2026-08-01 metro), not one cutover;
- the 25-measurement / 5-IP gate is not enforced by the pipeline, so it is
  documented as the paper's recommended, tunable filter;
- access: hermes_union is open to discuss@, the hermes lookup tables are on
  request; partition filters are enforced.

Corrections found while validating every query:
- the example queries returned nothing because metro values are full
  strings ("Chicago-Illinois-US");
- partitions carry the 7-day lookback (only ~12% of a partition's rows are
  measured that day), so counts and medians now filter on
  DATE(measurement_time) = partition_date and the schema explains why;
- the docs now use only the canonical hermes_union.events_enriched view and
  its names. It exposes every operational column since m-lab/hermes#47,
  including performance.tests and performance.analysis_day, and *_count was
  never a count (now *_significant).

Adds a worked tutorial on a real event (T-Mobile, Seattle, 2026-09-18/19),
with every quoted figure taken from the published queries on 2026-09-30.

All SQL on these pages was dry-run against the live view, and executed
where it quotes results. The Jekyll site was not built locally.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Document the corrected loop and unresponsive-within-AS flags

m-lab/hermes#48 fixed both path flags to measure what they were meant to (an
AS reappearing after a different AS, and after hops with no ASN) and is live
in hermes_union.events_enriched. The old descriptions matched neither the old
nor the new behaviour.

- schema: new definitions; reverse-path flags describe the path as measured
  (before truncation) and are NULL for dates processed before the fix; the
  operational columns are marked as incorrect before the fix.
- path-enrichment: loop_detected is a reason for caution, not a diagnosis.
- examples: the trustworthy-paths query is now runnable and quotes what each
  filter keeps on 2025-07-04 (9.4% overall; loops 0.7%, gaps 29.5%).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Changelog: statistical test corrections from 2026-09-30

m-lab/hermes#49 corrects the Welch t and Mann-Whitney U implementations and
removes a 20,000-sample short-circuit. Detection verdicts shift at the
deploy boundary (about 5% fewer RTT and 3% fewer download verdicts, measured
on 2026-08-04 with detection run both ways on identical inputs), so the
boundary is documented alongside the 2026-08-01 grouping change. Also records
that the reverse-path flags start on the same date.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Mark the HERMES pages ready: drop draft banners, fix IP-to-AS source

Every banner guarded something now done: all _[confirm]_ items are resolved
against m-lab/hermes and the live tables, every query on these pages has been
executed (or dry-run where it quotes nothing) against
hermes_union.events_enriched, and every documented field name, type and mode
matches the live view's INFORMATION_SCHEMA.

The IP-to-AS row named only "CAIDA IP-to-prefix mapping". hermes.unified_ip_to_as
is built from CAIDA's RouteViews prefix-to-AS datasets plus IXP peering-LAN
prefixes from PeeringDB and PCH, with hopannotation2 used where it applies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Pavlos Sermpezis <sermpezis.pavlos@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant