Skip to content

Benchmark sharding basic functionality - #1813

Open
cmdupuis3 wants to merge 58 commits into
UXARRAY:mainfrom
cmdupuis3:cmd/bench_shards
Open

cmdupuis3 wants to merge 58 commits into
UXARRAY:mainfrom
cmdupuis3:cmd/bench_shards

Conversation

@cmdupuis3

@cmdupuis3 cmdupuis3 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #1814

Overview

This PR reworks the ASV CI workflow to split the benchmarking suite into separate jobs, allowing for parallel benchmarking runs.

Basically, all the benchmarks are scored and are roughly load-balanced based on the number of threads available. Each job runs independently, so there is some overhead with redundant benchmark cache builds. This work can also support larger HPC-scale jobs, but I'm deferring the actual HPC scripts to another PR.

Current timings show benchmark suite wall time usually around 6-8 minutes, versus 20-30 minutes for the current baseline.

PR Checklist

General

  • An issue is created and linked
  • Added appropriate labels (if your uxarray repo permissions allow it)
  • Filled out Overview and Expected Usage (if applicable) sections

Testing & Benchmarking

  • There is adequate test coverage of changes from this PR (add new tests if needed)
  • If this PR could affect performance, ran ASV benchmarks and confirmed they show expected behavior (add a new benchmark if necessary)

Documentation and Examples

  • Docstrings updated with any function changes, and included in all new functions
  • User (public) functions added to docs/api.rst; internal (private) function names start with an underscore (_)

AI Disclosure

AI Usage: Claude Opus 5, Opus 5.5

  • I have tested and take responsibility for all AI-generated content in my PR.

cmdupuis3 and others added 30 commits August 21, 2026 12:07
asv preimports the benchmark suite before it runs any setup_cache
(asv/runner.py, spawner.preimport() ahead of the run loop), so on a cold
cache bench_connectivity's import-time preload_topologies is what fills
it -- serially, in the forkserver parent, before a single benchmark
starts. CachedFixtures.setup_cache then finds everything already built,
and prime(workers=...) never runs on the path it was written for.

Filling it from the CLI first puts those reads back in the parallel
prime. Worth a second or two on the GitHub runners, which only see the
oQU grids; worth rather more on a machine that can reach the four
dyamond grids on campaign storage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
cmdupuis3 and others added 17 commits August 28, 2026 15:49
Main now carries squash-merged versions of work this branch had unsquashed
(cached benchmark I/O UXARRAY#1700, lazy/cached neighborhood kernels UXARRAY#1708/UXARRAY#1768,
asv env installs UXARRAY#1776, connectivity peakmem benchmarks UXARRAY#1661). Conflicts
resolved in main's favor except where the branch adds sharding:

- asv-benchmarking-pr.yml: keep the sharded setup/shard/merge jobs; adopt
  UXARRAY#1776's asv env cache key (ci/environment.yml + asv.conf.json, no
  restore-keys) in both jobs; drop a duplicated CPU topology step.
- asv-benchmarking.yml: take main's unguarded fixture priming; apply the same
  env cache key fix.
- asv.conf.json: take main's matrix and comments; keep the branch's BLAS
  thread pins in env_nobuild.
- .gitignore: keep the per-shard config/results entries.
- bench_connectivity.py: accept deletion; superseded by connectivity.py.
- neighbors.py, _fixtures.py, _warmup.py, mpas_ocean.py, geometry_samebody*:
  take main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cmd/bench_io is stale: its content landed on main as UXARRAY#1700 and has since been
superseded there (UXARRAY#1768 neighborhood kernels, UXARRAY#1776 asv env). This branch
already carries main, so every conflict resolves to the branch's side and the
tree is unchanged; the merge only records bench_io's history so PR #2 against
it is mergeable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
asv continuous is the same interleaved run plus a comparison, and exits 1
when the comparison finds a regression -- the status it also uses for a
broken run -- so any slower benchmark failed its shard (shards 0 and 1 in
run 37510066776). Shards now only measure; the merge job compares.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The main merge left a duplicate warm_in_parent import and a docstring-only
diff against main; both files now match main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Text only: module docstrings cut to purpose, key constraints and usage;
comments to one or two lines of why. Fixes two that had drifted from the
code (a stale stderr claim in _threads.resolve, a contradictory matrix note
in asv.conf.hpc.json) and a comment above the wheels cache that described
the durations cache.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NeighborhoodReduce ran at r=15, where a 120km face has ~600 neighbors and
time_dataset_reduce alone took over a minute per round -- the class that
bounded the slowest shard. It now runs at 5, shared with NeighborhoodDask
as NEIGHBORHOOD_RADIUS.

NeighborhoodBuild swept r over 1, 5 and 15 for every benchmark. The
timings now skip 15 (1.3s per call at 120km) and the track_* benchmarks
skip 1, where peakmem and nbytes barely move (362k vs 384k at 480km).
asv reads params from the method first, so this stays one class and the
kernels still compile once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same shards, caches and pre-builds; about 630 fewer lines.

- _merge: union rows and max durations. Every shard's benchmarks.json
  already lists the whole suite (asv saves all it discovers), and the
  conflict, ordering and column-realignment paths were unreachable.
- _partition: group by class. No setup_cache outside _fixtures spans
  classes, so the union-find never merged more. --shard now implies
  --asv-args; the uncalled --bench-args and --config-out are gone, and
  the report lists the heaviest benchmarks in place of the workflow's
  inline script.
- _machine: write asv's detected specs under the pinned name directly,
  replacing `asv machine --yes` plus a rename.
- Drop the unused _threads.py and asv.conf.hpc.json. The latter had
  drifted from asv.conf.json (no conda_environment_file), so HPC runs
  built a different env than CI. stage.pbs now always sets
  NUMBA_NUM_THREADS (default 8; numba's default counts SMT siblings) and
  passes -a timeout=1800.
- local.sh pins each shard to whole physical cores in socket order, so
  concurrent shards no longer share SMT siblings on nodes that number
  them N and N + cores.
- PR workflow: YAML anchors for the steps repeated across jobs, one
  SHARDS setting, and a shorter baseline resolution.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cmdupuis3 cmdupuis3 self-assigned this Oct 7, 2026
@cmdupuis3 cmdupuis3 added benchmarking Related to benchmarks, memory usage, and/or time profiling run-benchmark Run ASV benchmark workflow labels Oct 7, 2026
@cmdupuis3
cmdupuis3 marked this pull request as ready for review October 7, 2026 23:35
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

ASV Benchmarking

Benchmark Comparison Results

Benchmarks that have stayed the same:

Change Before [140b91c] After [b6d2e7e] Ratio Benchmark (Parameter)
989±20μs 998±10μs 1.01 connectivity.Connectivity.time_edge_face('120km')
490±20μs 493±10μs 1.01 connectivity.Connectivity.time_edge_face('480km')
4.05±0.06ms 4.08±0.1ms 1.01 connectivity.Connectivity.time_edge_node('120km')
1.33±0.05ms 1.29±0.01ms 0.97 connectivity.Connectivity.time_edge_node('480km')
80.5±9μs 86.1±9μs 1.07 connectivity.Connectivity.time_face_edge('120km')
75.3±8μs 77.2±10μs 1.03 connectivity.Connectivity.time_face_edge('480km')
1.04±0.01ms 1.03±0.01ms 0.99 connectivity.Connectivity.time_face_face('120km')
419±20μs 424±20μs 1.01 connectivity.Connectivity.time_face_face('480km')
66.3±6μs 65.5±20μs 0.99 connectivity.Connectivity.time_face_node('120km')
69.5±10μs 67.7±10μs 0.97 connectivity.Connectivity.time_face_node('480km')
496±30μs 498±50μs 1 connectivity.Connectivity.time_n_nodes_per_face('120km')
422±9μs 443±30μs 1.05 connectivity.Connectivity.time_n_nodes_per_face('480km')
1.37±0.02ms 1.40±0.04ms 1.02 connectivity.Connectivity.time_node_edge('120km')
477±20μs 511±20μs 1.07 connectivity.Connectivity.time_node_edge('480km')
85.1±1ms 86.2±0.9ms 1.01 connectivity.Connectivity.time_node_face('120km')
5.17±0.03ms 5.17±0.05ms 1 connectivity.Connectivity.time_node_face('480km')
1.42M 1.42M 1 connectivity.ConnectivityTracemalloc.track_peakmem_edge_face('120km')
106k 106k 1 connectivity.ConnectivityTracemalloc.track_peakmem_edge_face('480km')
6.48M 6.48M 1 connectivity.ConnectivityTracemalloc.track_peakmem_edge_node('120km')
417k 417k 1 connectivity.ConnectivityTracemalloc.track_peakmem_edge_node('480km')
2.42k 2.42k 1 connectivity.ConnectivityTracemalloc.track_peakmem_face_edge('120km')
2.42k 2.42k 1 connectivity.ConnectivityTracemalloc.track_peakmem_face_edge('480km')
1.6M 1.6M 1 connectivity.ConnectivityTracemalloc.track_peakmem_face_face('120km')
101k 101k 1 connectivity.ConnectivityTracemalloc.track_peakmem_face_face('480km')
2.54k 2.54k 1 connectivity.ConnectivityTracemalloc.track_peakmem_face_node('120km')
2.54k 2.54k 1 connectivity.ConnectivityTracemalloc.track_peakmem_face_node('480km')
240k 240k 1 connectivity.ConnectivityTracemalloc.track_peakmem_n_nodes_per_face('120km')
25.8k 25.8k 1 connectivity.ConnectivityTracemalloc.track_peakmem_n_nodes_per_face('480km')
1.9M 1.9M 1 connectivity.ConnectivityTracemalloc.track_peakmem_node_edge('120km')
127k 127k 1 connectivity.ConnectivityTracemalloc.track_peakmem_node_edge('480km')
11.9M 11.9M 1 connectivity.ConnectivityTracemalloc.track_peakmem_node_face('120km')
747k 742k 0.99 connectivity.ConnectivityTracemalloc.track_peakmem_node_face('480km')
8.76±0.3ms 8.54±0.1ms 0.98 face_bounds.FaceBounds.time_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/mpas/QU/oQU480.231010.nc'))
2.82±0.1ms 2.75±0.05ms 0.98 face_bounds.FaceBounds.time_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/scrip/outCSne8/outCSne8.nc'))
10.2±0.1ms 10.3±10ms 1.01 face_bounds.FaceBounds.time_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/geoflow-small/grid.nc'))
1.54±0.02ms 1.56±0.03ms 1.01 face_bounds.FaceBounds.time_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/quad-hexagon/grid.nc'))
57.3k 57.3k 1 face_bounds.FaceBounds.track_nbytes_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/mpas/QU/oQU480.231010.nc'))
12.3k 12.3k 1 face_bounds.FaceBounds.track_nbytes_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/scrip/outCSne8/outCSne8.nc'))
123k 123k 1 face_bounds.FaceBounds.track_nbytes_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/geoflow-small/grid.nc'))
128 128 1 face_bounds.FaceBounds.track_nbytes_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/quad-hexagon/grid.nc'))
1.27M 1.27M 1 face_bounds.FaceBounds.track_nbytes_grid_with_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/mpas/QU/oQU480.231010.nc'))
50.1k 50.1k 1 face_bounds.FaceBounds.track_nbytes_grid_with_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/scrip/outCSne8/outCSne8.nc'))
1.48M 1.48M 1 face_bounds.FaceBounds.track_nbytes_grid_with_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/geoflow-small/grid.nc'))
712 712 1 face_bounds.FaceBounds.track_nbytes_grid_with_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/quad-hexagon/grid.nc'))
1.98M 1.98M 1 face_bounds.FaceBounds.track_peakmem_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/mpas/QU/oQU480.231010.nc'))
1.97M 1.97M 1 face_bounds.FaceBounds.track_peakmem_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/scrip/outCSne8/outCSne8.nc'))
2.13M 2.13M 1 face_bounds.FaceBounds.track_peakmem_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/geoflow-small/grid.nc'))
35.5k 35.4k 1 face_bounds.FaceBounds.track_peakmem_face_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/quad-hexagon/grid.nc'))
1.13±0.03μs 1.23±4μs 1.09 geometry_kernels.AccucrossKernels.time_accucross
2.61±0.04μs 2.56±0.03μs 0.98 geometry_kernels.AccucrossKernels.time_accucross_pair
481±20ns 451±9ns 0.94 geometry_kernels.EFTPrimitives.time_acc_sqrt_re
446±30ns 451±10ns 1.01 geometry_kernels.EFTPrimitives.time_diff_of_products
426±10ns 396±20ns 0.93 geometry_kernels.EFTPrimitives.time_two_prod
376±30ns 401±5ns 1.07 geometry_kernels.EFTPrimitives.time_two_sum
661±9ns 722±40ns 1.09 geometry_kernels.GCAConstLatIntersection.time_accux_constlat_kernel
747±9ns 781±20ns 1.05 geometry_kernels.GCAConstLatIntersection.time_gca_const_lat_intersection
827±20ns 867±40ns 1.05 geometry_kernels.GCAConstLatIntersection.time_try_gca_const_lat_intersection
792±10ns 832±20ns 1.05 geometry_kernels.GCAGCAIntersection.time_accux_gca_kernel
872±20ns 902±20ns 1.03 geometry_kernels.GCAGCAIntersection.time_gca_gca_intersection
1.05±0.02μs 1.15±0.6μs 1.1 geometry_kernels.GCAGCAIntersection.time_try_gca_gca_intersection
52.2±0.6μs 53.5±0.7μs 1.02 geometry_kernels.OrientPredicates.time_on_minor_arc
52.2±0.4μs 53.4±0.6μs 1.02 geometry_kernels.OrientPredicates.time_orient3d_on_sphere
3.05±0ms 3.06±0ms 1 geometry_samebody.SameBodyConstLat.time_accux_dispatch
1.16±0ms 1.16±0.01ms 1 geometry_samebody.SameBodyConstLat.time_accux_kernel
2.29±0ms 2.28±0ms 1 geometry_samebody.SameBodyConstLat.time_fp64_dispatch
147±0.6μs 148±0.6μs 1 geometry_samebody.SameBodyConstLat.time_fp64_kernel
22.7±0.03ms 22.8±0.02ms 1.01 geometry_samebody_gcagca.SameBodyGcaGca.time_accux_dispatch
5.53±0.01ms 5.54±0.01ms 1 geometry_samebody_gcagca.SameBodyGcaGca.time_accux_kernel
17.8±0.02ms 17.9±0.1ms 1.01 geometry_samebody_gcagca.SameBodyGcaGca.time_fp64_dispatch
697±2μs 701±4μs 1.01 geometry_samebody_gcagca.SameBodyGcaGca.time_fp64_kernel
690±8ms 687±6ms 1 import.Imports.timeraw_import_uxarray
1.52±0.01ms 1.53±0.01ms 1.01 mpas_ocean.CheckNorm.time_check_norm('120km')
1.20±0.02ms 1.20±0.01ms 1 mpas_ocean.CheckNorm.time_check_norm('480km')
782±2μs 789±10μs 1.01 mpas_ocean.ConnectivityConstruction.time_face_face_connectivity('120km')
384±7μs 381±3μs 0.99 mpas_ocean.ConnectivityConstruction.time_face_face_connectivity('480km')
521±6μs 516±6μs 0.99 mpas_ocean.ConnectivityConstruction.time_n_nodes_per_face('120km')
420±8μs 415±6μs 0.99 mpas_ocean.ConnectivityConstruction.time_n_nodes_per_face('480km')
3.40±0ms 3.40±0.01ms 1 mpas_ocean.ConstructFaceLatLon.time_cartesian_averaging('120km')
2.61±0.04ms 2.56±0.01ms 0.98 mpas_ocean.ConstructFaceLatLon.time_cartesian_averaging('480km')
79.1±0.4ms 78.7±0.4ms 0.99 mpas_ocean.ConstructFaceLatLon.time_welzl('120km')
7.22±0.05ms 7.49±0.3ms 1.04 mpas_ocean.ConstructFaceLatLon.time_welzl('480km')
16.2±0.03ms 16.1±0.02ms 1 mpas_ocean.ConstructTreeStructures.time_ball_tree('120km')
810±10μs 808±30μs 1 mpas_ocean.ConstructTreeStructures.time_ball_tree('480km')
8.18±0.02ms 8.18±0.01ms 1 mpas_ocean.ConstructTreeStructures.time_kd_tree('120km')
479±10μs 474±20μs 0.99 mpas_ocean.ConstructTreeStructures.time_kd_tree('480km')
411±4ms 414±2ms 1.01 mpas_ocean.CrossSections.time_const_lat('120km', 1)
208±0.8ms 213±2ms 1.02 mpas_ocean.CrossSections.time_const_lat('120km', 2)
108±0.4ms 108±1ms 1 mpas_ocean.CrossSections.time_const_lat('120km', 4)
364±2ms 371±5ms 1.02 mpas_ocean.CrossSections.time_const_lat('480km', 1)
186±2ms 186±1ms 1 mpas_ocean.CrossSections.time_const_lat('480km', 2)
93.3±0.5ms 94.7±2ms 1.01 mpas_ocean.CrossSections.time_const_lat('480km', 4)
339M 358M 1.06 mpas_ocean.CrossSectionsPeakMem.track_peakmem_const_lat('120km', 1)
338M 358M 1.06 mpas_ocean.CrossSectionsPeakMem.track_peakmem_const_lat('120km', 2)
339M 358M 1.06 mpas_ocean.CrossSectionsPeakMem.track_peakmem_const_lat('120km', 4)
18.2±0.1ms 18.1±0.2ms 1 mpas_ocean.DualMesh.time_dual_mesh_construction('120km')
2.01±0.01ms 2.03±0.04ms 1.01 mpas_ocean.DualMesh.time_dual_mesh_construction('480km')
14.5±0.8ms 14.5±0.5ms 1 mpas_ocean.FaceAreas.time_face_areas('120km')
4.54±0.2ms 4.51±0.3ms 0.99 mpas_ocean.FaceAreas.time_face_areas('480km')
229k 229k 1 mpas_ocean.FaceAreas.track_nbytes_face_areas('120km')
14.3k 14.3k 1 mpas_ocean.FaceAreas.track_nbytes_face_areas('480km')
2.12M 2.12M 1 mpas_ocean.FaceAreas.track_peakmem_face_areas('120km')
715k 715k 1 mpas_ocean.FaceAreas.track_peakmem_face_areas('480km')
234±1ms 228±0.6ms 0.97 mpas_ocean.GeoDataFrame.time_to_geodataframe('120km', False)
39.4±0.7ms 38.6±0.5ms 0.98 mpas_ocean.GeoDataFrame.time_to_geodataframe('120km', True)
28.4±0.3ms 28.0±0.6ms 0.98 mpas_ocean.GeoDataFrame.time_to_geodataframe('480km', False)
3.80±0.1ms 3.80±0.2ms 1 mpas_ocean.GeoDataFrame.time_to_geodataframe('480km', True)
13.6±0.1ms 13.1±0.1ms 0.96 mpas_ocean.Gradient.time_gradient('120km')
1.83±0.01ms 1.80±0.02ms 0.99 mpas_ocean.Gradient.time_gradient('480km')
457k 457k 1 mpas_ocean.Gradient.track_nbytes_gradient('120km')
28.7k 28.7k 1 mpas_ocean.Gradient.track_nbytes_gradient('480km')
3.2M 3.2M 1 mpas_ocean.Gradient.track_peakmem_gradient('120km')
204k 204k 1 mpas_ocean.Gradient.track_peakmem_gradient('480km')
334M 358M 1.07 mpas_ocean.GradientColdStartRss.track_peakmem_gradient('120km')
278±7μs 278±9μs 1 mpas_ocean.HoleEdgeIndices.time_construct_hole_edge_indices('120km')
150±4μs 143±2μs 0.96 mpas_ocean.HoleEdgeIndices.time_construct_hole_edge_indices('480km')
142±0.6μs 141±2μs 0.99 mpas_ocean.Integrate.time_integrate('120km')
129±9μs 129±0.8μs 1 mpas_ocean.Integrate.time_integrate('480km')
18.4M 18.4M 1 mpas_ocean.Integrate.track_nbytes_integrate('120km')
1.2M 1.2M 1 mpas_ocean.Integrate.track_nbytes_integrate('480km')
186±2ms 183±3ms 0.99 mpas_ocean.MatplotlibConversion.time_dataarray_to_polycollection('120km', 'exclude')
184±0.8ms 187±2ms 1.01 mpas_ocean.MatplotlibConversion.time_dataarray_to_polycollection('120km', 'include')
185±1ms 186±3ms 1.01 mpas_ocean.MatplotlibConversion.time_dataarray_to_polycollection('120km', 'split')
13.3±0.07ms 13.4±0.4ms 1.01 mpas_ocean.MatplotlibConversion.time_dataarray_to_polycollection('480km', 'exclude')
13.5±0.1ms 13.8±0.4ms 1.02 mpas_ocean.MatplotlibConversion.time_dataarray_to_polycollection('480km', 'include')
13.2±0.09ms 13.4±0.2ms 1.01 mpas_ocean.MatplotlibConversion.time_dataarray_to_polycollection('480km', 'split')
245±0.3ms 246±1ms 1 mpas_ocean.NeighborhoodBuild.time_build('120km', 1.0)
507±3ms 508±3ms 1 mpas_ocean.NeighborhoodBuild.time_build('120km', 5.0)
13.5±0.05ms 13.2±0.01ms 0.98 mpas_ocean.NeighborhoodBuild.time_build('480km', 1.0)
16.5±0.05ms 16.6±0.2ms 1.01 mpas_ocean.NeighborhoodBuild.time_build('480km', 5.0)
241±0.1ms 240±0.2ms 1 mpas_ocean.NeighborhoodBuild.time_query_radius('120km', 1.0)
503±4ms 499±0.6ms 0.99 mpas_ocean.NeighborhoodBuild.time_query_radius('120km', 5.0)
12.9±0.03ms 12.9±0.01ms 1 mpas_ocean.NeighborhoodBuild.time_query_radius('480km', 1.0)
16.2±0.1ms 16.1±0.05ms 0.99 mpas_ocean.NeighborhoodBuild.time_query_radius('480km', 5.0)
612.76 612.76 1 mpas_ocean.NeighborhoodBuild.track_mean_neighbors('120km', 15.0)
74.17 74.17 1 mpas_ocean.NeighborhoodBuild.track_mean_neighbors('120km', 5.0)
37.29 37.29 1 mpas_ocean.NeighborhoodBuild.track_mean_neighbors('480km', 15.0)
6.57 6.57 1 mpas_ocean.NeighborhoodBuild.track_mean_neighbors('480km', 5.0)
141M 141M 1 mpas_ocean.NeighborhoodBuild.track_nbytes_neighbors('120km', 15.0)
17.4M 17.4M 1 mpas_ocean.NeighborhoodBuild.track_nbytes_neighbors('120km', 5.0)
563k 563k 1 mpas_ocean.NeighborhoodBuild.track_nbytes_neighbors('480km', 15.0)
123k 123k 1 mpas_ocean.NeighborhoodBuild.track_nbytes_neighbors('480km', 5.0)
145M 145M 1 mpas_ocean.NeighborhoodBuild.track_peakmem_build('120km', 15.0)
21.5M 21.5M 1 mpas_ocean.NeighborhoodBuild.track_peakmem_build('120km', 5.0)
825k 825k 1 mpas_ocean.NeighborhoodBuild.track_peakmem_build('480km', 15.0)
384k 384k 1 mpas_ocean.NeighborhoodBuild.track_peakmem_build('480km', 5.0)
40.2±0.3ms 41.0±0.4ms 1.02 mpas_ocean.NeighborhoodDask.time_mean('120km', 'grid_chunks')
18.3±0.02ms 18.1±0.07ms 0.99 mpas_ocean.NeighborhoodDask.time_mean('120km', 'numpy')
37.1±0.6ms 36.9±0.6ms 0.99 mpas_ocean.NeighborhoodDask.time_mean('120km', 'time_chunks')
10.9±0.2ms 10.6±0.1ms 0.97 mpas_ocean.NeighborhoodDask.time_mean('480km', 'grid_chunks')
736±8μs 722±10μs 0.98 mpas_ocean.NeighborhoodDask.time_mean('480km', 'numpy')
7.34±0.2ms 7.48±0.2ms 1.02 mpas_ocean.NeighborhoodDask.time_mean('480km', 'time_chunks')
5.83M 5.82M 1 mpas_ocean.NeighborhoodDask.track_peakmem_mean('120km', 'grid_chunks')
2.75M 2.75M 1 mpas_ocean.NeighborhoodDask.track_peakmem_mean('120km', 'numpy')
5.68M 5.68M 1 mpas_ocean.NeighborhoodDask.track_peakmem_mean('120km', 'time_chunks')
679k 679k 1 mpas_ocean.NeighborhoodDask.track_peakmem_mean('480km', 'grid_chunks')
177k 177k 1 mpas_ocean.NeighborhoodDask.track_peakmem_mean('480km', 'numpy')
538k 547k 1.02 mpas_ocean.NeighborhoodDask.track_peakmem_mean('480km', 'time_chunks')
4.48±0s 4.50±0.01s 1 mpas_ocean.NeighborhoodReduce.time_dataset_reduce('120km', 'mean')
4.63±0.01s 4.63±0.01s 1 mpas_ocean.NeighborhoodReduce.time_dataset_reduce('120km', 'median')
122±0.2ms 122±0.1ms 1 mpas_ocean.NeighborhoodReduce.time_dataset_reduce('480km', 'mean')
123±0.09ms 124±0.2ms 1.01 mpas_ocean.NeighborhoodReduce.time_dataset_reduce('480km', 'median')
508±3ms 512±2ms 1.01 mpas_ocean.NeighborhoodReduce.time_neighborhood_reduce('120km', 'mean')
541±2ms 543±1ms 1 mpas_ocean.NeighborhoodReduce.time_neighborhood_reduce('120km', 'median')
16.6±0ms 16.7±0.03ms 1.01 mpas_ocean.NeighborhoodReduce.time_neighborhood_reduce('480km', 'mean')
17.1±0.1ms 17.0±0.04ms 0.99 mpas_ocean.NeighborhoodReduce.time_neighborhood_reduce('480km', 'median')
4.54±0.04ms 4.56±0.01ms 1 mpas_ocean.NeighborhoodReduce.time_reduce('120km', 'mean')
37.3±0.2ms 37.5±0.09ms 1.01 mpas_ocean.NeighborhoodReduce.time_reduce('120km', 'median')
244±10μs 254±10μs 1.04 mpas_ocean.NeighborhoodReduce.time_reduce('480km', 'mean')
532±10μs 539±10μs 1.01 mpas_ocean.NeighborhoodReduce.time_reduce('480km', 'median')
234k 234k 1 mpas_ocean.NeighborhoodReduce.track_peakmem_reduce('120km', 'mean')
235k 235k 1 mpas_ocean.NeighborhoodReduce.track_peakmem_reduce('120km', 'median')
19.5k 19.5k 1 mpas_ocean.NeighborhoodReduce.track_peakmem_reduce('480km', 'mean')
19.5k 19.4k 1 mpas_ocean.NeighborhoodReduce.track_peakmem_reduce('480km', 'median')
442±10μs 459±9μs 1.04 mpas_ocean.PointInPolygon.time_face_search_lonlat('120km')
409±10μs 430±20μs 1.05 mpas_ocean.PointInPolygon.time_face_search_lonlat('480km')
385±10μs 418±20μs 1.09 mpas_ocean.PointInPolygon.time_face_search_xyz('120km')
392±10μs 400±9μs 1.02 mpas_ocean.PointInPolygon.time_face_search_xyz('480km')
132±0.3ms 135±0.4ms 1.03 mpas_ocean.RemapDownsample.time_bilinear_remapping
17.7±0.09ms 17.9±0.2ms 1.01 mpas_ocean.RemapDownsample.time_inverse_distance_weighted_remapping
16.0±0.2ms 16.0±0.07ms 1 mpas_ocean.RemapDownsample.time_nearest_neighbor_remapping
1.44±0.01s 1.43±0.01s 1 mpas_ocean.RemapUpsample.time_bilinear_remapping
27.1±0.3ms 26.6±0.4ms 0.98 mpas_ocean.RemapUpsample.time_inverse_distance_weighted_remapping
12.1±0.1ms 12.3±0.1ms 1.01 mpas_ocean.RemapUpsample.time_nearest_neighbor_remapping
6.38±0.06ms 6.81±0.04ms 1.07 mpas_ocean.ZonalAverage.time_zonal_average('120km')
3.65±0.02ms 3.72±0.1ms 1.02 mpas_ocean.ZonalAverage.time_zonal_average('480km')
341M 357M 1.05 mpas_ocean.ZonalAveragePeakMem.track_peakmem_zonal_average('120km')
1.0256211772841293 1.0278681798861602 1 nogil_scaling.GILScaling.track_gil_scaling
7.56±0.02ms 7.52±0.04ms 0.99 quad_hexagon.QuadHexagon.time_open_dataset
6.42±0.1ms 6.45±0.2ms 1.01 quad_hexagon.QuadHexagon.time_open_grid
408 408 1 quad_hexagon.QuadHexagon.track_nbytes_open_dataset
392 392 1 quad_hexagon.QuadHexagon.track_nbytes_open_grid
73.1k 73.7k 1.01 quad_hexagon.QuadHexagon.track_peakmem_open_dataset
72.4k 72.9k 1.01 quad_hexagon.QuadHexagon.track_peakmem_open_grid

Benchmarks that have got worse:

Change Before [140b91c] After [b6d2e7e] Ratio Benchmark (Parameter)
+ 319M 358M 1.12 face_bounds.FaceBoundsColdStartRss.track_peakmem_open_and_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/mpas/QU/oQU480.231010.nc'))
+ 319M 358M 1.12 face_bounds.FaceBoundsColdStartRss.track_peakmem_open_and_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/scrip/outCSne8/outCSne8.nc'))
+ 320M 358M 1.12 face_bounds.FaceBoundsColdStartRss.track_peakmem_open_and_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/geoflow-small/grid.nc'))
+ 320M 358M 1.12 face_bounds.FaceBoundsColdStartRss.track_peakmem_open_and_bounds(PosixPath('/home/runner/work/uxarray/uxarray/test/meshfiles/ugrid/quad-hexagon/grid.nc'))
+ 276M 358M 1.3 import.Imports.track_peakmem_import_uxarray
+ 322M 358M 1.11 mpas_ocean.CrossSectionsPeakMem.track_peakmem_const_lat('480km', 1)
+ 322M 358M 1.11 mpas_ocean.CrossSectionsPeakMem.track_peakmem_const_lat('480km', 2)
+ 322M 358M 1.11 mpas_ocean.CrossSectionsPeakMem.track_peakmem_const_lat('480km', 4)
+ 314M 358M 1.14 mpas_ocean.GradientColdStartRss.track_peakmem_gradient('480km')
+ 323M 357M 1.1 mpas_ocean.ZonalAveragePeakMem.track_peakmem_zonal_average('480km')

…ns from main

asv_runner re-runs setup before every sample, about six times per benchmark
process. NeighborhoodReduce.time_reduce timed a 5ms kernel at 120km but took
27s, and NeighborhoodDask.time_mean 49s for 1-40ms calls, almost all setup.
That made NeighborhoodReduce alone larger than a quarter of the suite, which
capped a 4-shard split at 3.2x.

- NeighborhoodReduce and NeighborhoodDask build their fixture, warmup and
  neighborhood once per process. The timed calls only read them. The grid
  caches a single ball tree and time_dataset_reduce leaves it on edges, so
  setup restores the face tree each sample, as a fresh setup did.
- main's benchmark workflow saves the commit it just ran under the
  asv-results cache key. PR caches are scoped to their PR, so every PR's
  first run had no durations and split by count (this branch's first run:
  shards of 3:11, 4:48, 2:17 and 9:49).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@Sevans711 Sevans711 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great overall, I'm excited for very fast ASV benchmark runs! I read through all the changes and they look pretty reasonable to me. Because this is all for benchmarks, not anything user-facing, I didn't take the time to spin up a local or HPC environment to test it directly for myself. It seems to work, and I think that should be good enough? And, if there's some subtle bug that gets missed here, it could always be fixed in a follow-up PR, without users ever being affected directly.

I think there's only one important thing to fix before merging, and that would be the roughly 11% to 30% increases to track_peakmem for many benchmarks here. I would hope to see no changes to peakmem usage, to avoid re-introducing something like #1605. Any guesses on what is causing this?

Beyond that, I also have a few clarifying questions, just so I can understand this a bit better:

  1. Where do the sharding timing weights come from when you make a new github action ASV run? I can understand from the code that if you run things locally you are just using the previous job's weights to decide how to group jobs into shards. Are these being cached somewhere on github? What happens if the cache doesn't exist?
  2. Can you clarify, are the changes in mpas_ocean.py related to the sharding at all, or are they a completely separate improvement (i.e. is this part of the original issue or an expansion of PR scope)? I don't necessarily have any concerns about those changes, just trying to understand the motivation for them a bit better.

@cmdupuis3

cmdupuis3 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

@Sevans711 The memory issues are ongoing, I think this issue is something about subprocess_peak_rss on main.

  1. At the moment, they are basically cached and updated as a running average upon every benchmark run. To me, this makes sense on HPC, but I'm looking into possibly using a more robust mechanism, like having one of the bots here scan the repo's benchmark runs and re-weight them if they get too unbalanced.
  2. These are mostly unrelated, but are important because the neighborhood filters are otherwise the heaviest jobs atm, and they are harder to load-balance against. The two change sets are: (a) pruning expensive benchmarks by removing the large-radius jobs where we don't need them, and (b) running NeighborhoodReduce.setup per shard rather than per sample.

@jayeshkrishna jayeshkrishna left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would recommend fusing/melding commits together so that the history of changes are more tractable.

My understanding of the changes in this PR,

The benchmark test suite run using asv (Airspeed Velocity) in github CI workflow takes up to 20-30 mins to run. This PR shards (breaks up) the benchmark test suite into multiple independent jobs, containing batches of tests, that run in parallel. Running these test shards/batches in parallel reduces the overall time taken to run the benchmark. However since the project does not use custom testing frameworks that support sharding the mechanism of sharding/batching tests is implemented manually.

The PR contains the following changes,

  • A helper script, _machine.py, that captures the "machine/node name" from the environment (or platform)
  • Helper scripts for sharding/partitioning the tests into batches and gathering the test results
    ** Script _partition.py : Reads benchmark.json (asv output, one for each run) from result dirs to calculate mean "duration" (wallclock time for benchmark) for each benchmark. The benchmarks are binned greedily (sort benchmarks using duration, longest first, and add to bin with the lowest load) into separate shards/bins (number of bins is a parameter to script). The asv config file (parameter to script) is updated with the shard information.
    ** Script _merge.py : Since the shard workflow jobs are run in separate directories (and the results are in these directories) this script consolidates/merges the results (benchmarks.json) and machine files (machine.json) in each result directory into a separate result directory (benchmarks.json in this directory contains results from all shard jobs).
  • Since all shard jobs need to use the same input files (oQU*.nc), asv files (wheels) these files are cached
  • Misc fixes
    ** Adding env for OMP, MKL, BLAS
    ** Updates to MPAS neighborhood tests

env:
BASE: ${{ needs.setup.outputs.base }}
run: |
set -ex

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible to use pre-built actions here (actions/upload-artifact/merge@v4)?

merged = {}
for shard_dir in map(Path, shard_dirs):
for path in sorted(shard_dir.rglob("*.json")):
rel = path.relative_to(shard_dir)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the function is merging files with the same name across directories. Does the function also merge json files other than benchmarks.json ("results" and "duration" - referred later - seem to be just in the benchmarks.json file)?
Renaming rel ("relative file path/name") to something like "fname" makes the intent more clear

for path in sorted(shard_dir.rglob("*.json")):
rel = path.relative_to(shard_dir)
data = json.loads(path.read_text())
if rel not in merged or path.name in _WHOLE_FILES:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you need the extra check in _WHOLE_FILES?

for path in sorted(shard_dir.rglob("*.json")):
rel = path.relative_to(shard_dir)
data = json.loads(path.read_text())
if rel not in merged or path.name in _WHOLE_FILES:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adding a comment that you are trying to find the first file to append would be useful. You might also want to instead do append vs initialize based on whether the merged[fname] is set or not (make the intent more clear)



def load_weights(results_dirs):
"""Mean recorded duration per benchmark, in seconds.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not use median (to get rid of outliers) here too?

for results_dir in results_dirs:
for path in sorted(Path(results_dir).glob("*/*.json")):
data = json.loads(path.read_text())
columns = data.get("result_columns") or []

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like you want to read the "duration" for each "benchmark name" in "results". Avoiding accessing "result_columns" might simplify this code

class_cost = {owner: sum(cost[name] for name in names) for owner, names in classes.items()}

shards = [[] for _ in range(n_shards)]
loads = [0.0] * n_shards

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like you are binning greedily a sorted (longest duration in the front) list of benchmarks. Having a list of tuples, (benchmark_name, benchmark_cost), and iterating over that list would make the intent more clear. Also renaming classes/owner (or adding more comments) would help.

Comment thread benchmarks/mpas_ocean.py
params = DatasetBenchmark.params + [['mean', 'median']]

radius = 15.0
radius = NEIGHBORHOOD_RADIUS

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was this intentional (NEIGHBORHOOD_RADIUS is 5 now)?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, the radius parameter scales the algorithm cost as O(N^2), so these neighborhood algorithms are very expensive in general. I chose radii that will hopefully show the scaling of the radius parameter on the grids we use, without needlessly bloating the compute on the shard the neighborhood filter benchmarks land on.

if: always()
with:
name: asv-benchmark-results-${{ runner.os }}
name: asv-benchmark-results-Linux

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was this an intentional change?

@jayeshkrishna

Copy link
Copy Markdown

As the number of benchmark runs (previous/older) increase you might also want to investigate alternatives to scanning results from all previous runs to calculate mean benchmark runtimes (when binning the tests to shard jobs).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmarking Related to benchmarks, memory usage, and/or time profiling run-benchmark Run ASV benchmark workflow

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelized benchmark jobs

3 participants