Skip to content

perf: align dask chunks with stored chunks and keep conversions lazy - #527

Open
fneum wants to merge 3 commits into
masterfrom
perf-chunk-alignment
Open

fneum wants to merge 3 commits into
masterfrom
perf-chunk-alignment

Conversation

@fneum

@fneum fneum commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Follow-up to #497 and PyPSA/pypsa-eur#2137.

Changes proposed in this Pull Request

Conversions were slow mainly because dask chunks did not align with the chunks stored in the cutout file. This PR fixes that in atlite and removes a few other bottlenecks. Results do not change.

  • Chunk alignment. Existing cutouts now open with chunks="auto" by default (also after Cutout.prepare). Before, {"time": 100} split the stored chunks (e.g. 2190 time steps in the PyPSA-Eur cutouts), so each stored chunk was decompressed about 20 times, serialised by the HDF5 lock. An explicit chunks argument is still respected. PyPSA-Eur already passes chunks="auto", but atlite applied it via Dataset.chunk("auto"), which ignores the stored chunks; it now goes through open_dataset.
  • CSP. efficiency.interp(...) merged all time steps into one chunk and added four (time, y, x) coordinates to the result. It is now a chunk-wise RegularGridInterpolator (the same linear scheme).
  • PV. trigon_model="other" computed the full graph once for a diagnostic warning (now at DEBUG level). tracking="tilted_horizontal" used np.where, which computes eagerly; now xr.where.
  • Dynamic line rating. line_rating created one delayed task per line, and each task read and decompressed the cutout data again. It now evaluates all (line, cell) pairs in one lazy operation and takes the minimum per line with np.fmin.reduceat (NaN-skipping like the former .min("spatial")).

Verification

Old vs new code on an extract of the PyPSA-Eur cutout europe-2013-sarah3-era5.nc with the original storage layout (Q1 2013, 131 x 90 cells), 15 conversion cases, aggregated time series and per-cell means:

  • Aggregated time series: bit-identical in all cases. Dims, coords, attrs, names and dtypes identical.
  • Per-cell time means: relative differences around 1e-15 (summation order), around 1e-6 for float32 outputs (wind, runoff).
  • line_rating (300 lines, 1 month): bit-identical, including lines outside the cutout.

Wall time (16 threads, aggregated time series):

Conversion master this PR
pv 48.6 s 6.4 s
pv, trigon_model="other" 81.3 s 7.4 s
pv, tracking="tilted_horizontal" 118.1 s 7.6 s
wind 13.3 s 0.9 s
csp 31.6 s 3.0 s
line_rating, 300 lines x 1 month 44 s 0.8 s

For a full year, pv took 373 s on master and 28 s with this PR. Peak memory can be higher with larger chunks (full-year pv: 3.6 GB on master vs up to 8.4 GB); pass chunks to reduce it.

Checklist

  • Code changes are sufficiently documented; i.e. new functions contain docstrings and further explanations may be given in doc.
  • Unit tests for new features were added (if applicable).
  • Newly introduced dependencies are added to environment.yaml, environment_docs.yaml and setup.py (if applicable).
  • A note for the release notes doc/release_notes.rst of the upcoming release is included.
  • I consent to the release of this PR's code under the MIT license.

fneum added 2 commits October 7, 2026 13:08
- Open existing cutouts with chunks="auto" by default (also after prepare),
  so dask chunks align with the chunks stored in the NetCDF file. Before,
  each stored chunk was decompressed many times.
- Interpolate the CSP efficiency per chunk with RegularGridInterpolator.
- Keep pv lazy for trigon_model="other" and tracking="tilted_horizontal".
- Vectorise line_rating over all (line, cell) pairs instead of one dask
  task per line.
Keep rows, cols = I.nonzero() from master and use np.unique for the first pair of each line instead of CSR internals. Pass the panel config as a plain dict in the laziness test.
@fneum
fneum requested a review from FabianHofmann October 7, 2026 12:41

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant