Repository navigation
Conversation
- Open existing cutouts with chunks="auto" by default (also after prepare), so dask chunks align with the chunks stored in the NetCDF file. Before, each stored chunk was decompressed many times. - Interpolate the CSP efficiency per chunk with RegularGridInterpolator. - Keep pv lazy for trigon_model="other" and tracking="tilted_horizontal". - Vectorise line_rating over all (line, cell) pairs instead of one dask task per line.
5 tasks done
Keep rows, cols = I.nonzero() from master and use np.unique for the first pair of each line instead of CSR internals. Pass the panel config as a plain dict in the laziness test.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #497 and PyPSA/pypsa-eur#2137.
Changes proposed in this Pull Request
Conversions were slow mainly because dask chunks did not align with the chunks stored in the cutout file. This PR fixes that in atlite and removes a few other bottlenecks. Results do not change.
chunks="auto"by default (also afterCutout.prepare). Before,{"time": 100}split the stored chunks (e.g. 2190 time steps in the PyPSA-Eur cutouts), so each stored chunk was decompressed about 20 times, serialised by the HDF5 lock. An explicitchunksargument is still respected. PyPSA-Eur already passeschunks="auto", but atlite applied it viaDataset.chunk("auto"), which ignores the stored chunks; it now goes throughopen_dataset.efficiency.interp(...)merged all time steps into one chunk and added four(time, y, x)coordinates to the result. It is now a chunk-wiseRegularGridInterpolator(the same linear scheme).trigon_model="other"computed the full graph once for a diagnostic warning (now atDEBUGlevel).tracking="tilted_horizontal"usednp.where, which computes eagerly; nowxr.where.line_ratingcreated one delayed task per line, and each task read and decompressed the cutout data again. It now evaluates all (line, cell) pairs in one lazy operation and takes the minimum per line withnp.fmin.reduceat(NaN-skipping like the former.min("spatial")).Verification
Old vs new code on an extract of the PyPSA-Eur cutout
europe-2013-sarah3-era5.ncwith the original storage layout (Q1 2013, 131 x 90 cells), 15 conversion cases, aggregated time series and per-cell means:line_rating(300 lines, 1 month): bit-identical, including lines outside the cutout.Wall time (16 threads, aggregated time series):
pvpv,trigon_model="other"pv,tracking="tilted_horizontal"windcspline_rating, 300 lines x 1 monthFor a full year,
pvtook 373 s on master and 28 s with this PR. Peak memory can be higher with larger chunks (full-yearpv: 3.6 GB on master vs up to 8.4 GB); passchunksto reduce it.Checklist
doc.environment.yaml,environment_docs.yamlandsetup.py(if applicable).doc/release_notes.rstof the upcoming release is included.