Summary
On arm64 macOS, LLVM libomp defaults KMP_BLOCKTIME to 0.
Worker threads therefore go to sleep right after every #pragma omp parallel region,
and the next region has to wake them again through a condition variable.
The SW engine opens many small parallel regions per frame (one per image, or one per shape).
On this platform the wake/sleep cost exceeds the work being parallelized.
Setting only KMP_BLOCKTIME=1 (1 ms) doubles the CPU benchmark FPS.
We should either:
- (A) set
KMP_BLOCKTIME in the benchmark runs so the numbers reflect ThorVG and not the libomp idle policy, or
- (B) call
kmp_set_blocktime() inside the SW engine when OpenMP is enabled. ( LLVM/Intel libomp)
Environment
- Apple M5 Pro (6P + 12E cores), macOS 26 (Darwin 25)
- ThorVG CPU engine,
-Dextra=openmp (the meson default), libomp 22.1.x (Homebrew libomp / llvm)
- thorvg.benchmark,
--backend=cpu --scene=default, default --threads=4, 2560x1440
Results
multiimagebench (5000 instances from 25 assets, 1000 frames) x2.03
./build/multiimagebench_thorvg_sdl
KMP_BLOCKTIME=1 ./build/multiimagebench_thorvg_sdl
Precedent
These projects set blocktime or wait policy programmatically:
Libraries with many short regions per call (ncnn, NNAPI, TensorFlow, ggml) pick a small non-zero value and respect user overrides. Projects that pick 0 do so because they share cores with other pools, or for power reasons. ThorVG's case, thousands of small regions per frame, is the former.
Summary
On arm64 macOS, LLVM libomp defaults
KMP_BLOCKTIMEto 0.Worker threads therefore go to sleep right after every
#pragma omp parallelregion,and the next region has to wake them again through a condition variable.
The SW engine opens many small parallel regions per frame (one per image, or one per shape).
On this platform the wake/sleep cost exceeds the work being parallelized.
Setting only
KMP_BLOCKTIME=1(1 ms) doubles the CPU benchmark FPS.We should either:
KMP_BLOCKTIMEin the benchmark runs so the numbers reflect ThorVG and not the libomp idle policy, orkmp_set_blocktime()inside the SW engine when OpenMP is enabled. ( LLVM/Intel libomp)Environment
-Dextra=openmp(the meson default), libomp 22.1.x (Homebrewlibomp/llvm)--backend=cpu --scene=default, default--threads=4, 2560x1440Results
multiimagebench (5000 instances from 25 assets, 1000 frames) x2.03
Precedent
These projects set blocktime or wait policy programmatically:
set_kmp_blocktime()->kmp_set_blocktime()around eachExtractor::extract(), restored afterOption::openmp_blocktime)ScopedOpenmpSettings:kmp_set_blocktime(20), restored on destructionKMP_BLOCKTIMEsetenv("KMP_BLOCKTIME", "200", 0)when built with OpenMPOMP_WAIT_POLICY=ACTIVE+KMP_BLOCKTIME=30msunless the user set themKMP_BLOCKTIME=1#if __APPLE__ && _OPENMP: works around blocktime=0 on Apple Siliconkmp_set_blocktime(0)KMP_BLOCKTIME=0unless set, to avoid spinning next to TBB's poolKMP_BLOCKTIME=0unless set (Redis module sharing the host process)Libraries with many short regions per call (ncnn, NNAPI, TensorFlow, ggml) pick a small non-zero value and respect user overrides. Projects that pick 0 do so because they share cores with other pools, or for power reasons. ThorVG's case, thousands of small regions per frame, is the former.