Skip to content

cpu_engine: OpenMP regions are ~2x slower on Apple Silicon because libomp's default KMP_BLOCKTIME is 0 #11

Description

@Nor-s

Summary

On arm64 macOS, LLVM libomp defaults KMP_BLOCKTIME to 0.

Worker threads therefore go to sleep right after every #pragma omp parallel region,
and the next region has to wake them again through a condition variable.

The SW engine opens many small parallel regions per frame (one per image, or one per shape).
On this platform the wake/sleep cost exceeds the work being parallelized.

Setting only KMP_BLOCKTIME=1 (1 ms) doubles the CPU benchmark FPS.

We should either:

  • (A) set KMP_BLOCKTIME in the benchmark runs so the numbers reflect ThorVG and not the libomp idle policy, or
  • (B) call kmp_set_blocktime() inside the SW engine when OpenMP is enabled. ( LLVM/Intel libomp)

Environment

  • Apple M5 Pro (6P + 12E cores), macOS 26 (Darwin 25)
  • ThorVG CPU engine, -Dextra=openmp (the meson default), libomp 22.1.x (Homebrew libomp / llvm)
  • thorvg.benchmark, --backend=cpu --scene=default, default --threads=4, 2560x1440

Results

multiimagebench (5000 instances from 25 assets, 1000 frames) x2.03

./build/multiimagebench_thorvg_sdl
KMP_BLOCKTIME=1 ./build/multiimagebench_thorvg_sdl

Precedent

These projects set blocktime or wait policy programmatically:

project what value code
Tencent/ncnn set_kmp_blocktime() -> kmp_set_blocktime() around each Extractor::extract(), restored after 20 ms (Option::openmp_blocktime) cpu.cpp#L3284-L3291, net.cpp#L3216-L3217
Android NNAPI CPU executor RAII ScopedOpenmpSettings: kmp_set_blocktime(20), restored on destruction 20 ms CpuExecutor.cpp#L1874-L1901 (AOSP mirror)
TensorFlow (oneDNN/MKL build) sets blocktime once unless the user set KMP_BLOCKTIME 1 ms threadpool_device.cc#L90-L98
ggml / llama.cpp setenv("KMP_BLOCKTIME", "200", 0) when built with OpenMP 200 ms ggml-cpu.c#L3909-L3924
primecount app sets OMP_WAIT_POLICY=ACTIVE + KMP_BLOCKTIME=30ms unless the user set them 30 ms main.cpp#L85-L110
PyTorch / IPEX CPU launchers KMP_BLOCKTIME=1 1 ms run_cpu.py#L434-L438, launcher_base.py#L314-L317
DynEarthSol #if __APPLE__ && _OPENMP: works around blocktime=0 on Apple Silicon active dynearthsol.cxx#L602-L613
LAMMPS (INTEL) kmp_set_blocktime(0) 0 fix_intel.cpp#L153-L156
oneTBB python module KMP_BLOCKTIME=0 unless set, to avoid spinning next to TBB's pool 0 __init__.py#L323-L324
FalkorDB KMP_BLOCKTIME=0 unless set (Redis module sharing the host process) 0 module_init.rs#L232-L240

Libraries with many short regions per call (ncnn, NNAPI, TensorFlow, ggml) pick a small non-zero value and respect user overrides. Projects that pick 0 do so because they share cores with other pools, or for power reasons. ThorVG's case, thousands of small regions per frame, is the former.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions