You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
PYTHON-6065 Benchmark OpenTelemetry support - #3102
Adds tools/otel_bench.py, a benchmark harness covering the tracing config variants for DRIVERS-3620, plus per-op CPU-time recording in test/performance/ (PERF_CPU_TIME=1).
Hot-path optimizations: defer expensive span attributes until after the sampling decision, cache connection-static attributes, resolve the tracing option once at client construction, skip the duration timedelta when logging/APM are off.
Overhead is a fixed per-op cost (~50-55 µs/op at 1% sampling, ~130 µs/op always-on); span creation drops 18.0 → 3.0 µs/op; disabled tracing is at parity (≤2%) with the pre-OTel baseline. No public APIs changed.
16-core x86_64, MongoDB 8.0.4 localhost, CPython 3.9, 5 interleaved repetitions with rotated config order, medians; no exporter or span processor installed.
Throughput (median MB/s, overhead vs untraced baseline in parens):
config
sync insertOne
sync findOne
async insertOne
async findOne
off (baseline)
1.07
5.21
0.61
3.19
api-only
0.96 (+10.2%)
4.79 (+8.0%)
0.57 (+6.1%)
3.04 (+4.5%)
sdk-ratio (1%)
0.87 (+19.1%)
4.36 (+16.3%)
0.54 (+11.3%)
2.88 (+9.6%)
sdk-always
0.69 (+36.0%)
3.56 (+31.6%)
0.46 (+24.4%)
2.50 (+21.7%)
CPU time per operation (µs/op median, delta vs baseline in parens):
config
sync insertOne
sync findOne
async insertOne
async findOne
off (baseline)
161
222
396
462
api-only
187 (+25)
246 (+24)
422 (+26)
484 (+22)
sdk-ratio (1%)
214 (+53)
278 (+56)
447 (+51)
511 (+50)
sdk-always
291 (+130)
355 (+133)
526 (+131)
592 (+130)
Tracing disabled vs pre-tracing commit (d5934e6):
task
sync
async
SmallDocInsertOne
+1.8%
-0.8%
FindOneByID
+0.8%
+0.6%
Overhead barely depends on sampling rate: a span must be started to learn it won't record, so attribute collection and sampler machinery run on every command. Micro-benchmarks on ~160-460 µs ops magnify the fixed cost; real workloads amortize it.
Test Plan
New tests in test/asynchronous/test_otel.py (sync suite mirrored) cover the deferred attributes, the connection cache, and a hot-path regression test.
python tools/otel_bench.py --verify checks span wiring end-to-end; full run via python tools/otel_bench.py --reps 5.
…butes
Fix tracing.enabled so an explicit client value overrides the
OTEL_PYTHON_INSTRUMENTATION_MONGODB_ENABLED environment variable,
fix db.query.text truncation so budgets smaller than the "..."
marker still honor the bound, and fix collection-name extraction so
user and role management commands do not expose usernames as
db.collection.name.
Benchmark tracing overhead per the OpenTelemetry spec's performance
requirements with tools/otel_bench.py, and optimize the
per-command hot path based on the results:
- Defer expensive span attributes (db.query.summary, db.mongodb.lsid,
db.mongodb.txn_number, db.query.text) until after the sampler's
decision, so unsampled and no-op spans skip building them entirely.
db.query.text is the big win: it serialized the command to extended
JSON on every single command.
- Cache connection-static span attributes on the connection (keyed by
server connection id) instead of rebuilding them per command.
- Resolve the client's tracing option against the environment once, at
MongoClient construction, so no command consults the environment.
- Skip building the duration timedelta when neither command logging nor
APM events are enabled; a tracing-only client doesn't need it.
Add PERF_CPU_TIME to the DriverBench performance tests to record
per-operation CPU time, which otel_bench.py uses to report the fixed
CPU cost of each tracing configuration (tracing off, api-only, SDK with
TraceIdRatioBased sampling, SDK always-on).
- Never record a literal 0 for db.mongodb.cursor_id: an exhausted
getMore reply records the cursor id sent in the command, and a
cursor-creating command whose reply exhausts the cursor omits the
attribute.
- Preserve the construction-time tracing environment snapshot when an
SRV client rebuilds its ClientOptions.
- Key the benchmark's persisted pre-OTel comparison by API so the sync
and async rows both survive.
- Interpret --data-dir and --output-dir as absolute paths before
launching benchmark children that run in a different cwd.
- Pin opentelemetry-sdk>=1.20.0 for the test SDK so lowest-direct
resolution cannot drag opentelemetry-api below the driver's floor,
and skip the tracing env-deferral tests when opentelemetry is
not installed.
The failure path never sets the required error.type attribute on command spans. The OpenTelemetry specification requires it for every failed command: it should be the string server response code when code is present, otherwise the exception class name. Please set it alongside db.response.status_code before ending the span so server and transport failures are distinguishable in telemetry.
Avoid CPU timing overhead unless explicitly enabled
test/performance/async_perf_test.py:118
time.process_time() is called for every benchmark iteration even when PERF_CPU_TIME is unset. Because this timer surrounds the measured operation, the normal performance suite now pays the extra clock calls and its MB/s results include instrumentation that the opt-in CPU metric is supposed to avoid; only collect CPU timestamps when RECORD_CPU_TIME is enabled.
Avoid CPU timing overhead unless explicitly enabled
test/performance/perf_test.py:118
time.process_time() is called for every benchmark iteration even when PERF_CPU_TIME is unset. Because this timer surrounds the measured operation, the normal performance suite now pays the extra clock calls and its MB/s results include instrumentation that the opt-in CPU metric is supposed to avoid; only collect CPU timestamps when RECORD_CPU_TIME is enabled.
- Set error.type on failed command spans: the server response code
alongside db.response.status_code for server failures, otherwise the
exception class name, so server and transport failures are
distinguishable.
- Only collect CPU time in the performance tests when PERF_CPU_TIME is
enabled, so the standard suite no longer pays clock calls inside the
measured window.
Prototype the spec's operation-span layer in the driver: one CLIENT
span per driver operation, started before server selection, with the
operation's command spans nested beneath it via the current OTel
context. Wired at the _run_operation (queries) and _retryable_write
funnels with db/collection plumbed from the public call sites; getMore
and cursor iteration remain PYTHON-5947 work. On failure the span
records the exception and sets error.type to the exception class name.
This replaces the bench-side prototype: the benchmark's existing
tracing configurations now measure command and operation spans
together, and otel_bench --verify asserts the nesting. Also annotate
the failure-path test's options dict as TracingOptions, fixing 4 mypy
arg-type errors from the previous commit.
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PYTHON-6065
Changes in this PR
tools/otel_bench.py, a benchmark harness covering the tracing config variants for DRIVERS-3620, plus per-op CPU-time recording intest/performance/(PERF_CPU_TIME=1).tracingoption once at client construction, skip the durationtimedeltawhen logging/APM are off.Results
16-core x86_64, MongoDB 8.0.4 localhost, CPython 3.9, 5 interleaved repetitions with rotated config order, medians; no exporter or span processor installed.
Throughput (median MB/s, overhead vs untraced baseline in parens):
CPU time per operation (µs/op median, delta vs baseline in parens):
Tracing disabled vs pre-tracing commit (
d5934e6):Overhead barely depends on sampling rate: a span must be started to learn it won't record, so attribute collection and sampler machinery run on every command. Micro-benchmarks on ~160-460 µs ops magnify the fixed cost; real workloads amortize it.
Test Plan
test/asynchronous/test_otel.py(sync suite mirrored) cover the deferred attributes, the connection cache, and a hot-path regression test.python tools/otel_bench.py --verifychecks span wiring end-to-end; full run viapython tools/otel_bench.py --reps 5.Checklist
Checklist for Author
Checklist for Reviewer