Skip to content

PYTHON-6065 Benchmark OpenTelemetry support - #3102

Draft
blink1073 wants to merge 10 commits into
mongodb:mainfrom
blink1073:PYTHON-6065
Draft

blink1073 wants to merge 10 commits into
mongodb:mainfrom
blink1073:PYTHON-6065

Conversation

@blink1073

@blink1073 blink1073 commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

PYTHON-6065

Changes in this PR

  • Adds tools/otel_bench.py, a benchmark harness covering the tracing config variants for DRIVERS-3620, plus per-op CPU-time recording in test/performance/ (PERF_CPU_TIME=1).
  • Hot-path optimizations: defer expensive span attributes until after the sampling decision, cache connection-static attributes, resolve the tracing option once at client construction, skip the duration timedelta when logging/APM are off.
  • Overhead is a fixed per-op cost (~50-55 µs/op at 1% sampling, ~130 µs/op always-on); span creation drops 18.0 → 3.0 µs/op; disabled tracing is at parity (≤2%) with the pre-OTel baseline. No public APIs changed.
  • The first two commits are from PYTHON-5945 Add OpenTelemetry Simple Command Support #3071. The second two add the optimizations and the bench harness.

Results

16-core x86_64, MongoDB 8.0.4 localhost, CPython 3.9, 5 interleaved repetitions with rotated config order, medians; no exporter or span processor installed.

Throughput (median MB/s, overhead vs untraced baseline in parens):

config sync insertOne sync findOne async insertOne async findOne
off (baseline) 1.07 5.21 0.61 3.19
api-only 0.96 (+10.2%) 4.79 (+8.0%) 0.57 (+6.1%) 3.04 (+4.5%)
sdk-ratio (1%) 0.87 (+19.1%) 4.36 (+16.3%) 0.54 (+11.3%) 2.88 (+9.6%)
sdk-always 0.69 (+36.0%) 3.56 (+31.6%) 0.46 (+24.4%) 2.50 (+21.7%)

CPU time per operation (µs/op median, delta vs baseline in parens):

config sync insertOne sync findOne async insertOne async findOne
off (baseline) 161 222 396 462
api-only 187 (+25) 246 (+24) 422 (+26) 484 (+22)
sdk-ratio (1%) 214 (+53) 278 (+56) 447 (+51) 511 (+50)
sdk-always 291 (+130) 355 (+133) 526 (+131) 592 (+130)

Tracing disabled vs pre-tracing commit (d5934e6):

task sync async
SmallDocInsertOne +1.8% -0.8%
FindOneByID +0.8% +0.6%

Overhead barely depends on sampling rate: a span must be started to learn it won't record, so attribute collection and sampler machinery run on every command. Micro-benchmarks on ~160-460 µs ops magnify the fixed cost; real workloads amortize it.

Test Plan

  • New tests in test/asynchronous/test_otel.py (sync suite mirrored) cover the deferred attributes, the connection cache, and a hot-path regression test.
  • python tools/otel_bench.py --verify checks span wiring end-to-end; full run via python tools/otel_bench.py --reps 5.
  • Passing build

Checklist

Checklist for Author

  • Did you update the changelog (if necessary)?
  • Is there test coverage?
  • Is any followup work tracked in a JIRA ticket? If so, add link(s).

Checklist for Reviewer

  • Does the title of the PR reference a JIRA Ticket?
  • Do you fully understand the implementation? (Would you be comfortable explaining how this code works to someone else?)
  • Is all relevant documentation (README or docstring) updated?

…butes

Fix tracing.enabled so an explicit client value overrides the
OTEL_PYTHON_INSTRUMENTATION_MONGODB_ENABLED environment variable,
fix db.query.text truncation so budgets smaller than the "..."
marker still honor the bound, and fix collection-name extraction so
user and role management commands do not expose usernames as
db.collection.name.
Benchmark tracing overhead per the OpenTelemetry spec's performance
requirements with tools/otel_bench.py, and optimize the
per-command hot path based on the results:

- Defer expensive span attributes (db.query.summary, db.mongodb.lsid,
  db.mongodb.txn_number, db.query.text) until after the sampler's
  decision, so unsampled and no-op spans skip building them entirely.
  db.query.text is the big win: it serialized the command to extended
  JSON on every single command.
- Cache connection-static span attributes on the connection (keyed by
  server connection id) instead of rebuilding them per command.
- Resolve the client's tracing option against the environment once, at
  MongoClient construction, so no command consults the environment.
- Skip building the duration timedelta when neither command logging nor
  APM events are enabled; a tracing-only client doesn't need it.

Add PERF_CPU_TIME to the DriverBench performance tests to record
per-operation CPU time, which otel_bench.py uses to report the fixed
CPU cost of each tracing configuration (tracing off, api-only, SDK with
TraceIdRatioBased sampling, SDK always-on).
@codecov

codecov Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.75908% with 28 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
pymongo/_otel.py 88.60% 10 Missing and 12 partials ⚠️
pymongo/asynchronous/collection.py 60.00% 2 Missing ⚠️
pymongo/common.py 87.50% 1 Missing and 1 partial ⚠️
pymongo/synchronous/collection.py 66.66% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Cursor attributes, deferred SRV option resolution, relative paths, and persisted benchmark comparisons have correctness issues.

4 open findings
What changed in this PR

Adds optional OpenTelemetry command-span instrumentation and benchmarking support while minimizing disabled and unsampled tracing overhead.

Changes:

  • Implements configurable OpenTelemetry command spans and related tests.
  • Adds CPU-time metrics and an interleaved tracing benchmark harness.
  • Adds dependency, documentation, typing, and Evergreen integration.
File Description
.gitignore Ignores benchmark output.
.evergreen/​generated_configs/​variants.yml Adds the OTel test variant.
.evergreen/​scripts/​generate_config.py Generates the OTel variant.
.evergreen/​scripts/​setup_tests.py Installs OTel test dependencies and enables coverage.
.evergreen/​scripts/​utils.py Registers the OTel test type.
doc/​changelog.rst Documents command-span support.
justfile Adds OTel dependencies to typing runs.
pymongo/​_otel.py Implements command-span creation and completion.
pymongo/​_telemetry.py Integrates spans with command telemetry.
pymongo/​asynchronous/​command_runner.py Enables tracing and cancellation cleanup.
pymongo/​asynchronous/​mongo_client.py Documents tracing options.
pymongo/​asynchronous/​pool.py Caches connection span attributes.
pymongo/​client_options.py Resolves and exposes tracing configuration.
pymongo/​common.py Validates tracing options.
pymongo/​message.py Supports converting cancellation exceptions.
pymongo/​pool_shared.py Extends the connection telemetry protocol.
pymongo/​synchronous/​command_runner.py Mirrors command tracing integration.
pymongo/​synchronous/​mongo_client.py Mirrors tracing documentation.
pymongo/​synchronous/​pool.py Mirrors connection attribute caching.
pyproject.toml Registers the extra and pytest marker.
requirements/​opentelemetry.txt Declares the optional API dependency.
test/​asynchronous/​test_otel.py Tests asynchronous tracing behavior.
test/​performance/​async_perf_test.py Records asynchronous CPU-time metrics.
test/​performance/​perf_test.py Records synchronous CPU-time metrics.
test/​test_otel.py Provides the generated synchronous OTel tests.
tools/​otel_bench.py Adds the tracing benchmark harness.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread pymongo/_otel.py Outdated
Comment thread pymongo/client_options.py
Comment on lines +250 to +253
# Resolve the tracing option against the environment variables once:
# tracing cannot be changed after the client is constructed, so no
# command needs to consult the environment again.
self.__tracing = _otel._resolve_tracing_options(options.get("tracing"))
Comment thread tools/otel_bench.py Outdated
Comment thread tools/otel_bench.py
- Never record a literal 0 for db.mongodb.cursor_id: an exhausted
  getMore reply records the cursor id sent in the command, and a
  cursor-creating command whose reply exhausts the cursor omits the
  attribute.
- Preserve the construction-time tracing environment snapshot when an
  SRV client rebuilds its ClientOptions.
- Key the benchmark's persisted pre-OTel comparison by API so the sync
  and async rows both survive.
- Interpret --data-dir and --output-dir as absolute paths before
  launching benchmark children that run in a different cwd.
- Pin opentelemetry-sdk>=1.20.0 for the test SDK so lowest-direct
  resolution cannot drag opentelemetry-api below the driver's floor,
  and skip the tracing env-deferral tests when opentelemetry is
  not installed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Five moderate findings remain unresolved.

1 open finding
3 resolved since last review
Previously missed (3)

In code that hasn't changed since last review

Medium severity Set error.type on failed command spans

pymongo/​_otel.py:378

The failure path never sets the required error.type attribute on command spans. The OpenTelemetry specification requires it for every failed command: it should be the string server response code when code is present, otherwise the exception class name. Please set it alongside db.response.status_code before ending the span so server and transport failures are distinguishable in telemetry.

Medium severity Avoid CPU timing overhead unless explicitly enabled

test/​performance/​async_perf_test.py:118

time.process_time() is called for every benchmark iteration even when PERF_CPU_TIME is unset. Because this timer surrounds the measured operation, the normal performance suite now pays the extra clock calls and its MB/s results include instrumentation that the opt-in CPU metric is supposed to avoid; only collect CPU timestamps when RECORD_CPU_TIME is enabled.

Medium severity Avoid CPU timing overhead unless explicitly enabled

test/​performance/​perf_test.py:118

time.process_time() is called for every benchmark iteration even when PERF_CPU_TIME is unset. Because this timer surrounds the measured operation, the normal performance suite now pays the extra clock calls and its MB/s results include instrumentation that the opt-in CPU metric is supposed to avoid; only collect CPU timestamps when RECORD_CPU_TIME is enabled.

🧠 Review effort: Lite

- Set error.type on failed command spans: the server response code
  alongside db.response.status_code for server failures, otherwise the
  exception class name, so server and transport failures are
  distinguishable.
- Only collect CPU time in the performance tests when PERF_CPU_TIME is
  enabled, so the standard suite no longer pays clock calls inside the
  measured window.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Three moderate issues remain in pymongo/_otel.py, covering admin-database attributes, sampling-path test coverage, and getMore cursor IDs.

1 open finding

🧠 Review effort: Lite

Prototype the spec's operation-span layer in the driver: one CLIENT
span per driver operation, started before server selection, with the
operation's command spans nested beneath it via the current OTel
context. Wired at the _run_operation (queries) and _retryable_write
funnels with db/collection plumbed from the public call sites; getMore
and cursor iteration remain PYTHON-5947 work. On failure the span
records the exception and sets error.type to the exception class name.

This replaces the bench-side prototype: the benchmark's existing
tracing configurations now measure command and operation spans
together, and otel_bench --verify asserts the nesting. Also annotate
the failure-path test's options dict as TracingOptions, fixing 4 mypy
arg-type errors from the previous commit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants