Repository navigation
feat(integrations): add MongoDB-backed session and memory services - #7291
theshanbhag wants to merge 10 commits into
Conversation
Add MongoDbSessionService, which persists sessions, events, and app:/user:-scoped state in MongoDB with optimistic-concurrency revisions, and MongoDbMemoryService, which ingests session events into a keyword-indexed memory collection for long-term recall. Register the MONGODB_* experimental features (enabled by default) and export both services from google.adk.integrations.mongodb. Also add a README documenting the MongoDB integration (search toolset, both services, and MongoDbToolSettings) and add mongomock to the test dependencies for the new unit tests. NO_UNIT_GUIDE=The integration README added in this commit documents all MongoDB units; per-unit guides will follow when the experimental API stabilizes.
# Conflicts: # constraints-3.10.txt # constraints-3.11.txt # constraints-3.12.txt # constraints-3.13.txt # constraints-3.14.txt
The previous merge of main into add-memory-to-mongodb left MONGODB_SESSION_SERVICE's FeatureConfig unclosed (syntax error) and duplicated the MONGODB_TOOLSET enum member and registry entry.
…memory-to-mongodb
The merge combined upstream's regenerated pins with the branch's mongomock addition textually; regenerate with scripts/update_constraints.sh so the '# via' annotations match the stabilized generation flow CI checks for. No version pins change.
…memory-to-mongodb
…ngoDB - Atlas Automated Embedding (Preview): `use_mongodb_auto_embedding` and `mongodb_auto_embedding_model` on MongoDbToolSettings send the query as plain text (`query.text`) so Atlas embeds it with the index's Voyage AI model; the default mode keeps embedding queries with Google models via the genai client. - MongoDbSessionService now runs multi-document writes in a transaction (replica set / Atlas) with a cached fallback to sequential writes plus optimistic-concurrency revision checks on deployments without transaction support, and serializes appends with per-session locks. - Optional `ensure_indexes` on both services creates the recommended secondary indexes on first use. - MongoDbToolset, MongoDbSessionService and MongoDbMemoryService are picklable when constructed with `connection_string`, so agents using them can deploy to Agent Engine (cloudpickle boundary): the client is dropped at pickle time and rebuilt on the runtime. - Drop the MONGODB_* experimental feature flags; the integration is always on. - README: document embedding modes, secondary indexes, Agent Engine deployment, and a required-permissions (least-privilege) section covering the model-chosen collection risk and a before_tool_callback allowlist.
…memory-to-mongodb
|
Add MongoDbSessionService, which persists sessions, events, and app:/user:-scoped state in MongoDB with optimistic-concurrency revisions and multi-document transactions (with a cached fallback to sequential writes on deployments without transaction support), and MongoDbMemoryService, which ingests session events into a keyword-indexed memory collection for long-term recall. The search tools now support both embedding modes: Google embedding models via the genai client (default, MongoDbToolset, MongoDbSessionService, and MongoDbMemoryService are picklable when constructed with The integration README documents the embedding modes, optional secondary indexes ( NO_UNIT_GUIDE=The integration README added in this change documents all MongoDB units; per-unit guides will follow when the API stabilizes. Please ensure you have read the contribution guide before creating a pull request. Link to Issue or Description of Change1. Link to an existing issue (if applicable):
2. Or, if no issue exists, describe the change: Problem: ADK's MongoDB integration provided search tools only: no MongoDB-backed session or memory persistence, and no way to deploy agents using them to Agent Engine, where apps cross a cloudpickle boundary that a live Solution:
Testing PlanUnit Tests:
Manual End-to-End (E2E) Tests: Verified against a live MongoDB Atlas cluster seeded with a 10-product catalog (
Checklist
Additional contextWiring |
|
Thanks for the thorough PR, @theshanbhag! This is too big for us to review as one change: it's four separate features in about 2.4k lines. Could you split it into separate PRs that we can review and land one at a time? Roughly:
Some other notes:
Happy to review each piece as it comes in. |
Add MongoDbSessionService, which persists sessions, events, and app:/user:-scoped state in MongoDB with optimistic-concurrency revisions and multi-document transactions (with a cached fallback to sequential writes on deployments without transaction support), and MongoDbMemoryService, which ingests session events into a keyword-indexed memory collection for long-term recall.
The search tools now support both embedding modes: Google embedding models via the genai client (default,
queryVector), and Atlas Automated Embedding (Preview), whereMongoDbToolSettings.use_mongodb_auto_embeddingsends the query as plain text (query.text) and Atlas embeds it with the index's Voyage AI model — no Google embedding credentials needed.MongoDbToolset, MongoDbSessionService, and MongoDbMemoryService are picklable when constructed with
connection_string, so agents using them deploy cleanly to Agent Engine (which packages apps with cloudpickle): the client is dropped at pickle time and rebuilt on the runtime. The MONGODB_* experimental feature flags are dropped; the integration is always on.The integration README documents the embedding modes, optional secondary indexes (
ensure_indexes), Agent Engine deployment, and a required-permissions (least-privilege) section — including the fact that the model chooses the collection at run time, with abefore_tool_callbackallowlist example. mongomock is added to the test dependencies for the new unit tests.NO_UNIT_GUIDE=The integration README added in this change documents all MongoDB units; per-unit guides will follow when the API stabilizes.
Please ensure you have read the contribution guide before creating a pull request.
Link to Issue or Description of Change
1. Link to an existing issue (if applicable):
2. Or, if no issue exists, describe the change:
Problem:
ADK's MongoDB integration provided search tools only: no MongoDB-backed session or memory persistence, and no way to deploy agents using them to Agent Engine, where apps cross a cloudpickle boundary that a live
pymongo.MongoClient(sockets, locks, background threads) cannot survive. Query embedding also required Google embedding credentials even when the Atlas cluster could embed queries itself.Solution:
MongoDbSessionService(sessions, events, and app:/user:-scoped state with optimistic-concurrency revisions; multi-document writes run in transactions where the deployment supports them, with per-session locking and a cached fallback otherwise) andMongoDbMemoryService(keyword-indexed long-term recall with idempotent ingestion).MongoDbToolSettings.use_mongodb_auto_embedding/mongodb_auto_embedding_model, alongside the default Google embedding mode.connection_string, they drop the live client at pickle time and rebuild it on the destination, so Agent Engine deployments work unchanged.ensure_indexes, Agent Engine deployment, and least-privilege database setup (read-only search user vs. read-write state user, per-collection scoping, and abefore_tool_callbackcollection allowlist).Testing Plan
Unit Tests:
pytest tests/unittests/integrations/mongodb/ tests/unittests/features/→ 126 passed (98 MongoDB integration tests, mongomock-backed: search tools in both embedding modes, toolset exposure/injection, pickle round-trips, session-service transactions and fallback, memory-service recall/idempotency; plus the 28 feature-registry tests). All pre-commit hooks (pyink, isort, ruff, mdformat, codespell, compliance checks) pass.Manual End-to-End (E2E) Tests:
Verified against a live MongoDB Atlas cluster seeded with a 10-product catalog (
productswith anautoEmbedvector index,products_googlewithtext-embedding-005embeddings under a regular vector index, both with a full-text index):mongodb_vector_searchandmongodb_hybrid_searchin Atlas auto-embedding mode return ranked, filtered results withsearch_score(e.g. "cordless robot vacuum for pet hair" ranks "Robot Vacuum Pro" first; acategoryfilter restricts hits).text-embedding-005) return equivalent rankings.pickle.loads(pickle.dumps(root_agent))round-trips the agent → toolset chain with tools intact and clients rebuilt.Checklist
Additional context
Wiring
MongoDbSessionService/MongoDbMemoryServiceinto Agent Engine deployments directly (as managed-service alternatives) is planned for a follow-up; today Agent Engine deployments use the managed Vertex AI session and memory services, while the MongoDB services back self-hosted runners (adk run,adk web, Cloud Run).