Skip to content

Bound memory and make per-mention cost independent of data size - #19

Merged
gkostkowski merged 12 commits into
OP-TED:developfrom
meaningfy-ws:develop
Sep 29, 2026
Merged

gkostkowski merged 12 commits into
OP-TED:developfrom
meaningfy-ws:develop

Conversation

@schivmeister

Copy link
Copy Markdown
Contributor

Summary

  • Prevent per-request Splink and DuckDB state from accumulating; use a configurable, disk-backed DuckDB store shared with Splink.
  • Persist and reload the trained model and search space so resolution can warm-start after a restart.
  • Process queued mentions in bounded batches, reducing repeated scoring and persistence overhead.
  • Add name-similarity blocking and update clustering behavior to reduce unnecessary candidate comparisons.
  • Add payload-free lifecycle logs, including Redis queue wait time from enqueue to ERE dequeue and timings for each processing stage.
  • Validate LinkTable column lengths to prevent silent link loss.

Compatibility

The Redis queue and response format remain unchanged. Responses are emitted in request order, in batches. Redis 6.2 or newer is required.

Matching behavior changes: the cluster threshold changes from 0.20 to 0.7, and weak-tail candidates may no longer be returned.

Logging and queue-delay measurement

The previous log-based measurement -- comparing when a request was sent to Redis with when the ERE received it -- is no longer possible from the old log entries; the raw request/receipt log format has been replaced. The new payload-free ERE lifecycle logs report queue-wait time when a request is dequeued, along with timings for subsequent processing stages. Use that queue-wait value to measure Redis-to-ERE delay.

Upgrade warning

Stop the ERE and remove the existing DuckDB file before updating. This version does not migrate the old database; it rebuilds the database, discarding its persisted ERE state. Back up or plan for a clean rebuild before proceeding.

Validation

Unit, feature, integration, and end-to-end tests, along with internal profiling, were run. Internal results indicate improved memory stability and throughput.

meanigfy and others added 12 commits September 17, 2026 20:32
Set up spec-driven work in this repository so changes are shaped, specified and
tracked as durable artifacts instead of ad-hoc task files.

- openspec/config.yaml pins the `meaningfy` schema and injects the project
  context (cosmic-python layering, src/ere package, develop branch, commit and
  artifact conventions) plus the three per-artifact rules.
- openspec/schemas/meaningfy/ is a pinned copy of the schema and its templates
  (OpenSpec 1.4.1); openspec/specs/ is the durable truth, openspec/changes/ the
  in-flight work.
- Install the /opsx:* commands and skills (core profile: propose, explore,
  apply, sync, archive).
- Makefile: `make check-specs` runs `openspec validate --all --strict`, and
  `make all-quality-checks` now includes it.
- AGENTS.md becomes the canonical agent instruction file and documents the
  golden thread (EPIC -> PLAN -> specs -> commit); CLAUDE.md imports it.
…ecs, plan)

The ERE slowed down during OP's nightly run (0.2 s to 0.9 s per request, queue
waits up to 47 min) and grew in memory with uptime. This records the shaped bet
and the measurements behind it.

- proposal.md: the EPIC in two parts. Part 1 stops the leak and makes
  per-mention cost independent of the stored corpus; Part 2 lowers that cost.
  Decisions DEC-1..DEC-28 carry their rationale: single-threaded consumption,
  no crash recovery, no clustering-correctness guarantee, train once and freeze,
  on-disk DuckDB by default, no retention policy, blocking by name similarity,
  cluster threshold 0.7, no `country_code` comparison, database reset on upgrade.
- design.md: how (D-1..D-23), the measurement plan with pass and cancel
  criteria, test-pinned interfaces, error matrix, risks and review findings.
- specs/: four capabilities as RFC-2119 requirements with Given/When/Then
  scenarios: resolution-resource-bounds, resolution-model-lifecycle,
  mention-lifecycle-tracing, candidate-scoring, batched-resolution.
- tasks.md: the executable plan, ticked as the work landed.
- inputs/: the evidence. The original memory analysis, the blocking experiments
  (two rounds on 5,497 real organisations), the Splink batch-cost spike, and the
  25k profiles before and after.
…of corpus size

The ERE leaked a DuckDB table per request, forgot its search space on restart,
and scored every stored mention of the same country. Per-mention latency grew
from 87 ms to 352 ms over 25k mentions and memory grew with uptime. It now stays
flat, and a backlog is drained in bites instead of one request at a time.

Resource bounds
- Release Splink's per-request prediction table, new-record view and query
  trackers after every scoring call; unregister temporary frames; a cleanup
  failure never fails the request.
- Index the mention and cluster lookups; build response candidates from the
  links just computed instead of re-reading stored links; drop the unconditional
  full membership load.
- Store only each mention's best `top_n` links and keep `similarities` unindexed:
  it is written on the request path and read only by re-submissions.
- DuckDB now runs on disk by default and is configurable per deployment through
  ERE_DUCKDB_STORAGE, ERE_DUCKDB_MEMORY_LIMIT, ERE_DUCKDB_THREADS and
  ERE_DUCKDB_TEMP_DIR. Splink shares that database, so the corpus is stored once
  and DuckDB can spill to the data volume instead of the read-only working
  directory.

Model lifecycle
- Splink reads the search space from the persisted `mentions` table, so a restart
  no longer empties it and mentions stored earlier are matched again.
- Training runs once at 200 mentions, is persisted next to the database and
  reloaded on start; stored similarity scores are never recomputed. Background
  training hands its parameters to the resolver thread, which never shares a
  DuckDB connection across threads.

Candidate scoring
- Block on country plus Jaro-Winkler similarity of a stored normalised legal
  name. On the 5,497-organisation reference corpus this scores 2.4 pairs per
  mention instead of 267, keeps every reference pair, and cuts merges of
  unrelated organisations from 93 % to below 0.1 %.
- Raise the cluster threshold to 0.7 and remove the `country_code` comparison,
  which only inflated scores because blocking already guarantees it agrees.

Throughput
- Take waiting requests in bites sized by recent processing speed (about two
  seconds of work), capped by count and payload bytes, with the rest left in
  Redis. One Splink call, one transaction and one pipelined response push per
  bite; mentions in a bite can match each other; a failed bite rolls back and is
  resolved one request at a time so failures stay isolated.
- A backlog of 25k requests drains at ~320 mentions/s, against ~12/s before.

Observability
- Structured lifecycle events per request (dequeued, parsed, guard, scored,
  clustered, persisted, responded) with per-stage timings, queue wait and one
  summary line, plus one event per bite. Request payloads and mention attributes
  are never logged, at any level.

Architecture
- Ports live with the domain in ere.models.ports; ere.entrypoints.bootstrap is
  the composition root that reads the environment and wires concrete adapters;
  Redis list mechanics moved to an adapter. A new import-linter contract forbids
  ere.services from importing ere.adapters.

Configuration
- resolver_compound.yaml and resolver_multirule.yaml could never start the ERE
  (no `entity_fields`), blocked on a `city` field the RDF mapping does not
  provide, and assumed 30 % of random pairs match. Fixed, and every shipped
  configuration is now exercised by a test.

Tests: 314 unit and BDD tests (96 % coverage), Gherkin features per capability,
Splink-backed scenarios on real organisations, and profiling and spike tools
under test/stress.

BREAKING CHANGE: databases created before this change lack the normalised-name
column; the ERE refuses to start on them and the file must be deleted.
BREAKING CHANGE: blocking, cluster threshold and comparisons changed, so cluster
assignments and weak-tail candidates differ from previous runs.
BREAKING CHANGE: the raw request body is no longer logged at INFO.
The skills now call `node .gitnexus/run.cjs <command>` instead of `npx`, and
document how to regenerate the gitignored runner and how to work around the npm
11.x install crash.

Adds the flat `.claude/skills/gitnexus-*` layout shipped by the newer release
alongside the existing nested `.claude/skills/gitnexus/` copies; the two hold the
same skills, so one layout should be removed once it is clear which the tooling
reads.
Three generations of the same skills were present: `.claude/skills/gitnexus/<topic>/`,
`.claude/skills/gitnexus/gitnexus-<topic>/` and the flat `.claude/skills/gitnexus-<topic>/`
that the tooling reads today. Remove the nested directory and point the last
references in .claude/CLAUDE.md at the flat paths.
Re-runs the single-path and backlog profiles after the review refactors and
compares them with the Part 2 run: throughput is within 1.5 % at every
checkpoint, single-request mean latency is 1-5 % lower, RSS within 9 MB and the
catalog stays at 4 objects. The per-stage split is now valid for the whole run
and confirms the Splink call is 89 % of a request.
Gathers every measurement of the memory-improvement change into one file:
the leaking baseline, the post-Part 1 and post-Part 2 profiles at 25k, the
post-refactor rerun and the S1/S5 spikes, with a stage-by-stage summary and
the open caveats. Linked from tasks.md.
Add post-initialization check for equal length of LinkTable columns.

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Corrected the abbreviation 'ie' to 'i.e.' and updated the class reference from 'ere.models.ere.ERERequest' to 'erspec.models.ere.ERERequest'.

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
feat(resolution): bound memory and make per-mention cost independent of corpus size
@gkostkowski
gkostkowski merged commit b7d67b2 into OP-TED:develop Sep 29, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants