Skip to content

Reframe the vocabulary-estimation series around its learner purpose #22

Description

@jamiepratt

Goal

Reframe the opening vocabulary-estimation articles as a Jamie-first curriculum that also stands alone for any educated lay reader. Make the learner-facing purpose explicit, use the simpler pair-count estimand as a deliberate learning step, and preserve the complete executable theory-to-algorithm history.

Decisions

  • Number the opening sequence from 1: workflow, purpose, Bayes, Proposal 1, Proposal 2.
  • The intermediate estimand is receptive knowledge of lemma–surface-form pairs.
  • The eventual learner-facing metric is estimated receptive Polish lemmas.
  • Call the research scorers Proposal 1 and Proposal 2 in reader-facing prose while retaining exact algorithm IDs.
  • Keep the current LexiBench scorer distinct from both proposals.
  • Give every article a concise standalone narrative, article-specific structure, progressive technical disclosure, responsive contents, and compact series navigation.
  • Retain every existing simulation and add a vocabulary bridge to the Bayes introduction.
  • Deliver Articles 1–4 as coordinated vertical slices; Proposal 2 remains governed by Explain why continuous-frequency V2 failed without discarding its signal #16.

Scope

The Civitas publication repository owns executable article source, interactive assets, screenshot assets, tests, and rendered publication inputs. This workflow repository owns orchestration, durable series guidance, and the published submodule pointer.

Non-goals

  • Design how pair evidence becomes latent lemma knowledge.
  • Decide detailed question-bank sampling or how many forms and contexts each lemma needs.
  • Design or deploy adaptive selection.
  • Change either historical scoring proposal or immutable V2 evidence.
  • Claim that simulations validate representative learners.
  • Claim that Proposal 1 or Proposal 2 powers LexiBench.
  • Resolve global Civitas configuration ownership tracked by Decide ownership of the global Quarto *.cljc resource rule #12.

Acceptance Criteria

  • The series order is workflow → purpose → Bayes → Proposal 1 → Proposal 2.
  • Pair knowledge is consistently presented as an intermediate estimand and estimated receptive Polish lemmas as the eventual learner metric.
  • The deployed LexiBench scorer, Proposal 1, and Proposal 2 cannot be confused.
  • Articles 1–4 use concise standalone narratives, progressive technical disclosure, responsive article contents, and compact series navigation.
  • All existing simulations remain available and the Bayes article adds the agreed vocabulary bridge.
  • Proposal 1 remains the current research target but not the deployed scorer; Proposal 2 remains non-promoted under Explain why continuous-frequency V2 failed without discarding its signal #16.
  • Model, software, and publication validation are reported separately.
  • Durable series decisions are recorded in root-owned guidance without duplicating future work outside GitHub Issues.

Orchestration

Issue kind: parent PRD

Children:

Related:

Notes

The live LexiBench product currently estimates recognized Polish lemmas in sentence context using a separate deployed scorer. Production screenshots are current-product evidence only and must be stored immutably with capture provenance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions