Skip to content

If there is no title in the output, we should add a default fallback of "Data Management Plan" because title is a required field in DMP Tool schema #2

Description

@bofstein

I noticed while testing that sometimes the PDF came out without any title, but that is a required field in the DMP Tool schema. We should add a post-processing step to enter a default title if there isn't one at the end.

I plan try to take this one on myself with AI-assisted coding just as a learning experience - feel free to reject anything that's not appropriate, and please review carefully anything I submit from it! I started the process with ChatGPT but am now leaving for the day; I'm entering the text here for myself to follow up on next week.

My prompt: Some of the PDFs end up not having any title in the JSON, which is required for our purpose. DMP Bridge has a step that applies a rules-based cleanup to the LLM-labeled text to get it to meet basic requirements. If I wanted to submit a PR that would add a rule to this step that says "if the 'title' field for the whole DMP is empty, use the text 'Data Management Plan' instead of leaving it blank", how would I go about doing that?

ChatGPT response:
Yes. This is a small, well-contained change. The safest implementation is to apply the default after the existing spreadsheet-based rules have finished.

That placement matters: the current rules use the document title as an input when filling certain section and question fields. Setting the default too early could unintentionally put “Data Management Plan” into those fields as well.

1. Get a Git-enabled copy of the repository

Because your existing copy came from a ZIP, use a fresh clone for the PR.

  1. Install [GitHub Desktop](https://desktop.github.com/) and sign in.
  2. Open the [DMPBridge repository](https://github.com/fairdataihub/dmpbridge).
  3. Click Code → Open with GitHub Desktop.
  4. Choose a local folder and click Clone.
  5. In GitHub Desktop, select Current branch → New branch.
  6. Name it:
fix/default-empty-dmp-title

If you do not have write access to the repository, GitHub Desktop will help create a fork when you publish the branch. GitHub also documents the [clone/fork workflow](https://docs.github.com/en/desktop/adding-and-cloning-repositories/cloning-and-forking-repositories-from-github-desktop?platform=windows).

2. Point Python at the cloned copy

Open the cloned repository in PowerShell. In GitHub Desktop, Repository → Open in PowerShell may do this for you.

Run:

py -3.14 -m pip install -e .

This changes your editable DMPBridge installation from the ZIP copy to your new Git clone.

If you want to edit through JupyterLab, run:

py -3.14 -m jupyter lab

Use JupyterLab’s file browser and text editor to edit the following Python files. You do not need to create another notebook.

3. Add the fallback rule

Open:

dmpbridge/evaluation/annotation_rules.py

Near the other module-level constants, add:

DEFAULT_DOCUMENT_TITLE = "Data Management Plan"

Then find the end of apply_new_annotation_rules(). After its section/question processing loops, but immediately before return data, add:

    # Apply the document-level fallback after the per-question rules.
    if not (template.get("title") or "").strip():
        template["title"] = DEFAULT_DOCUMENT_TITLE

    return data

The indentation is important: the if must be outside the for section ... loop and at the same indentation level as that loop.

The relevant existing implementation is in [annotation_rules.py](https://github.com/fairdataihub/dmpbridge/blob/main/dmpbridge/evaluation/annotation_rules.py).

This condition treats all of these as empty:

  • ""
  • None
  • " "—a title containing only spaces

It leaves every meaningful existing title unchanged.

4. Update the tests

Open:

tests/test_annotation_rules.py

There is currently a test named approximately:

test_document_title_is_never_modified

That test reflects the old behavior, so replace or split it. Add tests along these lines:

@pytest.mark.parametrize("title", ["", None, "   "])
def test_empty_document_title_gets_default(title):
    out = ar.apply_new_annotation_rules(_doc(title=title))

    assert (
        out["narrative"]["template"]["title"]
        == "Data Management Plan"
    )


@pytest.mark.parametrize(
    "title",
    ["Existing title", "Specific Project DMP"],
)
def test_existing_document_title_is_not_modified(title):
    out = ar.apply_new_annotation_rules(_doc(title=title))

    assert out["narrative"]["template"]["title"] == title

I would also add this regression test to prove that the default does not leak into empty section or question titles:

def test_default_document_title_does_not_change_empty_question_fields():
    out = ar.apply_new_annotation_rules(
        _doc(
            title="",
            sec_title="",
            sec_desc="",
            q_text="",
        )
    )

    template = out["narrative"]["template"]

    assert template["title"] == "Data Management Plan"
    assert template["section"][0]["title"] == ""
    assert template["section"][0]["question"][0]["text"] == ""

You may need to adjust the argument names if the existing _doc() test helper uses slightly different names. Follow the spelling already used elsewhere in [test_annotation_rules.py](https://github.com/fairdataihub/dmpbridge/blob/main/tests/test_annotation_rules.py).

5. Run the tests

From PowerShell in the cloned repository:

py -3.14 -m pip install pytest
py -3.14 -m pytest tests/test_annotation_rules.py

If those pass, run the entire test suite:

py -3.14 -m pytest

These tests should not need Ollama or Gemma because they test only the deterministic cleanup function.

If the full suite has unrelated Python 3.14 compatibility errors, record those in the PR, but make sure the annotation-rules tests pass.

6. Commit and submit the PR

In GitHub Desktop:

  1. Review the changed files.
  2. Enter the commit summary:
Add fallback for empty DMP titles
  1. Click Commit to fix/default-empty-dmp-title.
  2. Click Publish branch.
  3. Click Create Pull Request.
  4. Confirm that the base branch is main.

GitHub’s documentation also covers [creating a pull request from a fork](https://docs.github.com/en/pull-requests/how-tos/create-pull-requests/creating-a-pull-request-from-a-fork).

A suitable PR description would be:

## Summary

- Sets an empty, null, or whitespace-only DMP title to
  "Data Management Plan".
- Preserves existing non-empty document titles.
- Applies the fallback after the existing per-question rules so it does
  not affect section or question fields.
- Adds regression tests for the new behavior.

## Testing

- `py -3.14 -m pytest tests/test_annotation_rules.py`
- `py -3.14 -m pytest`

One design note: the current Rules.xlsx describes a 16-row truth table for copying text among title, section, and question fields. This new behavior is a document-level default rather than another truth-table combination. I would keep it separate in the code, but mention that choice in the PR and ask whether the maintainers also want the spreadsheet updated as human-readable documentation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions