A scraper that collects and aggregates conference CFP and posting metadata from public sources.
This repository contains a conference aggregator that crawls public conference listing sites and produces machine-readable conference datasets at the repository root (conference-postings.json without CFP body text, and conference-postings-full.json with conferenceText).
The collector implements scrapers for these sources:
wiki.cfp(WikiCFP)wikidata.orgconfident-conference.org(ConfIDent)
Extracted fields include conference name, acronym, year, dates, location, website URL, CFP text snippets, and categories metadata.
- Crawls listing and detail pages for the source(s) specified.
- Parses conference metadata and normalizes dates and acronyms.
- Merges new postings into JSON databases at
conference-postings.json(slim) andconference-postings-full.json(includesconferenceText). - Provides single-site and full-collection CLI commands.
Scripts run main.ts with --site set to all or a source id (wiki.cfp, wikidata.org, confident-conference.org).
Each collector crawls listings, parses detail pages, and returns rows. main.ts deduplicates by conference name, then writes conference-postings-full.json (includes conferenceText) and conference-postings.json (same records, conferenceText omitted).
- One source: load the full file, delete that source’s old rows, add the new crawl, dedupe, save both files. Other sources are untouched.
all: run every collector, deduplicate, save both files (full refresh).
Normalization happens in two layers: while scraping, and when merging duplicates.
During scraping
- Dates are parsed into ISO
YYYY-MM-DDwhere possible (utils.ts). - Each scraper sets
_sourcesto the site id (for examplewiki.cfp). conferenceAcronymis passed throughnormalizeConferenceAcronym()inutils.ts. Values that look like titles, URLs, or full sentences are dropped (null), with a console log explaining the rejection.
Merge and deduplication
After collection, deduplicate.ts groups postings by normalized conference name (trimmed, collapsed whitespace, case-insensitive). Each group becomes one record.
For most fields, the merged value is taken from the most trusted source that has a non-empty value. Trust order is DEDUP_CONFIG.sourceOrder in collection-config.ts (first entry is most trusted). Tie-breakers: more populated merge fields, then lowest stable id.
Special cases:
- Start and end dates: postings with both dates set rank above partial dates; then the usual source order applies so start and end usually come from the same row.
- Categories: union of unique category strings across the group (order follows source priority when iterating).
_sources: union of all sources in the group.collectionDate: newest date in the group.id: taken from the winning row forconferenceName.
scripts/conference-collecting/main.ts: CLI entrypoint (--site)scripts/conference-collecting/collection-config.ts: per-source crawl limits and dedupsourceOrderscripts/conference-collecting/wikicfp.ts: WikiCFP scraperscripts/conference-collecting/wikidata.ts: Wikidata SPARQL collectorscripts/conference-collecting/confident-conference.ts: ConfIDent via MediaWikiaction=ask+action=query(SMW)scripts/conference-collecting/schema.ts: record/DB shapesscripts/conference-collecting/storage.ts: load/save conference JSON exports
conference-postings.json(repo root): consolidated output withoutconferenceText(for lightweight consumers).conference-postings-full.json: same records withconferenceTextpreserved.storage/(repo root): Crawlee internal crawl state (KV stores + request queues). Deleting it resets crawl progress/state.
You will need the following installed on your system:
-
Clone the repository
git clone https://github.com/fairdataihub/posters-science.git cd posters-science -
Trust and install the required tool versions
mise trust mise install
-
Install dependencies
pnpm install
-
Add your environment variables
cp .env.example .env
-
Start the development server
pnpm dev
-
Open the application at http://localhost:3000 or appropriate port if you have it configured differently.
Run the collectors (examples):
- Collect all configured sources (local runs use test limits; GitHub Actions use full automatically):
pnpm run collect:all- Collect only one site (for example wiki.cfp):
pnpm run collect:wiki.cfpEdit LIMITS.test / LIMITS.full in scripts/conference-collecting/collection-config.ts. Mode is full when GITHUB_ACTIONS=true, otherwise test.
The collectors write both JSON files at the repository root. Load conference-postings.json for dropdowns and lists; use conference-postings-full.json when you need CFP text.
Collection limits and modes are in scripts/conference-collecting/collection-config.ts.
After changes, run pnpm run collect:* to update conference-postings.json.
Contributions welcome — open issues or pull requests for new collector sources, bug fixes, or improvements to normalization/merging logic.