Enrich the monster catalog with per-creature harvesting tables #132

Merged
nasandre merged 1 commit from feature/monster-harvesting-data into main 2026-10-03 10:00:30 +00:00
Owner

Scrapes thievesguild.cc/harvest into a committed dataset and merges it into every served monster, so statblocks (GM screen, encounters, story tracker) show per-creature harvesting rules.

What ships

  • Dataset (backend/src/data/harvesting/thievesguild-harvest.json): 1,021 creatures scraped, 941 with harvesting tables, 2,542 rows - skill by creature type (Arcana/Nature/Medicine/Investigation), per-item DC, value, weight, expiration, alchemy-ingredient and craft-into references, meat yields.
  • Scraper (backend/scripts/scrape_thievesguild_harvest.py): resumable, fetches through FlareSolverr (site is behind a Cloudflare managed challenge). FlareSolverr corrupts concurrent requests sharing a session - caught by the built-in page verification after 119 files were hit - so workers now use strictly per-session queues and every cached page is verified against the index before parsing.
  • Serve-time merge (harvestCatalogService): normalized-name matching + comma-variant fallback ("Aboleth, Sage" -> Aboleth) + a 2024-rename alias table ("Azer Sentinel" -> "Azer"). Applied in mcpDataService after the cache layer, because the creature catalog snapshot is wholesale replaced on every Open5E refresh - inline enrichment would be wiped.

Coverage

  • 88.6% of the served catalog is enriched (302/341 live). The remainder are NPC statblocks (Archmage, Bandit, Guard...) and 2024-only creatures the 2014-era source does not cover; every harvestable creature resolves.
  • Preserved through the homebrew-monster reshaping in /dnd/creatures (which would otherwise have dropped the field); an empty MCP fetch_monster result no longer shadows the catalog fallback for srd-2024_* keys.

Provenance

Monster stats are SRD 5.1/5.2 (CC-BY); harvest tables derive from Hamund's Harvesting Handbook. Private, non-commercial campaign tool, attribution shown in the statblock tooltip.

Validation

Backend 34 suites / 740 tests, frontend 32 files / 361 tests, tsc clean x3. New tests: parser fixtures against a committed HTML page (DC/value/weight/expiry/alchemy), service units (normalization, variants, aliases, idempotency, missing dataset), dataset integrity (every alias target resolves). Live-verified: /api/dnd/creatures serves 302 enriched; Aboleth carries its 4-row Arcana table.

Known notes

  • FlareSolverr (2 containers on :8191/:8192) is only needed to re-scrape; the app itself needs nothing extra.
  • Old encounter snapshots embed pre-enrichment statblocks until refreshSnapshots re-resolves them.
  • Phase 2 of the crafting/harvesting plan (docs/crafting-harvesting-shop-plan.md, updated) consumes these tables for the encounter-end harvest flow.
Scrapes thievesguild.cc/harvest into a committed dataset and merges it into every served monster, so statblocks (GM screen, encounters, story tracker) show per-creature harvesting rules. ## What ships - **Dataset** (`backend/src/data/harvesting/thievesguild-harvest.json`): 1,021 creatures scraped, 941 with harvesting tables, 2,542 rows - skill by creature type (Arcana/Nature/Medicine/Investigation), per-item DC, value, weight, expiration, alchemy-ingredient and craft-into references, meat yields. - **Scraper** (`backend/scripts/scrape_thievesguild_harvest.py`): resumable, fetches through FlareSolverr (site is behind a Cloudflare managed challenge). FlareSolverr corrupts concurrent requests sharing a session - caught by the built-in page verification after 119 files were hit - so workers now use strictly per-session queues and every cached page is verified against the index before parsing. - **Serve-time merge** (`harvestCatalogService`): normalized-name matching + comma-variant fallback ("Aboleth, Sage" -> Aboleth) + a 2024-rename alias table ("Azer Sentinel" -> "Azer"). Applied in `mcpDataService` after the cache layer, because the creature catalog snapshot is wholesale replaced on every Open5E refresh - inline enrichment would be wiped. ## Coverage - **88.6% of the served catalog is enriched** (302/341 live). The remainder are NPC statblocks (Archmage, Bandit, Guard...) and 2024-only creatures the 2014-era source does not cover; every harvestable creature resolves. - Preserved through the homebrew-monster reshaping in `/dnd/creatures` (which would otherwise have dropped the field); an empty MCP `fetch_monster` result no longer shadows the catalog fallback for `srd-2024_*` keys. ## Provenance Monster stats are SRD 5.1/5.2 (CC-BY); harvest tables derive from Hamund's Harvesting Handbook. Private, non-commercial campaign tool, attribution shown in the statblock tooltip. ## Validation Backend 34 suites / **740 tests**, frontend 32 files / **361 tests**, `tsc` clean x3. New tests: parser fixtures against a committed HTML page (DC/value/weight/expiry/alchemy), service units (normalization, variants, aliases, idempotency, missing dataset), dataset integrity (every alias target resolves). Live-verified: `/api/dnd/creatures` serves 302 enriched; Aboleth carries its 4-row Arcana table. ## Known notes - FlareSolverr (2 containers on :8191/:8192) is only needed to re-scrape; the app itself needs nothing extra. - Old encounter snapshots embed pre-enrichment statblocks until `refreshSnapshots` re-resolves them. - Phase 2 of the crafting/harvesting plan (docs/crafting-harvesting-shop-plan.md, updated) consumes these tables for the encounter-end harvest flow.
Scrapes thievesguild.cc/harvest (1,021 creatures, 941 with harvesting
tables, 2,542 rows) into a committed dataset and merges it into every
served monster, so the GM screen and encounter views show per-creature
harvest rules: skill by creature type (Arcana/Nature/Medicine/
Investigation), per-item DC, value, weight, expiration and alchemy/craft
references, plus meat yields.

- backend/scripts/scrape_thievesguild_harvest.py: resumable scraper that
  fetches through FlareSolverr (the site sits behind a Cloudflare managed
  challenge), parses the harvesting/meat blocks and writes the dataset.
  FlareSolverr corrupts concurrent requests sharing a session (page A
  delivered for URL B - 119 files were hit before verification caught
  it), so workers get strictly per-session queues and every cached page
  is verified against the index before parsing.
- harvestCatalogService merges by normalized name with comma-variant
  fallback ("Aboleth, Sage" -> Aboleth) and a 2024-rename alias table
  ("Azer Sentinel" -> "Azer"); 88.6% of the served catalog is enriched -
  the remainder are NPC statblocks and 2024-only creatures the source
  does not cover. Enrichment is applied after the cache layer and the
  dataset ships in the image, because the creature catalog snapshot is
  wholesale replaced on every Open5E refresh.
- The legacy /dnd/monsters path is untouched; /dnd/creatures +
  /dnd/creatures/:key carry `harvesting`, preserved through the
  homebrew-monster reshaping (which previously would have dropped it).
  An empty MCP fetch_monster result no longer shadows the catalog
  fallback for srd-2024_* keys.
- StatblockDisplay renders the table with DC/value/weight/expiry and the
  provenance tooltip: monster stats SRD 5.1/5.2 (CC-BY), harvest tables
  derived from Hamund's Harvesting Handbook - private, non-commercial
  use.

Tests: parser fixtures (DC/value/weight/expiry/alchemy parsing against a
committed HTML page), service units (normalization, comma variants,
aliases, idempotency, missing dataset), and dataset integrity (every
alias target resolves).
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
nasandre/dnd-character-generator!132
No description provided.