built a simple scraper for training data we currently scrape:

  • open github projects open for LLM training
  • internal code and docs
  • wikipedia's LLM training page

these are all going to be used to train the first iteration of edith agents


Restoration note: Restored with original attachments from the Stardance backup. Publication date reconstructed from project metadata, original log order, and DeltaTime activity; the export did not retain devlog timestamps. Hours are reconstructed allocations of surviving DeltaTime history.