built a simple scraper for training data we currently scrape:
- open github projects open for LLM training
- internal code and docs
- wikipedia's LLM training page
these are all going to be used to train the first iteration of edith agents
Restoration note: Restored with original attachments from the Stardance backup. Publication date reconstructed from project metadata, original log order, and DeltaTime activity; the export did not retain devlog timestamps. Hours are reconstructed allocations of surviving DeltaTime history.

Sign in to comment and vote.