About This Project
This project is built for fun and learning. The main goal is to learn how to crawl, store,
and analyze large datasets in a practical setup with MySQL and S3-compatible storage.
Parts of it are "vibe-coded". Years ago there was a handcrafted PHP version; it worked,
but it was much less structured and harder/slower to maintain than this rebuild. I know this contributes to the noise on the internet. That’s fine, and I’m not sorry but you can always remove your site.
Why It Exists
- Learn scalable crawling patterns.
- Practice data modeling and deduplication.
- Measure storage growth over time.
- Keep architecture stateless for container environments.
- Testing ZFS compression and dedup.
- All on slow homelab hardware, so it has to be very efficient.
Current Dataset
| URLs | 2209079 |
| Fetches | 2404360 |
| Screenshots | 1053473 |
| Site details | 2114131 |
Storage Footprint
| MySQL total size | 9.1 GiB (9781608448 bytes) |
| S3 total size | 1.7 GiB (1860067365 bytes) |
| S3 total files | 80354 |
| S3 snapshot updated | 2026-07-12 18:06:00 |
S3 Breakdown
| Prefix | Files | Bytes | Readable |
| screenshots/ |
80354 |
1860067365 |
1.7 GiB |
Top MySQL Tables By Size
| Table | Rows (est.) | Size |
| site_enrichments |
2048114 |
4.9 GiB |
| fetches |
2309453 |
940.8 MiB |
| urls |
2284522 |
783.4 MiB |
| links |
2435545 |
470.5 MiB |
| site_profiles |
1930943 |
465.9 MiB |
| domain_scores |
1982081 |
413.1 MiB |
| render_snapshots |
894579 |
286.9 MiB |
| change_events |
1762791 |
280.6 MiB |
| screenshots |
940896 |
260.8 MiB |
| fetch_observations |
971324 |
259.0 MiB |
`Rows (est.)` comes from MySQL metadata and is approximate.