Knows.nl

Technical website intelligence, screenshots, and crawl history.

About This Project

This project is built for fun and learning. The main goal is to learn how to crawl, store, and analyze large datasets in a practical setup with MySQL and S3-compatible storage. Parts of it are "vibe-coded". Years ago there was a handcrafted PHP version; it worked, but it was much less structured and harder/slower to maintain than this rebuild. I know this contributes to the noise on the internet. That’s fine, and I’m not sorry but you can always remove your site.

Why It Exists

  • Learn scalable crawling patterns.
  • Practice data modeling and deduplication.
  • Measure storage growth over time.
  • Keep architecture stateless for container environments.
  • Testing ZFS compression and dedup.
  • All on slow homelab hardware, so it has to be very efficient.

Current Dataset

URLs2209079
Fetches2404360
Screenshots1053473
Site details2114131

Storage Footprint

MySQL total size9.1 GiB (9781608448 bytes)
S3 total size1.7 GiB (1860067365 bytes)
S3 total files80354
S3 snapshot updated2026-07-12 18:06:00

S3 Breakdown

PrefixFilesBytesReadable
screenshots/ 80354 1860067365 1.7 GiB

Top MySQL Tables By Size

TableRows (est.)Size
site_enrichments 2048114 4.9 GiB
fetches 2309453 940.8 MiB
urls 2284522 783.4 MiB
links 2435545 470.5 MiB
site_profiles 1930943 465.9 MiB
domain_scores 1982081 413.1 MiB
render_snapshots 894579 286.9 MiB
change_events 1762791 280.6 MiB
screenshots 940896 260.8 MiB
fetch_observations 971324 259.0 MiB

`Rows (est.)` comes from MySQL metadata and is approximate.