Crawl a site
1
Create the spider
books_spider.py
2
Prepare Scrapy
3
Crawl the site
The destination name must be unused. Restore installs the deny-by-default allowlist, connection cap, restricted guest profile, and lifetime bound before boot. Copy the spider into the root-owned The crawler and the microVM policy both constrain navigation to the target host. Change the spider, start URL, and network rule together when adapting the example, and respect the site’s terms and robots policy.
/work directory and make it read-only before running it as the unprivileged user; no host directory is exposed.4
Copy out the result
5
Clean up