Website Content Scraper
Turn any website into clean, AI-ready text. Paste a URL and CrawlHawk crawls the site and extracts every page's main content as Markdown — navigation, footers and cookie banners stripped away — downloadable as one file per page plus a single combined file. Feed it to ChatGPT, Claude, NotebookLM or any AI workflow. 500 pages free, no signup.
What the Content Scraper Does
It answers a need that barely existed two years ago and is now everywhere: "I need this website's content as text, so I can give it to an AI." CrawlHawk crawls the site within the scope you choose, isolates each page's main content — the article, the documentation, the product description, not the theme around it — and converts it to clean Markdown. The output preserves the structure that matters (headings, lists, links, emphasis) and drops the noise that wastes AI context windows (menus, sidebars, cookie banners, footers).
One File Per Page — Plus One Combined File
Multi-page crawls download as a ZIP containing a .md file for every page and a single combined file of the whole site. The combined file is the practical winner: tools like NotebookLM, ChatGPT Projects and Claude Projects limit how many files you can upload — one file containing the entire site fits where five hundred files won't. The per-page files serve the other workflows: version-controlled docs, selective imports, page-level processing.
Common Use Cases
AI assistants and knowledge bases: load your documentation, help center or company site into a custom GPT, Claude Project or NotebookLM so it answers from your actual content. RAG pipelines without the pipeline: clean Markdown is the standard ingestion format — get it as a download instead of building a scraper. Site migrations: content extracted from the old site, ready to import anywhere — CMS-independent by design. Content audits: every page's actual text in one place, for review, deduplication and rewriting projects. Translation preparation: the complete site as text files, ready to hand to translators or translation tooling. Research and archiving: a structured text capture of any public site.
Scope Control: the Whole Site or One Section
Extract the entire domain, a subdomain, a single path (just /docs/ or /blog/) or one page — see Crawl Scope Explained for choosing. Path scoping is the workhorse here: an AI assistant for your product usually needs the documentation section, not the careers page.
Free Tier and Credits
Content extraction is standard crawling: 1 credit per page, and the free tier covers 500 pages per crawl with no signup and no card — enough for most documentation sites and blogs, free. JavaScript-heavy or protected sites use additional credits per page, shown before the crawl starts. Larger sites use pay-once credit packs; credits never expire. See Pricing.
Frequently Asked Questions
How do I feed a whole website into ChatGPT or Claude?
Extract the site with the Content Scraper and upload the combined Markdown file to a ChatGPT Project, custom GPT, Claude Project or NotebookLM. One file, the whole site, inside the tool's upload limits.
How do I convert a website to Markdown?
Paste the URL above, keep Markdown as the output, and run. Each page's main content is converted to clean Markdown; the download is a ZIP with per-page files plus a combined file.
Does it strip menus, footers and cookie banners?
Yes — extraction targets each page's main content and drops repeated theme elements, so the output is the text that matters, not the chrome around it.
Is it free?
Up to 500 pages per crawl, yes — no signup, no card. That covers most blogs and documentation sites. Bigger sites use pay-once credits that never expire.
Can it extract text from JavaScript-heavy sites?
Yes — with rendering enabled, dynamically loaded content is captured, at an additional credit cost per page shown before the crawl starts.
Can I scrape someone else's website content?
CrawlHawk extracts publicly accessible content only. The text remains its author's intellectual property — you are responsible for how you use it and for compliance with the site's terms, copyright law and the Acceptable Use Policy. Extracting your own site, or content you have rights to use, is the intended use.
Start extracting — 500 pages free, no credit card required →
Related tools: WordPress Scraper · Custom Link Crawler · AI Product Scraper · XML Sitemap Generator