WordPress Scraper
Extract content, links, images and files from any WordPress site. Paste the site URL and CrawlHawk crawls it into structured, downloadable data — link inventories as CSV or Excel, page content as clean Markdown ready for migrations and AI workflows. For WooCommerce product data, see the dedicated WooCommerce Scraper. No plugin, no admin access, no coding.
What the WordPress Scraper Does
WordPress runs more of the web than any other CMS, so "scrape a WordPress site" means different things to different people — and CrawlHawk covers the spectrum: content extraction (posts and pages as clean Markdown, navigation and widgets stripped), link auditing (every internal, external, broken and orphan link with source and anchor text), media inventory (every image and file URL, including the uploads library as it's actually referenced), and XML sitemap generation. It works from the public site — no plugin installed, no wp-admin access, no REST API keys.
Content as Clean Markdown
The export format that changes WordPress workflows: each crawled page's main content as Markdown, with theme chrome, sidebars, cookie banners and footers stripped away. That's the right input for site migrations (WordPress to a static generator, a headless CMS, or another WordPress), for feeding a site's knowledge into an AI assistant or RAG pipeline, and for content audits where you want the words, not the theme. Multi-page crawls download as a ZIP with one .md file per page plus a combined file.
Scope It Like a WordPress Site
WordPress's predictable structure makes scoping natural: crawl the whole site, or restrict to a path — /blog/ for the posts, a category prefix, or a single page. Tag and category archives can multiply URLs on large blogs; a path-scoped crawl keeps the run focused on the content itself. See Crawl Scope Explained for the scope-to-task mapping.
Common Use Cases
Site migrations: full content export before moving off (or between) WordPress installs — the Markdown output imports anywhere. SEO audits: the complete link graph of a WordPress site, broken links included, without installing yet another plugin on a production site. AI knowledge bases: turn a documentation or blog site into LLM-ready Markdown. Research and archiving: structured capture of a public WordPress site's content and media inventory.
Looking for Products? That's WooCommerce
If the WordPress site is a shop, it's running WooCommerce — and product extraction (titles, prices, variations, SKUs) has its own tool with its own page: the WooCommerce Scraper. This page covers the content, link and media side of WordPress.
Credits and the Free Tier
Standard WordPress pages crawl at 1 credit per URL — and the free tier's 500 URLs per crawl cover exactly this, so most blogs and small sites can be crawled free, no card required. JavaScript-heavy or protected sites use additional credits per URL, shown before the crawl starts. Credit packs are pay-once and never expire. See Pricing.
Frequently Asked Questions
How do I scrape a WordPress site?
Paste the site's URL into the crawler above, choose what to collect (links, images, files) and the output format (CSV, Excel, JSON, Markdown), and run. Results download when the crawl completes.
Do I need a plugin or wp-admin access?
No — CrawlHawk crawls the public site like a visitor. Nothing is installed, which also means it works on WordPress sites you don't own.
Can I export a WordPress site's content as Markdown?
Yes — each page's main content is extracted as clean Markdown with theme elements stripped, downloadable per page and as a combined file. It's the practical format for migrations and AI pipelines.
Is this free?
Standard crawling is free up to 500 URLs per crawl, with no signup and no card — enough for most blogs. Larger sites and advanced modes use pay-once credits that never expire.
Can it scrape membership or password-protected content?
No — CrawlHawk crawls publicly accessible pages only and does not bypass logins or access controls, per the Acceptable Use Policy.
Does it handle WordPress sites behind Cloudflare?
Yes — protected pages are crawled with browser rendering at an additional credit cost per URL, shown before the crawl starts.
Start crawling — 500 URLs free, no credit card required →
Related tools: WooCommerce Scraper · Custom Link Crawler · XML Sitemap Generator · Broken Link Checker