Substack Scraper
Newsletter archives, full text, reactions, and paywall status
What it does
Scrape any Substack newsletter's full archive: headline, standfirst, byline, publish date, full post text, emoji reaction breakdown, comment count, and a paywall flag. Reads each publication's own archive and per-post endpoints, works on subdomains and custom domains alike. No login, no API key, pay per post.
Built in
- ✓ Headline, standfirst, byline, and publish date per post
- ✓ Full post text pulled from the per-post endpoint, not just the archive summary
- ✓ Emoji reaction breakdown, aggregated reaction score, and comment count
- ✓ Paywall flag and a truncation flag for text shorter than the reported word count
- ✓ Works on subdomains and custom domains, no login required
Common questions
Does this need a Substack login? +
No. It reads each publication's public archive and per-post endpoints over plain HTTP, no login, no browser automation.
Does it work on newsletters with a custom domain, not just *.substack.com? +
Yes. Both subdomains and custom domains are supported.
Can I get paywalled posts? +
You get the metadata and a paywall flag for every post. Full text comes back for whatever the publication makes publicly readable; paywalled body text stays paywalled.
What are the main use cases? +
Newsletter competitive research by ranking posts on engagement, topic and framing research across an archive, tracking publishing cadence and free-vs-paid mix, and building a RAG index from newsletter prose.
How does pricing work? +
Pay-per-event at $0.0015 per post delivered.
Try it on Apify.
Pay only for results delivered. No credit card required to sign up.