How to Scrape Reddit Without an API Key in 2026
The old reddit.com .json endpoints now return 403 and commercial API access is enterprise-only. Every method that still works in 2026, with code you can use today.
The actor referenced in this article. Pay only for results delivered.
Reddit’s API shutdown in June 2023 ended the era of free, unauthenticated JSON endpoints, and in 2026 the door closed completely. The old trick of appending .json to any Reddit URL now returns HTTP 403 even for public pages at casual volume (we re-verified while updating this post in July 2026). For any data collection, from a single subreddit feed to thousands of posts with full comment trees, you need one of the approaches below.
TL;DR: Reddit’s Android app uses a public OAuth client ID (
ohXpoqrZYub1kg) that any developer can use with theinstalled_clientgrant: no account registration, same 100 req/min rate limit as registered apps, full API access. For volume above that, a managed pay-per-result scraper is the practical choice.
Here is every method that still works in 2026, from lightest to most robust.
The Old .json Trick Is Dead
Before 2023, https://www.reddit.com/r/python.json returned clean paginated data with no auth. Reddit first rate-limited these endpoints aggressively (429s within minutes), and as of 2026 they return 403 Forbidden outright, even a single request from a normal browser user-agent is refused. Every script and tutorial built on the .json suffix is now broken. If that’s what brought you here, skip straight to the OAuth method or a managed scraper below.
Method 1: OAuth with a Reddit Developer App
The official path. You register an app at reddit.com/prefs/apps, get a client_id and client_secret, and exchange them for a bearer token via the password flow or authorization code flow.
import requests
CLIENT_ID = 'your_client_id'
CLIENT_SECRET = 'your_client_secret'
auth = requests.auth.HTTPBasicAuth(CLIENT_ID, CLIENT_SECRET)
data = {'grant_type': 'password', 'username': 'u', 'password': 'p'}
headers = {'User-Agent': 'MyApp/0.1'}
token = requests.post(
'https://www.reddit.com/api/v1/access_token',
auth=auth, data=data, headers=headers
).json()['access_token']
posts = requests.get(
'https://oauth.reddit.com/r/python/hot',
headers={**headers, 'Authorization': f'Bearer {token}'}
).json()
Limits: 100 requests per minute per OAuth token. Enough for moderate volumes. The app registration page uses reCAPTCHA, which is annoying to get through.
Method 2: The Public Reddit Android Client ID
This is the technique used in production by most Reddit scrapers. Reddit’s official Android app has a public client_id that anyone can use with the installed_client grant type, no developer account required.
const CLIENT_ID = 'ohXpoqrZYub1kg'; // Reddit Android public client_id
const USER_AGENT = 'android:com.reddit.frontpage:v2024.45.0 (by /u/yourusername)';
async function getToken() {
const deviceId = crypto.randomUUID();
const res = await fetch('https://www.reddit.com/api/v1/access_token', {
method: 'POST',
headers: {
'Authorization': 'Basic ' + btoa(`${CLIENT_ID}:`),
'Content-Type': 'application/x-www-form-urlencoded',
'User-Agent': USER_AGENT,
},
body: `grant_type=https://oauth.reddit.com/grants/installed_client&device_id=${deviceId}`,
});
return (await res.json()).access_token;
}
async function getPosts(subreddit, token) {
const res = await fetch(`https://oauth.reddit.com/r/${subreddit}/hot?limit=100`, {
headers: {
'Authorization': `Bearer ${token}`,
'User-Agent': USER_AGENT,
}
});
return (await res.json()).data.children.map(c => c.data);
}
This approach gives you the same rate limits as an authenticated user, 100 req/min, without account registration. All calls go to oauth.reddit.com, not www.reddit.com.
Important: Always set a descriptive User-Agent. Reddit blocks generic or missing user agents.
Method 3: Pushshift (Partially Restored)
Pushshift.io was the go-to for historical Reddit data until it was shut down in 2023. Its practical successor for bulk historical analysis is Arctic Shift, which serves downloadable archives; coverage is strongest pre-2023 and it has no global keyword search. For historical backfills it remains the right free tool; for live and recent data it doesn’t compete with the methods above.
Method 4: Use a Managed Scraper
If you need volume (millions of posts), historical backfills, or comment trees without managing session rotation, a managed solution is the practical choice. Our Reddit Scraper on Apify handles auth rotation, proxies, and rate limit backoff automatically. You pay per post scraped, and failed or empty results are never charged.
from apify_client import ApifyClient
client = ApifyClient('YOUR_API_TOKEN')
run = client.actor('themineworks/reddit-scraper').call(run_input={
'mode': 'subreddit',
'subreddits': ['machinelearning', 'LocalLLaMA'],
'sortBy': 'hot',
'maxPosts': 500,
'includeComments': True,
'maxCommentsPerPost': 50,
})
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item['title'], item['score'])
Handling Rate Limits and Bans
Regardless of method, follow these rules:
- One request every 0.6 seconds minimum: Reddit’s ToS requires this
- Rotate tokens if running parallel workers: don’t share a single token across threads
- Respect
Retry-Afterheaders: Reddit will 429 you on bursts and tell you how long to wait - Use descriptive user agents: include your username as Reddit recommends
Getting Full Comment Trees
Comments hold most of the useful signal and are the hardest part to collect completely. A post’s comments come back as a recursive tree where each comment has its own replies listing. Once a thread gets large, Reddit swaps deep or overflow branches for a more placeholder that holds only the IDs of the comments it did not send:
{
"kind": "more",
"data": {
"id": "xyz",
"children": ["def456", "ghi789"]
}
}
Using the token and USER_AGENT from Method 2, fetch the post with its first batch of comments. depth=10 asks for up to ten levels of nesting, well past the shallow default, though Reddit still truncates big branches:
async function getPostWithComments(postId, token) {
const res = await fetch(
`https://oauth.reddit.com/comments/${postId}?limit=200&sort=top&depth=10`,
{ headers: { 'Authorization': `Bearer ${token}`, 'User-Agent': USER_AGENT } }
);
const [postListing, commentListing] = await res.json();
return {
post: postListing.data.children[0].data,
comments: commentListing.data.children,
};
}
Walk that tree into a flat list and collect every more ID on the way:
function flattenTree(items, out = { comments: [], moreIds: [] }, depth = 0) {
for (const item of items) {
if (item.kind === 'more') {
out.moreIds.push(...item.data.children);
continue;
}
const c = item.data;
out.comments.push({
id: c.id,
parent_id: c.parent_id,
body: c.body,
author: c.author,
score: c.score,
created_utc: c.created_utc,
depth,
is_deleted: ['[deleted]', '[removed]'].includes(c.body),
});
if (c.replies?.data?.children) {
flattenTree(c.replies.data.children, out, depth + 1);
}
}
return out;
}
Then expand the placeholders through /api/morechildren, which accepts up to 100 IDs per call. Its results can contain further more objects, so treat the IDs as a queue and keep going until it is empty:
async function expandMore(postId, moreIds, token) {
const queue = [...moreIds];
const expanded = [];
while (queue.length) {
const batch = queue.splice(0, 100);
const res = await fetch('https://oauth.reddit.com/api/morechildren', {
method: 'POST',
headers: {
'Authorization': `Bearer ${token}`,
'User-Agent': USER_AGENT,
'Content-Type': 'application/x-www-form-urlencoded',
},
body: new URLSearchParams({
link_id: `t3_${postId}`,
children: batch.join(','),
sort: 'top',
api_type: 'json',
}),
});
const { json } = await res.json();
for (const thing of json.data.things) {
if (thing.kind === 't1') expanded.push(thing.data);
else if (thing.kind === 'more') queue.push(...thing.data.children);
}
await new Promise(r => setTimeout(r, 600)); // stay under 100 req/min
}
return expanded;
}
const { post, comments } = await getPostWithComments('abc123', token);
const tree = flattenTree(comments);
const allComments = [...tree.comments, ...(await expandMore('abc123', tree.moreIds, token))];
Expanded comments arrive as a flat list, so rebuild the thread from each comment’s parent_id (t1_ points to a comment, t3_ to the post) rather than from the order they come back in. Every morechildren call counts against the rate limit, and a thread with 5,000 comments can need 50 or more of them.
Deleted and removed comments come back with [deleted] or [removed] as the body. [deleted] means the author removed it and [removed] means a moderator did. Either way the text is gone from the API, so filter both out before any NLP work or your model will learn the [deleted] token.
The managed scraper in Method 4 returns the tree already nested. Set maxDepth (1 for top-level only, up to 10) and maxCommentsPerPost to cap how much of each thread you pull, and replies arrive inside each comment’s replies array.
What Data You Can Get
The Reddit OAuth API provides:
- Posts: title, body, score, upvote ratio, awards, flair, author, created timestamp
- Comments: full nested tree, author, score, gilded status
- Subreddit metadata: subscriber count, active users, description
- User profiles: post history, comment karma, account age
What it does not provide: deleted content, shadow-banned users, private subreddits.
Sending the Data to Claude
Once posts and comments are structured JSON, a model can do the reading. A common setup is a brand monitor: pull recent mentions, have Claude classify each post, alert on the few that need a person, and roll the rest into a weekly brief.
Collect mentions with the scraper’s search mode. If keyword monitoring is all you need, Reddit Search Scraper runs the same search with a single input, at $1 per 1,000 posts:
import os
from apify_client import ApifyClient
apify = ApifyClient(os.environ['APIFY_TOKEN'])
def fetch_mentions(query: str, max_posts: int = 100) -> list[dict]:
run = apify.actor('themineworks/reddit-scraper').call(run_input={
'mode': 'search',
'searchQuery': query,
'sortBy': 'new',
'maxPosts': max_posts,
'includeComments': True,
'maxCommentsPerPost': 20,
})
return list(apify.dataset(run['defaultDatasetId']).iterate_items())
Classify each post in one call that returns every field you need. Keyword classifiers struggle with sarcasm and understatement, which developer subreddits are full of, and they miss a feature request tucked inside a compliment. A model reading the whole post and its top comments handles both.
import json
from concurrent.futures import ThreadPoolExecutor
import anthropic
claude = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
def classify_post(post: dict) -> dict:
content = f"Title: {post['title']}\n\nBody: {(post.get('selftext') or '')[:1000]}"
top_comments = [(c.get('body') or '')[:200] for c in post.get('comments', [])[:5]]
if top_comments:
content += '\n\nTop comments:\n' + '\n'.join(top_comments)
response = claude.messages.create(
model='claude-haiku-4-5',
max_tokens=500,
messages=[{
'role': 'user',
'content': f"""Classify this Reddit post. Return only a JSON object with:
- sentiment: "positive" | "negative" | "neutral" | "mixed"
- intent: "complaint" | "praise" | "question" | "discussion" | "feature_request" | "comparison"
- urgency: 1-5 (5 = needs attention today)
- key_entities: products, companies or people mentioned
- summary: one sentence
- action_required: true or false
Post:
{content}""",
}],
)
text = response.content[0].text.strip()
if text.startswith('`'): # drop a markdown code fence if the model adds one
text = text.strip('`').removeprefix('json')
try:
result = json.loads(text)
except json.JSONDecodeError:
return {'id': post['id'], 'error': 'parse_failed'}
return {**result, 'id': post['id'], 'url': post['url'],
'score': post['score'], 'title': post['title']}
def classify_all(posts: list[dict]) -> list[dict]:
with ThreadPoolExecutor(max_workers=5) as pool:
results = list(pool.map(classify_post, posts))
return sorted(results, key=lambda r: r.get('urgency', 0), reverse=True)
Five workers is plenty for a few hundred posts a week, and the Anthropic SDK retries rate-limit errors twice by default. Then decide what cannot wait for the weekly report:
def alert_reasons(p: dict) -> list[str]:
reasons = []
if p.get('urgency', 0) >= 4:
reasons.append(f"urgency {p['urgency']}/5")
if p.get('score', 0) > 500:
reasons.append(f"{p['score']} upvotes")
if p.get('action_required'):
reasons.append('action required')
if p.get('intent') == 'complaint' and p.get('score', 0) > 50:
reasons.append('complaint gaining traction')
return reasons
posts = classify_all(fetch_mentions('"your product" OR "competitor name"'))
alerts = [(p, reasons) for p in posts if (reasons := alert_reasons(p))]
Send the alerts to Slack or email the same day. For the weekly brief, pass the sentiment and intent counts plus the one-line summaries of the top 50 posts back to Claude and ask for a short written report on themes, recurring complaints, competitor mentions, feature requests and suggested actions. Run the whole script on a schedule (cron or Apify Schedules) and nobody has to read Reddit by hand.
The same pipeline works for other jobs. Product teams get the bug reports their support queue never sees. Investors can run six months of history on a company before a deal. Content teams get a list of real questions to answer, and anyone tracking competitors sees a negative thread while it is still gaining traction.
Summary
| Method | Registration Required | Rate Limit | Best For |
|---|---|---|---|
.json trick | No | Very low | Testing only |
| Developer OAuth app | Yes (reCAPTCHA) | 100 req/min | Small projects |
| Android client ID | No | 100 req/min | Production scraping |
| Managed scraper | No | Effectively unlimited | High volume |
For most production use cases, the Android client ID approach is the sweet spot: no account required, full API access, same rate limits as registered apps.
Frequently Asked Questions
Can you still scrape Reddit without an API key in 2026?
Yes. Reddit’s official Android app uses a public client ID (ohXpoqrZYub1kg) that works with the installed_client grant type, no developer account registration required. This gives you the same 100 requests per minute as a registered app and full access to the OAuth API endpoints.
What happened to the Reddit .json trick after the 2023 API change?
It no longer works at all. After the 2023 change Reddit heavily rate-limited unauthenticated .json requests; in 2026 those endpoints return 403 Forbidden even for single public-page requests. Use the OAuth path or a managed scraper instead.
How much does the Reddit API cost at scale?
Reddit charges $0.24 per 1,000 API calls via the official developer OAuth path. At moderate volume (10,000 posts/month) that is $2.40. At scale, 1 million posts per month with comment trees averaging 3 API calls each, costs reach approximately $720/month before infrastructure.
What is Reddit’s API rate limit?
The Reddit OAuth API enforces 100 requests per minute per token, regardless of whether you use a registered developer app or the Android public client ID. For parallel scraping, you must rotate across multiple tokens and never share a single token across threads.
Is scraping Reddit legal?
The 2022 hiQ Labs v. LinkedIn ruling established that scraping publicly visible data does not constitute unauthorized access under the CFAA. Reddit’s terms of service prohibit automated access, but ToS violations are civil contract matters, not criminal offenses. For commercial operations at scale, review the applicable laws in your jurisdiction.
Explore the scraper referenced in this article: inputs, outputs, and pricing, then run it on Apify.
Reddit Data for Market Research After the API Changes: What Still Works in 2026
Reddit locked down its API: enterprise-only commercial access, no self-serve pricing, and the old .json endpoints gone. Here are the working options for market research teams, with honest costs.
Reddit Official API vs Reddit Scraper in 2025: Costs, Limits, and What You Actually Get
Reddit changed its API pricing in 2023 to $0.24 per 1,000 calls. Here is what that means for data collection workloads, and how scraping compares on cost and data coverage.
Reddit Sentiment Analysis Pipeline: From Raw Posts to Actionable Insights
How to build a production sentiment analysis pipeline on Reddit data, from scraping and preprocessing to classification.
Reddit Data for LLM Fine-Tuning: Quality, Licensing, and What Actually Works
How to use Reddit data for LLM training and fine-tuning: quality patterns, filtering strategies, preference pairs and evaluation sets.