The Mine Works
How to Scrape Reddit Without an API Key in 2026
← All posts
tutorial May 12, 2025 · 11 min read Updated September 17, 2026

How to Scrape Reddit Without an API Key in 2026

The old reddit.com .json endpoints now return 403 and commercial API access is enterprise-only. Every method that still works in 2026, with code you can use today.

Try the scraper

The actor referenced in this article. Pay only for results delivered.

View the scraper →

Reddit’s API shutdown in June 2023 ended the era of free, unauthenticated JSON endpoints, and in 2026 the door closed completely. The old trick of appending .json to any Reddit URL now returns HTTP 403 even for public pages at casual volume (we re-verified while updating this post in July 2026). For any data collection, from a single subreddit feed to thousands of posts with full comment trees, you need one of the approaches below.

TL;DR: Reddit’s Android app uses a public OAuth client ID (ohXpoqrZYub1kg) that any developer can use with the installed_client grant: no account registration, same 100 req/min rate limit as registered apps, full API access. For volume above that, a managed pay-per-result scraper is the practical choice.

Here is every method that still works in 2026, from lightest to most robust.

The Old .json Trick Is Dead

Before 2023, https://www.reddit.com/r/python.json returned clean paginated data with no auth. Reddit first rate-limited these endpoints aggressively (429s within minutes), and as of 2026 they return 403 Forbidden outright, even a single request from a normal browser user-agent is refused. Every script and tutorial built on the .json suffix is now broken. If that’s what brought you here, skip straight to the OAuth method or a managed scraper below.

Method 1: OAuth with a Reddit Developer App

The official path. You register an app at reddit.com/prefs/apps, get a client_id and client_secret, and exchange them for a bearer token via the password flow or authorization code flow.

import requests

CLIENT_ID = 'your_client_id'
CLIENT_SECRET = 'your_client_secret'

auth = requests.auth.HTTPBasicAuth(CLIENT_ID, CLIENT_SECRET)
data = {'grant_type': 'password', 'username': 'u', 'password': 'p'}
headers = {'User-Agent': 'MyApp/0.1'}

token = requests.post(
    'https://www.reddit.com/api/v1/access_token',
    auth=auth, data=data, headers=headers
).json()['access_token']

posts = requests.get(
    'https://oauth.reddit.com/r/python/hot',
    headers={**headers, 'Authorization': f'Bearer {token}'}
).json()

Limits: 100 requests per minute per OAuth token. Enough for moderate volumes. The app registration page uses reCAPTCHA, which is annoying to get through.

Method 2: The Public Reddit Android Client ID

This is the technique used in production by most Reddit scrapers. Reddit’s official Android app has a public client_id that anyone can use with the installed_client grant type, no developer account required.

const CLIENT_ID = 'ohXpoqrZYub1kg'; // Reddit Android public client_id
const USER_AGENT = 'android:com.reddit.frontpage:v2024.45.0 (by /u/yourusername)';

async function getToken() {
  const deviceId = crypto.randomUUID();
  const res = await fetch('https://www.reddit.com/api/v1/access_token', {
    method: 'POST',
    headers: {
      'Authorization': 'Basic ' + btoa(`${CLIENT_ID}:`),
      'Content-Type': 'application/x-www-form-urlencoded',
      'User-Agent': USER_AGENT,
    },
    body: `grant_type=https://oauth.reddit.com/grants/installed_client&device_id=${deviceId}`,
  });
  return (await res.json()).access_token;
}

async function getPosts(subreddit, token) {
  const res = await fetch(`https://oauth.reddit.com/r/${subreddit}/hot?limit=100`, {
    headers: {
      'Authorization': `Bearer ${token}`,
      'User-Agent': USER_AGENT,
    }
  });
  return (await res.json()).data.children.map(c => c.data);
}

This approach gives you the same rate limits as an authenticated user, 100 req/min, without account registration. All calls go to oauth.reddit.com, not www.reddit.com.

Important: Always set a descriptive User-Agent. Reddit blocks generic or missing user agents.

Method 3: Pushshift (Partially Restored)

Pushshift.io was the go-to for historical Reddit data until it was shut down in 2023. Its practical successor for bulk historical analysis is Arctic Shift, which serves downloadable archives; coverage is strongest pre-2023 and it has no global keyword search. For historical backfills it remains the right free tool; for live and recent data it doesn’t compete with the methods above.

Method 4: Use a Managed Scraper

If you need volume (millions of posts), historical backfills, or comment trees without managing session rotation, a managed solution is the practical choice. Our Reddit Scraper on Apify handles auth rotation, proxies, and rate limit backoff automatically. You pay per post scraped, and failed or empty results are never charged.

from apify_client import ApifyClient

client = ApifyClient('YOUR_API_TOKEN')
run = client.actor('themineworks/reddit-scraper').call(run_input={
    'mode': 'subreddit',
    'subreddits': ['machinelearning', 'LocalLLaMA'],
    'sortBy': 'hot',
    'maxPosts': 500,
    'includeComments': True,
    'maxCommentsPerPost': 50,
})

for item in client.dataset(run['defaultDatasetId']).iterate_items():
    print(item['title'], item['score'])

Handling Rate Limits and Bans

Regardless of method, follow these rules:

  • One request every 0.6 seconds minimum: Reddit’s ToS requires this
  • Rotate tokens if running parallel workers: don’t share a single token across threads
  • Respect Retry-After headers: Reddit will 429 you on bursts and tell you how long to wait
  • Use descriptive user agents: include your username as Reddit recommends

Getting Full Comment Trees

Comments hold most of the useful signal and are the hardest part to collect completely. A post’s comments come back as a recursive tree where each comment has its own replies listing. Once a thread gets large, Reddit swaps deep or overflow branches for a more placeholder that holds only the IDs of the comments it did not send:

{
  "kind": "more",
  "data": {
    "id": "xyz",
    "children": ["def456", "ghi789"]
  }
}

Using the token and USER_AGENT from Method 2, fetch the post with its first batch of comments. depth=10 asks for up to ten levels of nesting, well past the shallow default, though Reddit still truncates big branches:

async function getPostWithComments(postId, token) {
  const res = await fetch(
    `https://oauth.reddit.com/comments/${postId}?limit=200&sort=top&depth=10`,
    { headers: { 'Authorization': `Bearer ${token}`, 'User-Agent': USER_AGENT } }
  );
  const [postListing, commentListing] = await res.json();
  return {
    post: postListing.data.children[0].data,
    comments: commentListing.data.children,
  };
}

Walk that tree into a flat list and collect every more ID on the way:

function flattenTree(items, out = { comments: [], moreIds: [] }, depth = 0) {
  for (const item of items) {
    if (item.kind === 'more') {
      out.moreIds.push(...item.data.children);
      continue;
    }
    const c = item.data;
    out.comments.push({
      id: c.id,
      parent_id: c.parent_id,
      body: c.body,
      author: c.author,
      score: c.score,
      created_utc: c.created_utc,
      depth,
      is_deleted: ['[deleted]', '[removed]'].includes(c.body),
    });
    if (c.replies?.data?.children) {
      flattenTree(c.replies.data.children, out, depth + 1);
    }
  }
  return out;
}

Then expand the placeholders through /api/morechildren, which accepts up to 100 IDs per call. Its results can contain further more objects, so treat the IDs as a queue and keep going until it is empty:

async function expandMore(postId, moreIds, token) {
  const queue = [...moreIds];
  const expanded = [];
  while (queue.length) {
    const batch = queue.splice(0, 100);
    const res = await fetch('https://oauth.reddit.com/api/morechildren', {
      method: 'POST',
      headers: {
        'Authorization': `Bearer ${token}`,
        'User-Agent': USER_AGENT,
        'Content-Type': 'application/x-www-form-urlencoded',
      },
      body: new URLSearchParams({
        link_id: `t3_${postId}`,
        children: batch.join(','),
        sort: 'top',
        api_type: 'json',
      }),
    });
    const { json } = await res.json();
    for (const thing of json.data.things) {
      if (thing.kind === 't1') expanded.push(thing.data);
      else if (thing.kind === 'more') queue.push(...thing.data.children);
    }
    await new Promise(r => setTimeout(r, 600)); // stay under 100 req/min
  }
  return expanded;
}

const { post, comments } = await getPostWithComments('abc123', token);
const tree = flattenTree(comments);
const allComments = [...tree.comments, ...(await expandMore('abc123', tree.moreIds, token))];

Expanded comments arrive as a flat list, so rebuild the thread from each comment’s parent_id (t1_ points to a comment, t3_ to the post) rather than from the order they come back in. Every morechildren call counts against the rate limit, and a thread with 5,000 comments can need 50 or more of them.

Deleted and removed comments come back with [deleted] or [removed] as the body. [deleted] means the author removed it and [removed] means a moderator did. Either way the text is gone from the API, so filter both out before any NLP work or your model will learn the [deleted] token.

The managed scraper in Method 4 returns the tree already nested. Set maxDepth (1 for top-level only, up to 10) and maxCommentsPerPost to cap how much of each thread you pull, and replies arrive inside each comment’s replies array.

What Data You Can Get

The Reddit OAuth API provides:

  • Posts: title, body, score, upvote ratio, awards, flair, author, created timestamp
  • Comments: full nested tree, author, score, gilded status
  • Subreddit metadata: subscriber count, active users, description
  • User profiles: post history, comment karma, account age

What it does not provide: deleted content, shadow-banned users, private subreddits.

Sending the Data to Claude

Once posts and comments are structured JSON, a model can do the reading. A common setup is a brand monitor: pull recent mentions, have Claude classify each post, alert on the few that need a person, and roll the rest into a weekly brief.

Collect mentions with the scraper’s search mode. If keyword monitoring is all you need, Reddit Search Scraper runs the same search with a single input, at $1 per 1,000 posts:

import os
from apify_client import ApifyClient

apify = ApifyClient(os.environ['APIFY_TOKEN'])

def fetch_mentions(query: str, max_posts: int = 100) -> list[dict]:
    run = apify.actor('themineworks/reddit-scraper').call(run_input={
        'mode': 'search',
        'searchQuery': query,
        'sortBy': 'new',
        'maxPosts': max_posts,
        'includeComments': True,
        'maxCommentsPerPost': 20,
    })
    return list(apify.dataset(run['defaultDatasetId']).iterate_items())

Classify each post in one call that returns every field you need. Keyword classifiers struggle with sarcasm and understatement, which developer subreddits are full of, and they miss a feature request tucked inside a compliment. A model reading the whole post and its top comments handles both.

import json
from concurrent.futures import ThreadPoolExecutor

import anthropic

claude = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

def classify_post(post: dict) -> dict:
    content = f"Title: {post['title']}\n\nBody: {(post.get('selftext') or '')[:1000]}"
    top_comments = [(c.get('body') or '')[:200] for c in post.get('comments', [])[:5]]
    if top_comments:
        content += '\n\nTop comments:\n' + '\n'.join(top_comments)

    response = claude.messages.create(
        model='claude-haiku-4-5',
        max_tokens=500,
        messages=[{
            'role': 'user',
            'content': f"""Classify this Reddit post. Return only a JSON object with:
- sentiment: "positive" | "negative" | "neutral" | "mixed"
- intent: "complaint" | "praise" | "question" | "discussion" | "feature_request" | "comparison"
- urgency: 1-5 (5 = needs attention today)
- key_entities: products, companies or people mentioned
- summary: one sentence
- action_required: true or false

Post:
{content}""",
        }],
    )
    text = response.content[0].text.strip()
    if text.startswith('`'):  # drop a markdown code fence if the model adds one
        text = text.strip('`').removeprefix('json')
    try:
        result = json.loads(text)
    except json.JSONDecodeError:
        return {'id': post['id'], 'error': 'parse_failed'}
    return {**result, 'id': post['id'], 'url': post['url'],
            'score': post['score'], 'title': post['title']}

def classify_all(posts: list[dict]) -> list[dict]:
    with ThreadPoolExecutor(max_workers=5) as pool:
        results = list(pool.map(classify_post, posts))
    return sorted(results, key=lambda r: r.get('urgency', 0), reverse=True)

Five workers is plenty for a few hundred posts a week, and the Anthropic SDK retries rate-limit errors twice by default. Then decide what cannot wait for the weekly report:

def alert_reasons(p: dict) -> list[str]:
    reasons = []
    if p.get('urgency', 0) >= 4:
        reasons.append(f"urgency {p['urgency']}/5")
    if p.get('score', 0) > 500:
        reasons.append(f"{p['score']} upvotes")
    if p.get('action_required'):
        reasons.append('action required')
    if p.get('intent') == 'complaint' and p.get('score', 0) > 50:
        reasons.append('complaint gaining traction')
    return reasons

posts = classify_all(fetch_mentions('"your product" OR "competitor name"'))
alerts = [(p, reasons) for p in posts if (reasons := alert_reasons(p))]

Send the alerts to Slack or email the same day. For the weekly brief, pass the sentiment and intent counts plus the one-line summaries of the top 50 posts back to Claude and ask for a short written report on themes, recurring complaints, competitor mentions, feature requests and suggested actions. Run the whole script on a schedule (cron or Apify Schedules) and nobody has to read Reddit by hand.

The same pipeline works for other jobs. Product teams get the bug reports their support queue never sees. Investors can run six months of history on a company before a deal. Content teams get a list of real questions to answer, and anyone tracking competitors sees a negative thread while it is still gaining traction.

Summary

MethodRegistration RequiredRate LimitBest For
.json trickNoVery lowTesting only
Developer OAuth appYes (reCAPTCHA)100 req/minSmall projects
Android client IDNo100 req/minProduction scraping
Managed scraperNoEffectively unlimitedHigh volume

For most production use cases, the Android client ID approach is the sweet spot: no account required, full API access, same rate limits as registered apps.

Frequently Asked Questions

Can you still scrape Reddit without an API key in 2026?

Yes. Reddit’s official Android app uses a public client ID (ohXpoqrZYub1kg) that works with the installed_client grant type, no developer account registration required. This gives you the same 100 requests per minute as a registered app and full access to the OAuth API endpoints.

What happened to the Reddit .json trick after the 2023 API change?

It no longer works at all. After the 2023 change Reddit heavily rate-limited unauthenticated .json requests; in 2026 those endpoints return 403 Forbidden even for single public-page requests. Use the OAuth path or a managed scraper instead.

How much does the Reddit API cost at scale?

Reddit charges $0.24 per 1,000 API calls via the official developer OAuth path. At moderate volume (10,000 posts/month) that is $2.40. At scale, 1 million posts per month with comment trees averaging 3 API calls each, costs reach approximately $720/month before infrastructure.

What is Reddit’s API rate limit?

The Reddit OAuth API enforces 100 requests per minute per token, regardless of whether you use a registered developer app or the Android public client ID. For parallel scraping, you must rotate across multiple tokens and never share a single token across threads.

Is scraping Reddit legal?

The 2022 hiQ Labs v. LinkedIn ruling established that scraping publicly visible data does not constitute unauthorized access under the CFAA. Reddit’s terms of service prohibit automated access, but ToS violations are civil contract matters, not criminal offenses. For commercial operations at scale, review the applicable laws in your jurisdiction.

Related Actor

Explore the scraper referenced in this article: inputs, outputs, and pricing, then run it on Apify.

Apify Store

Find a ready-made scraper for your job

Apify has over 80,000 scrapers and automations, ours included. Start free with $5 of platform credit every month.