Media, Social & Knowledge

How do you scrape Reddit posts via API without getting blocked?

Call ReefAPI's reddit subreddit_posts or search action and read posts with id, title, author, score, upvote_ratio, comment count, selftext and timestamps back as JSON. The single setting that decides whether search is useful is sort: with sort 'new' across all of Reddit, your keywords stop being enforced.

Reddit engineLive JSON5 steps1,000 free credits

This guide demonstrates the real Reddit API engine with a captured response from . The example is only published because the engine passed the SEO snapshot gate.

Use case

Reddit post monitoring, community research, audience intelligence and social listening.

Step by step

Call the live endpoint

  1. 1

    Prefer subreddit_posts over site-wide search

    Pass subreddit, sort ('top', 'new', 'hot') and time ('hour' through 'all'). Scoped listings do not relax your intent the way keyword search does.

  2. 2

    If you must search, use sort relevance

    Site-wide search with sort 'new' returned zero on-topic titles for our three-word query; the same query with 'relevance' returned eight of ten. Add subreddit to tighten it further.

  3. 3

    Store url and permalink separately

    They are the same string on self posts and different on link posts. permalink is the discussion, url is what was linked.

  4. 4

    Page on the fullname cursor

    pagination.next_cursor is a t3_ fullname. Pass it as after or cursor, and stop on has_more false rather than on an empty page.

  5. 5

    Judge posts on score plus upvote_ratio

    Reddit fuzzes the raw vote counts but publishes the ratio, so the two together tell you whether a high score is agreement or controversy.

Code

Copy the request

These snippets use the captured request params for reddit/v1/search.

curl -X POST https://api.reefapi.com/reddit/v1/search \
  -H "x-api-key: $REEF_KEY" \
  -H "content-type: application/json" \
  -d '{"subreddit":"python","limit":10}'
MCP one-liner
Ask your MCP-connected assistant: call reefapi.reddit.search with {"subreddit":"python","limit":10}.
Real response

Captured output from ReefAPI

Captured on UTC. The response below is the committed snapshot, including the API envelope and metadata.

Captured request
{
  "method": "POST",
  "url": "https://api.reefapi.com/reddit/v1/search",
  "headers": {
    "x-api-key": "$REEF_KEY",
    "content-type": "application/json"
  },
  "body": {
    "subreddit": "python",
    "limit": 10
  }
}
Captured response
{
  "ok": true,
  "meta": {
    "api": "reddit",
    "endpoint": "search",
    "mode": "live",
    "latency_ms": 2600.9,
    "record_count": 10,
    "bytes": 42214,
    "cache_hit": false,
    "completeness_pct": 100,
    "requests": 3,
    "source_used": "arctic",
    "pagination": {
      "next_cursor": null,
      "has_more": false
    },
    "charged_credits": 1,
    "version": "0.4.0"
  },
  "data": {
    "results": [
      {
        "id": "1woijc7",
        "fullname": "t3_1woijc7",
        "title": "CryptoClipGuard: A real-time Windows clipboard-hijacking mitigation tool for crypto addresses",
        "author": "Traditional_Bear5492",
        "author_fullname": "t2_opcw5jkmi",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 1,
        "created_utc": 1790198234,
        "url": "https://www.reddit.com/r/Python/comments/1woijc7/cryptoclipguard_a_realtime_windows/",
        "permalink": "https://www.reddit.com/r/Python/comments/1woijc7/cryptoclipguard_a_realtime_windows/",
        "selftext": "[removed]",
        "flair": "Showcase",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1woh8y4",
        "fullname": "t3_1woh8y4",
        "title": "What's new in Python 3.15?",
        "author": "AlSweigart",
        "author_fullname": "t2_369qx",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 0,
        "created_utc": 1790195217,
        "url": "https://www.reddit.com/r/Python/comments/1woh8y4/whats_new_in_python_315/",
        "permalink": "https://www.reddit.com/r/Python/comments/1woh8y4/whats_new_in_python_315/",
        "selftext": "https://docs.python.org/3.15/whatsnew/3.15.html\n\n> Summary – Release highlights\n\n> Python 3.15 will be the latest stable release of the Python programming language, with a mix of changes to the language, the implementation, and the standard library. The biggest changes include lazy imports, frozendict and sentinel builtins, UTF-8 as the default encoding, unpacking in comprehensions, and a stable ABI for free-threaded builds.\n\n> The library changes include a new profiling package with Tachyon, a high-frequency statistical sampling profiler, more color in command-line output, as well as the usual deprecations and removals, and improvements in user-friendliness and correctness.\n\n> This article doesn’t attempt to provide a complete specification of all new features, but instead gives a convenient overview. For full details refer to the documentation, such as the Library Reference and Language Reference. To understand the complete implementation and design rationale for a change, refer to the PEP for a particular new feature; but note that PEPs usually are not kept up-to-date once a feature has been fully implemented.",
        "flair": "News",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1woct5v",
        "fullname": "t3_1woct5v",
        "title": "Built a post-quantum crypto lib in Python (ML-KEM + AES). Feedback?",
        "author": "Bratva250",
        "author_fullname": "t2_2nahm0tiji",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 1,
        "created_utc": 1790185319,
        "url": "https://www.reddit.com/r/Python/comments/1woct5v/built_a_postquantum_crypto_lib_in_python_mlkem/",
        "permalink": "https://www.reddit.com/r/Python/comments/1woct5v/built_a_postquantum_crypto_lib_in_python_mlkem/",
        "selftext": "[removed]",
        "flair": "Discussion",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1woca2r",
        "fullname": "t3_1woca2r",
        "title": "Best Fun Beginners Guides for Python - 2026",
        "author": "vkailas",
        "author_fullname": "t2_3gdc7",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 0,
        "created_utc": 1790184130,
        "url": "https://www.reddit.com/r/Python/comments/1woca2r/best_fun_beginners_guides_for_python_2026/",
        "permalink": "https://www.reddit.com/r/Python/comments/1woca2r/best_fun_beginners_guides_for_python_2026/",
        "selftext": "[removed]",
        "flair": "Discussion",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1wo7f3o",
        "fullname": "t3_1wo7f3o",
        "title": "jobHunt: a stdlib-only job search + CV tailoring tool with a local web UI",
        "author": "TackleAlone5814",
        "author_fullname": "t2_1q26nyrzzw",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 1,
        "created_utc": 1790173380,
        "url": "https://www.reddit.com/r/Python/comments/1wo7f3o/jobhunt_a_stdlibonly_job_search_cv_tailoring_tool/",
        "permalink": "https://www.reddit.com/r/Python/comments/1wo7f3o/jobhunt_a_stdlibonly_job_search_cv_tailoring_tool/",
        "selftext": "[removed]",
        "flair": "Showcase",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1wo7a20",
        "fullname": "t3_1wo7a20",
        "title": "CutCutCodec: Streamlined Video Processing – A Signal Processing Oriented MoviePy Alternative",
        "author": "robinechuca",
        "author_fullname": "t2_22akfn4w7r",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 0,
        "created_utc": 1790173058,
        "url": "https://www.reddit.com/r/Python/comments/1wo7a20/cutcutcodec_streamlined_video_processing_a_signal/",
        "permalink": "https://www.reddit.com/r/Python/comments/1wo7a20/cutcutcodec_streamlined_video_processing_a_signal/",
        "selftext": "After two years of work, *entierly hand-coded*, the audio and video editing library is finally up and running!  \nHere is the project documentation: [https://cutcutcodec.readthedocs.io/latest/](https://cutcutcodec.readthedocs.io/latest/)\n\nThis project is aimed at developers and the signal processing and machine learning community. It is based on Torch and rigorously implements the standards of the International Telecommunication Union.\n\nUnlike moviepy, which is based on the concept of clips, CutCutCodec is based on the concepts of media streams and editing graphs, which are more powerful, but slightly more verbose.\n\n*Let see a little basic example:*\n\n    import cutcutcodec\n    container = cutcutcodec.read(\"input_video.mp4\")  # open the file\n    stream = (\n        container.out_select(\"video\")[0]  # first video stream\n        .apply_video_subclip(0, 10)  # keep the 10 first seconds\n        .apply_video_equation(\"r0\", \"g0\", \".5*b0*sin(2*pi*0.5*t)+.5\")  # blink blue\n    )\n    streams_settings = [{\"encodec\": \"libx264\", \"rate\": 30, \"shape\": (480, 720)}]  # optional\n    cutcutcodec.write([stream], \"output_video.mp4\", streams_settings=streams_settings)\n\nIf you have any feedback for me, I’d love to hear it!",
        "flair": "Resource",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1wo6jek",
        "fullname": "t3_1wo6jek",
        "title": "Sanka: open-source DRF → FastAPI migration engine with AI bench",
        "author": "ageo89",
        "author_fullname": "t2_5fzbfacc",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 0,
        "created_utc": 1790171347,
        "url": "https://www.reddit.com/r/Python/comments/1wo6jek/sanka_opensource_drf_fastapi_migration_engine/",
        "permalink": "https://www.reddit.com/r/Python/comments/1wo6jek/sanka_opensource_drf_fastapi_migration_engine/",
        "selftext": "**What this guides show you**\n\nSanka is a migration engine plus CLI. \n\nFive commands take a Django REST Framework app to FastAPI: \n\n`scan` reads the app the way Django runs it (URLconf, routers, viewsets, serializers), \n\n`plan` writes a hash-locked plan you review, \n\n`apply` accepts only that exact hash and writes a separate FastAPI project (async handlers, Pydantic models, Tortoise ORM models on your existing tables) without touching your source, \n\n`test` runs the generated app's own suite, and \n\n`verify` replays requests against both apps and diffs the responses.\n\nRoutes outside the supported envelope are reported as manual work, never silently bridged. Everything runs locally; no account.\n\nRepo: [https://github.com/sankaHQ/sanka](https://github.com/sankaHQ/sanka)\n\nGuide (each step also runs in the browser): [https://sanka.com/docs/developers/migrate/django-to-fastapi/](https://sanka.com/docs/developers/migrate/django-to-fastapi/)\n\nWalkthrough video: [https://www.youtube.com/watch?v=iH1dgWGRn6g](https://www.youtube.com/watch?v=iH1dgWGRn6g)\n\n**Target Audience**\n\nTeams moving a DRF API to FastAPI (or Flask) who want a reviewable plan and a parity check instead of a one-off script.\n\nProduction intent; this is the first public release. Python 3.12+, uv, macOS/Linux installer. Engine AGPL-3.0, CLI/SDK Apache-2.0.\n\n**Comparison**\n\n\\- Hand port: no plan artifact, no differential verify; \"looks right\" is the review.\n\n\\- django-ninja: you stay on Django, which is the right call if you can; Sanka is for teams that have decided to leave.\n\n\\- One-shot LLM rewrite: no reviewable plan and no parity check. We measured the difference on 17 migrations with 8 models, same prompts with and without the CLI: 74.3% → 85.3% pass, and the weakest models gain the most (https://sanka.com/bench, evaluator and tasks open at https://github.com/sankaHQ/bench).",
        "flair": "Tutorial",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1wo6ak8",
        "fullname": "t3_1wo6ak8",
        "title": "Should dependency scanners flag unmaintained packages, or stick to CVEs?",
        "author": "No_Wedding2230",
        "author_fullname": "t2_qwwzaa2x",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 0,
        "created_utc": 1790170759,
        "url": "https://www.reddit.com/r/Python/comments/1wo6ak8/should_dependency_scanners_flag_unmaintained/",
        "permalink": "https://www.reddit.com/r/Python/comments/1wo6ak8/should_dependency_scanners_flag_unmaintained/",
        "selftext": "Should a dependency scanner tell you a package is unmaintained, or only that it has a known vulnerability? \n\nThe case for CVEs only: a vulnerability is a fact, and \"unmaintained\" is a guess. The case against: a CVE-only scanner stays silent about a package with no maintainer, right up until someone finds a bug in it and nobody is there to fix it.\n\nI measured this on the pinned dependencies of 100 well-known open source Python projects (Django, FastAPI, Airflow, Home Assistant, Sentry, vLLM and others) on 23 September. 98 had dependency files I could read, with 4,791 distinct packages between them.\n\n**Release age is almost useless on its own.** 988 of those packages, about 1 in 5, haven't shipped a release in two years. Most are simply finished. Flag them all and nobody reads the report.\n\n**Some signals are facts rather than guesses.** 75 packages have no known vulnerability in the pinned version, but their source repository is archived, or the maintainer has marked them Inactive on PyPI. A CVE scanner says nothing about any of them. The most common:\n\n* `backoff`: archived, in 16 of the projects\n* `nest-asyncio`: 14\n* `bleach`: 11, deprecated by its maintainers\n* `click-plugins`: 10\n\n**Even the facts need judgement.** An archived repo isn't always an abandoned package: `google-cloud-bigquery`'s repo is archived because the code moved into a monorepo, and it still ships monthly. And `backoff` retries function calls, so nobody left to patch it matters much less than it would for an HTML sanitiser.\n\nSo what should a scanner do with maintenance signals?\n\n1. Nothing. CVEs only\n2. Report archived/deprecated packages, never fail a build\n3. Fail the build, but only when the abandoned package handles untrusted input (parsers, auth, sanitisers)\n4. A single health score",
        "flair": "Discussion",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1wo3o2r",
        "fullname": "t3_1wo3o2r",
        "title": "Anyone wanna be an AI Engineer , Can I use LangChain framework and Doing without SDK or else ??",
        "author": "a9a_k8",
        "author_fullname": "t2_114b605yyz",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 1,
        "created_utc": 1790163833,
        "url": "https://www.reddit.com/r/Python/comments/1wo3o2r/anyone_wanna_be_an_ai_engineer_can_i_use/",
        "permalink": "https://www.reddit.com/r/Python/comments/1wo3o2r/anyone_wanna_be_an_ai_engineer_can_i_use/",
        "selftext": "",
        "flair": "Help",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      },
      {
        "id": "1wo314v",
        "fullname": "t3_1wo314v",
        "title": "Someone hijacked MemoryOS PyPI releases by replacing the build backend",
        "author": "BattleRemote3157",
        "author_fullname": "t2_ewn6cu80w",
        "subreddit": "Python",
        "score": 1,
        "upvote_ratio": 1,
        "num_comments": 0,
        "created_utc": 1790161930,
        "url": "https://www.reddit.com/r/Python/comments/1wo314v/someone_hijacked_memoryos_pypi_releases_by/",
        "permalink": "https://www.reddit.com/r/Python/comments/1wo314v/someone_hijacked_memoryos_pypi_releases_by/",
        "selftext": "The attacker swapped in a custom `pyproject.toml` build backend that grabbed the PyPI token before the real upload ran. Then used that token to push the backdoored package themselves.\n\ncomplete - [safedep.io/memtensor-sckit-worm-npm-pypi](http://safedep.io/memtensor-sckit-worm-npm-pypi)",
        "flair": "News",
        "over_18": false,
        "spoiler": false,
        "stickied": false,
        "locked": false,
        "is_self": true,
        "is_video": false,
        "domain": "self.Python",
        "thumbnail": null,
        "num_crossposts": 0,
        "total_awards_received": 0,
        "edited": false
      }
    ],
    "count": 10,
    "type": "post"
  }
}
Manual way

Why this is hard manually

Reddit's site-wide search relaxes your query when you sort by recency, and it does not tell you it has done so. We searched 'web scraping api' with sort 'new' and got ten results of which zero had the word 'scraping' anywhere in the title - the top hit was a fan-fiction post from r/TheLast20DollarsLore. The identical query with sort 'relevance' returned eight of ten with 'scraping' in the title, from r/webscraping, r/Python and r/learnprogramming. Same query, same limit, one parameter apart.

This matters more than it sounds, because 'newest posts matching my keyword' is exactly what a monitoring job wants, and it is the one combination that silently returns noise. Every result still comes back with ok:true, a full payload and a plausible-looking score. There is no relevance field to threshold on.

The second thing worth knowing before you store anything: on Reddit a post's url and its permalink are not the same field, and for link posts they point at different places entirely.

ReefAPI way

Why ReefAPI solves it

Sort deliberately, and scope when you can. For monitoring, the combination that works is sort 'relevance' with a subreddit set: 'web scraping api' inside r/webdev returned eight of ten on-topic titles including 'Web Scraping instead of Reddit API?' and 'Protecting public JSON API responses from scraping'. If you need recency, pull the subreddit's own new feed via subreddit_posts and filter client-side rather than asking site-wide search for it.

url is the destination, permalink is the discussion. On a self post they are the same Reddit URL. On a link post they diverge: one of our results returned url 'https://i.redd.it/waj0w3t8kslh1.png' with permalink pointing at the reddit.com comments thread. If you are building a feed reader you want permalink; if you are archiving the linked content you want url. Storing only one of them loses information you cannot recover later.

Every post carries both an id and a fullname - '1vyikxi' and 't3_1vyikxi'. The t3_ prefix is Reddit's type tag for a link/post (comments are t1_, subreddits t5_), and it is what the pagination cursor is made of: pagination.next_cursor came back as 't3_1vvdyqq'. Feed that back as after or cursor for the next page, and read pagination.has_more rather than guessing from result count.

upvote_ratio is the field that stops score from lying to you. A post with score 524 and upvote_ratio 0.96 is genuinely well received; the same score at 0.55 is a fight. Reddit fuzzes raw up and down counts but publishes the ratio, so the pair is the honest signal. num_comments sits beside it, and created_utc is a float epoch in UTC - not a formatted string, so no parsing and no timezone guesswork.

selftext arrives in full in the listing response rather than behind a per-post fetch, which means one call gives you the complete text corpus for a subreddit page. That is what makes cheap classification and keyword monitoring practical: our r/webdev top-of-week call returned ten posts with full bodies in 858 ms.

Eleven actions share the engine: subreddit_posts, multi_subreddit_posts, post_comments, load_more_comments, search, user, user_search, subreddit_about, subreddit_extras, communities and trending. multi_subreddit_posts takes a list of subreddits and a since_utc watermark in one call, which is the right shape for a monitor watching twenty communities. No Reddit account, app registration or OAuth token is involved.

FAQ

Questions developers ask

Do I need a Reddit account or OAuth app?

No. You send a ReefAPI key. There is no Reddit app registration, no OAuth flow and no per-app rate budget to manage.

My search returned completely unrelated posts. What did I do wrong?

Almost certainly sort 'new' on a site-wide search. We reproduced it: 'web scraping api' with sort 'new' returned zero titles containing 'scraping'; with sort 'relevance' it returned eight of ten. Reddit relaxes the query when it sorts by recency and does not flag that it has.

How do I get recent posts about a keyword then?

Pull the communities you care about with subreddit_posts or multi_subreddit_posts sorted by new, and filter on your keyword yourself. multi_subreddit_posts accepts a since_utc watermark so a monitor can ask only for what is new.

What is the difference between id and fullname?

fullname is the id with Reddit's type prefix - 't3_1vyikxi' for a post, t1_ for a comment, t5_ for a subreddit. Pagination cursors are fullnames, so keep it rather than reconstructing it.

Why do url and permalink differ on some posts?

Because a link post points somewhere else. One of our results had url 'https://i.redd.it/...png' and a permalink to the comments thread. Self posts have them identical. Store both.

Is score a reliable popularity measure?

Only with upvote_ratio beside it. Reddit deliberately fuzzes the raw up and down counts but publishes the ratio, so 524 at 0.96 and 524 at 0.55 are very different posts.

Do I get the full post text or just a preview?

The full selftext arrives in the listing response. There is no second fetch per post, which is what makes classifying a whole subreddit page in one call affordable.