Reddit RAG Dataset — LLM Training Data

Scrape reddit.com — build clean LLM and RAG datasets from posts with full comment threads as ready-to-chunk text · HTML & Markdown · only text-bearing records with parent/child thread structure. Mix subreddit, keyword-search and direct-post inputs in one run.

Publisher description, verbatim
Last 7 dayson this actor
Each one reaches followers by email the morning after.Follow this actor
Key figures30 readings
Price / 1k$2.00per 1,000 results
Users / 30d0→ 0 over 30 readings
Users, all time2↑ 1 over 30 readings
Runs / 30d31↑ 2 over 30 readings
Runs / user / moruns per monthly user
Users / 30d30 readings
08-10→ 009-08
Standing
#507of 652

on Reddit, by users in the last 30 days among actors classified to this platform.

Recorded changes

showing 3 of 5 in the last 30 days · followers get all 5 by email
  • 2026-08-24contentREADME updated+671 characters+671characters
  • 2026-08-19searchKeyword rank moveleft top-10 for “dataset” (#10 → outside top 100)on record
  • 2026-08-18searchKeyword rank moveentered top-10 for “dataset” (outside top 100 → #10)on record

Follow this actor, compare it against its rivals, and get every overnight change in one morning report. Publish it yourself? Claim your publisher page and its full history is yours free, forever.

Open in the app Publish this actor? Claim the page

Numbers on this page are measured only by an independent census of the Apify Store — not self-reported.