Reddit RAG Dataset — LLM Training Data
Scrape reddit.com — build clean LLM and RAG datasets from posts with full comment threads as ready-to-chunk text · HTML & Markdown · only text-bearing records with parent/child thread structure. Mix subreddit, keyword-search and direct-post inputs in one run.
Publisher description, verbatim
Standing
#507of 652
on Reddit, by users in the last 30 days among actors classified to this platform.
Recorded changes
showing 3 of 5 in the last 30 days · followers get all 5 by email
- 2026-08-24contentREADME updated+671 characters+671characters
- 2026-08-19searchKeyword rank moveleft top-10 for “dataset” (#10 → outside top 100)—on record
- 2026-08-18searchKeyword rank moveentered top-10 for “dataset” (outside top 100 → #10)—on record
Follow this actor, compare it against its rivals, and get
every overnight change in one morning report.
Publish it yourself? Claim your publisher page and its full history is
yours free, forever.
Open in the app
Publish this actor? Claim the page
Numbers on this page are measured only by an independent census of the Apify Store — not self-reported.