A collection of high quality datasets from huggingface.
Highly filtered pretraining datasets.
Warning: Writing is extremely underrepresented. I am in the process of looking for writing datasets.
- https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu
- https://huggingface.co/datasets/HuggingFaceFW/finewiki
Both finepdfs-edu and finewiki contain dense multilingual data that has been collected into language-specific subsets (en for eng finewiki, eng_Latn for eng finepdfs). Finewiki appears to be organized into markdown, finepdfs-edu does not appear particularly formatted. In finepdfs-edu, the list fw_edu_scores should contain the education scores.
Contains a generic mixture of exerts from websites. The language is under the 'language' column, score is under the 'score' column. Hopefully a fineweb2-edu will be out soon to replace this.
An alternative to fineweb-edu by openbmb (ultrachat, minicpm authors). Language subsets (en for eng) and 'score' column.
Compared to finemath which is full of noise and honestly quite terrible, l2 is quite decent for a filtered math dataset.
Surprisingly decent.
https://huggingface.co/datasets/allenai/dolma3_longmino_mix-100B-1125
Could be good for long context.
WIP, worth paying attention to:
- https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3
- https://huggingface.co/datasets/math-ai/AutoMathText-2.1
Honorable mentions (unfiltered):
- https://huggingface.co/datasets/ibm-granite/GneissWeb
- https://huggingface.co/datasets/openbmb/DCAD-2000
These datasets can be slotted into a pretraining run at the end for curriculum learning or mixed throughout. Remember that midtraining datasets must be very large but can be lower quality; SFT is the opposite.
Warning: I extensively cite NVIDIA datasets which have been known to produce extremely dry models.
- https://huggingface.co/datasets/nvidia/Nemotron-Cascade-SFT-Stage-2, https://huggingface.co/datasets/nvidia/Nemotron-Cascade-SFT-Stage-1
Nemotron Cascade is extremely extensive, with the one issue being that responses are generated with Deepseek R1/V3 instead of newer models such as GLM 5 or even DSV3.2 (thus I would not recommend it for SFT).
- https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-SFT-Data (Math subset only for DSV3.2 Speciale responses)
- https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v3 (DSV3.2)
- https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v2 (GLM 4.6, Kimi K2)
- https://huggingface.co/datasets/nvidia/Nemotron-SFT-SWE-v2 (Qwen3-Coder-480B-A35B-Instruct)
- https://huggingface.co/datasets/nvidia/Nemotron-SFT-Agentic-v2 (Deepseek 3.2, GLM 4.6)
These are a bit more niche compared to Nemotron-Cascade-SFT, and might be biased towards certain domains.
Also note that in my opinion, Kimi K2.5 is way better than Kimi K2, and GLM 5 is way better than 4.6, so these datasets have depreciated quite a bit.
Generic multiturn conversational data for training.
-
https://huggingface.co/datasets/stepfun-ai/Step-3.5-Flash-SFT (Contains a mixture of multiturn, reasoning, and agentic)
-
https://huggingface.co/datasets/nothingiisreal/Kalomaze-Opus-Instruct-25k-filtered
-
https://huggingface.co/datasets/allenai/WildChat-4.8M (Warning: GPT-4 era models, language roughly listed under 'language')
-
https://huggingface.co/datasets/lmarena-ai/arena-expert-5k/ (Language listed under language column, many models)
-
https://huggingface.co/datasets/lmarena-ai/arena-human-preference-140k (Language listed under language column, many models)
Honorable Mentions:
- https://huggingface.co/datasets/bjoernp/oasst25-08-23-filtered (Contains language tags within each chat block.)
- https://huggingface.co/datasets/HuggingFaceH4/no_robots
- https://huggingface.co/datasets/databricks/databricks-dolly-15k
These three are human written, but are quite outdated and low quality compared to synthetic LLM data.
- https://huggingface.co/datasets/NousResearch/Hermes-3-Dataset (Honestly outdated and a bit mediocre.)
- https://huggingface.co/datasets/CohereLabs/aya_dataset, https://huggingface.co/datasets/CohereLabs/xP3x (For multilingual, translation)
Math and science datasets. I will only put the most noteworthy ones here. If you need a lot of math data, check the midtraining data which will be sufficient.
-
https://huggingface.co/datasets/microsoft/mediflow (Warning: GPT4o)
- https://huggingface.co/datasets/allenai/Sera-4.6-Lite-T2
- https://huggingface.co/datasets/deepmind/code_contests
- https://huggingface.co/datasets/lmarena-ai/webdev-arena-preference-10k
- https://huggingface.co/datasets/facebook/BigOBench
- https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Heimdall-v1.1
- https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.0
- https://huggingface.co/datasets/LipengCS/Table-GPT
- https://huggingface.co/datasets/deepmind/narrativeqa
- https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1
NSFW
- https://huggingface.co/datasets/nothingiisreal/Reddit-Dirty-And-WritingPrompts
- https://huggingface.co/datasets/Squish42/bluemoon-fandom-1-1-rp-cleaned
Non-NSFW
- https://huggingface.co/datasets/llm-aes/writing-prompts (Note: Prompts only, also punctuation is excessively spaced)
- https://huggingface.co/datasets/euclaise/writingprompts, https://huggingface.co/datasets/RLAIF/WritingPrompts-Filtered
- https://huggingface.co/datasets/ResplendentAI/bluemoon
Honestly, all of the above suck.
- https://huggingface.co/datasets/zjunlp/PredictBeforeExecute
- https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-405b-preference-mixture
- https://huggingface.co/datasets/facebook/menlo
- https://huggingface.co/datasets/facebook/community-alignment-dataset
- https://huggingface.co/datasets/Dahoas/full-hh-rlhf
- https://huggingface.co/datasets/internlm/OREAL-RL-Prompts
- https://huggingface.co/datasets/internlm/Lean-Workbook
- https://huggingface.co/datasets/microsoft/rStar-Coder
- https://huggingface.co/datasets/m-a-p/AetherCode
- https://github.com/open-thought/reasoning-gym
- https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-2-RL
- https://huggingface.co/datasets/facebook/principia-collection
- https://huggingface.co/datasets/internlm/SWE-Fixer-Train-110K
- https://huggingface.co/datasets/NousResearch/SWE-smith-oracle
- https://huggingface.co/datasets/zai-org/DeepDive
- https://huggingface.co/datasets/inclusionAI/ASearcher-train-data
- https://huggingface.co/datasets/inclusionAI/AReaL-boba-2-RL-Code
- https://huggingface.co/datasets/marianna13/zlib
- https://huggingface.co/datasets/marianna13/fanfics
- https://huggingface.co/datasets/marianna13/the-eye
- https://huggingface.co/datasets/marianna13/vault_text
- https://huggingface.co/datasets/marianna13/research_gate
- https://huggingface.co/datasets/marianna13/superuser
- https://huggingface.co/datasets/m-a-p/COIG-Writer
- https://huggingface.co/datasets/RWKV/RWKV-World-Listing
- https://huggingface.co/datasets/allenai/SciRIFF (Note: Prompts only)
- https://huggingface.co/EssentialAI/datasets
- https://huggingface.co/datasets/LLM360/TxT360
- https://huggingface.co/datasets/Kassadin88/GLM-5.1-1000000x
- https://huggingface.co/datasets/Roman1111111/claude-sonnet-4.6-120000x