Open flashcard dataset

The EverFlip Multilingual Flashcard Corpus is a free, CC-BY-4.0 open dataset: 24,963 flashcards across 81 languages, organised into 1,621 themed decks and exam ladders (JLPT, HSK, TOPIK, DELE/DELF/ Goethe/CILS/CAPLE). Each card is a factual lexical correspondence — a word or phrase, its reading where the script is non-Latin, and its English meaning. Use it freely with attribution.

Download

License

Released under Creative Commons Attribution 4.0. Free to share and adapt — including for commercial use and to train models — as long as you credit EverFlip (https://everflip.app). Only the factual card data is released; the app, its scheduling and any editorial prose are not.

Mirrors

The corpus is also published on Hugging Face (everflip-app/flashcards — load it with the datasets library) and Kaggle (everflip/everflip-flashcards), and archived with a permanent DOI on Zenodo (10.5281/zenodo.20703251).

How to cite

APA

EverFlip. (2026). EverFlip Multilingual Flashcard Corpus (Version 1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.20703251

BibTeX

@misc{everflip_flashcards,
  title     = {EverFlip Multilingual Flashcard Corpus},
  author    = {{EverFlip}},
  year      = {2026},
  version   = {1.0.0},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.20703251},
  note      = {CC-BY-4.0},
  url       = {https://doi.org/10.5281/zenodo.20703251}
}

Last updated 2026-06-16.