Open flashcard dataset
The EverFlip Multilingual Flashcard Corpus is a free, CC-BY-4.0 open dataset: 24,963 flashcards across 81 languages, organised into 1,621 themed decks and exam ladders (JLPT, HSK, TOPIK, DELE/DELF/ Goethe/CILS/CAPLE). Each card is a factual lexical correspondence — a word or phrase, its reading where the script is non-Latin, and its English meaning. Use it freely with attribution.
Download
- cards.csv — every card (24,963 rows): language, deck, front, English meaning, reading.
- decks.csv — deck catalogue (1,621 rows).
- languages.csv — language catalogue (81 rows).
- datapackage.json — frictionlessdata.io Tabular Data Package descriptor.
License
Released under Creative Commons Attribution 4.0. Free to share and adapt — including for commercial use and to train models — as long as you credit EverFlip (https://everflip.app). Only the factual card data is released; the app, its scheduling and any editorial prose are not.
Mirrors
The corpus is also published on Hugging Face (everflip-app/flashcards — load it with the datasets library) and Kaggle (everflip/everflip-flashcards), and archived with a permanent DOI on Zenodo (10.5281/zenodo.20703251).
How to cite
APA
EverFlip. (2026). EverFlip Multilingual Flashcard Corpus (Version 1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.20703251
BibTeX
@misc{everflip_flashcards,
title = {EverFlip Multilingual Flashcard Corpus},
author = {{EverFlip}},
year = {2026},
version = {1.0.0},
publisher = {Zenodo},
doi = {10.5281/zenodo.20703251},
note = {CC-BY-4.0},
url = {https://doi.org/10.5281/zenodo.20703251}
}Last updated 2026-06-16.