TOPIK Vocabulary: Two Official Numbers, Then Four Guesses
The study unit, defined. TOPIK vocabulary is the word stock that the Test of Proficiency in Korean, the exam Korea's education ministry has run since 1997, samples its six levels from. Exactly two of those six levels carry the exam's own word counts: Level 1, about 800 words; Level 2, 1,500 to 2,000. Above Level 2 the official paperwork stops counting, and prep-industry estimates for Level 6 spread from "6,000 plus" to "10,000 to 11,000." The same exam, counted by different people, and the answers differ by nearly double at the summit. The verdict this guide defends: the level you registered for is the only count that matters, the frequency core of that level is where the marks concentrate, and any TOPIK Korean vocabulary list you print is raw material for a scheduler, never the plan itself. Same disclosure as every guide in this series: I build Piccard, a flashcard tool, and English speakers learning Korean are one of its lanes.
Two papers, six levels
One sitting, one paper. Since the July 2014 reformat there are two exams. TOPIK I is the beginner paper: 30 listening questions in 40 minutes, then 40 reading questions in 60 minutes, all multiple choice, 200 points total; 80 earns Level 1, 140 earns Level 2 (Wikipedia, reproducing the administrator's grading criteria). TOPIK II covers the other four levels in one 300-point sitting: 50 listening questions, a 50-minute writing section of four tasks, two of them essays, and 50 reading questions in 70 minutes, with pass marks stepping 120, 150, 190, and 230 across Levels 3 to 6. Scores stay valid for two years.
The vocabulary consequence splits at the same seam. TOPIK I asks you only to recognize — hear the word, pick the match; read the blank, pick the filler — so recognition built on cards maps onto the whole paper. TOPIK II still runs on recognition for 100 of its listening points and most of its reading, but the writing section wants production, and no amount of flipping cards teaches you to produce. Cards carry you to the door of Level 3; the writing tasks need sentences you have written on purpose.
How many words does each TOPIK level want?
For Levels 1 and 2 you can quote the exam itself; for Levels 3 through 6 you are quoting somebody's reconstruction. The Level 1 descriptor, in the official criteria as Wikipedia reproduces them:
> "Able to create simple sentences based on about 800 basic vocabulary items and possess understanding of basic grammar."
And Level 2, from the same criteria:
> "Able to use about 1,500 to 2,000 vocabulary and understand personal and familiar subjects in certain order, such as paragraphing."
Then the descriptors for Levels 3 to 6 talk about "public facilities," "news broadcasts," "politics, economics, society, and culture" — functions and topics, not counts. The exam never told anyone how many words Level 4 wants, and as SEQ Toolkit puts it:
> "The official TOPIK administrator (NIIED) has never published a vocabulary list. Every list you see — including this one — is reverse-engineered from corpus analysis of past papers."
So the honest table of TOPIK vocabulary by level looks like this:
| Level | Paper | Words | Who's counting |
|---|---|---|---|
| 1 | TOPIK I | ~800 | the exam's own descriptor |
| 2 | TOPIK I | 1,500–2,000 | the exam's own descriptor |
| 3 | TOPIK II | ~3,000 | prep estimates, broadly agreeing |
| 4 | TOPIK II | 4,000–5,000 | prep estimates, broadly agreeing |
| 5 | TOPIK II | ~6,000–6,500 | prep estimates |
| 6 | TOPIK II | 6,000+ vs 10,000–11,000 | prep estimates, flatly disagreeing |
The lower half is TOPIK Note's set of targets; the summit is where TOPIK Note says "6,000 plus" and SEQ Toolkit says 10,000–11,000. Neither is wrong in a way you can prove — there is no official list to arbitrate. SEQ Toolkit also reports that a pre-2006 draft of about 5,800 words once circulated from the administrator and was never adopted, which is why older textbooks still quote one confident number.
What the counts understate is how concentrated the exam is. TOPIK Note ran a morpheme analysis over the reading passages of 20 past exams:
> "The 1,000 most frequent words cover 98.6% of TOPIK I reading passages."
Even the top 500 covered 86%. A Level 2 sitter who owns the frequency core has already read most of the paper's words; breadth past that earns points at a steeply declining rate. Japanese learners have the sharper version of the no-list problem — no descriptors with counts at all — covered in JLPT N5 Vocabulary: The List That Isn't Official.
Where the printed list stops helping
Search "TOPIK vocabulary list pdf" and you will find printable lists within a minute, and printing one is a reasonable first move. The paper is where you meet the words: you can see the level's shape, mark what you half-know, cross off what K-dramas deposited for free. What the paper cannot do is decide when you see a word again. A highlighted row returns on your motivation's schedule, which is to say it doesn't return. By week three of a 2,000-word list, the first 300 have quietly left.
So use the list as input instead of as the course. Paste it, or drop the pdf itself, and you do not get 1,800 rows echoed back: the model mines the strongest candidates — roughly twenty-five from a list-sized import, up to 100 when a long document gets worked through in sections — laid out as one checklist with select-all available. Every candidate is tagged high, medium, or low for how much studying it deserves, the high tier already checked, and one press of the confirm button enrolls up to 40 of them. To work deeper into the list, paste in the next stretch of it; re-importing the same text proposes the same strongest terms again. The tag is the model's priority call, not a guess at what you already know, so unchecking the words your dramas paid for stays your job. Energy converts at roughly one unit per card kept — a full Level 1 build of about 800 words runs about 800 energy across however many slices it takes. The reviews themselves are where the real work moves.
Once the cards exist, FSRS-6 schedules them (the same scheduler family Anki ships, fixed rather than configurable). Set your pace against the sitting you booked: at 15 new cards a day, Level 1's 800 is about 53 days of intake, and Level 3's ~3,000 at 20 a day is about five months. The test runs six times a year in Korea and less often abroad, so the date is knowable months out; in the final two weeks the intake valve closes, and the review queue is the whole job.
Captions, hover lookups, and the timestamp that survives
A list row pairs 사람 with "person." It does not pair it with a voice, a scene, or a sentence doing real work, and TOPIK I is half listening. Captioned video closes that gap, and here is exactly what happens in the extension on a Korean YouTube video with captions: the cues sit there one line at a time, you click one, and the model reads exactly that cue, nothing before or after it, then proposes candidates from it, each carrying the word, the meaning, an example sentence, and the pronunciation. Confirm the ones you want; this path and the file path share a tighter ceiling than document imports, one confirm enrolling at most 20 cards. The saved card keeps its source, so when it comes due in review, it replays the second of the video the line came from. The word comes back with the voice that said it, not as an orphaned pair. And a video with no caption track leaves the click nowhere to start; the extension tells you that straight instead of making something up.
Away from YouTube, the same extension doubles as a Korean dictionary: a double-click on a word, wherever you happen to be reading, opens a popup with the meaning and the part of speech, plus a save button. The dictionary is offline, so it works with no connection. The popup serves Japanese and Chinese targets as well; readings are the feature those languages add. And listening material you own can go in as a file: drop in a Korean podcast recording and it processes in Korean; drop in a clip in the wrong language and a checker rejects it (the check runs after real transcription compute, so a small detection fee applies).
One boundary worth stating at Level 3 and up: cards train recognition, and TOPIK II's writing section grades production. Recognition at speed is its own argument — the decoding-versus-retrieval pile — and it has a full write-up in Practice Reading Korean Words.
The short version
TOPIK's own numbers stop after Level 2: 800, then 1,500–2,000. Everything above is reconstruction, agreeing near 3,000–5,000 for Levels 3–4 and splitting between 6,000-plus and 10,000–11,000 at the summit. Study the level you are sitting, own its frequency core first (the top 1,000 words carry 98.6% of TOPIK I reading), let captioned video attach voices to the words, and hand the review calendar to a scheduler the day the list enters the app.
The free tier is 150 energy (about 150 cards), then prepaid top-ups at $5 / $10 / $20; no subscription, and topped-up energy never expires. To start: Piccard on the Chrome Web Store.
All guides: Piccard Blog.
Try Piccard for free
Bring your word list — paste the text or drop the PDF, and it becomes FSRS-scheduled cards. 150 free energy to start, no credit card.