What is in the data
The corpus behind every name page on this site, counted on 2026-08-29. Each figure has one definition, and all of them are on the methodology page. When two numbers seem to disagree, they are counting different things.
How big is it? That depends on what you count.
One corpus, four units. Pick a rung to see what it counts and what drops out on the way down to it.
Named-entity spellings
0
Unique spellings labelled person, organization or place.
Deduplicate: billions of rows collapse into distinct (spelling, language) pairs.(−17.6B)
Drop the language axis. The same spelling in twelve languages is one string.(−630.3M)
Keep only spellings labelled person, organization or place; dictionary words and unknowns fall away.(−66.2M)
Bar lengths are logarithmic. The top rung is about 96× the bottom one.
191.2M distinct spellings in scope
Country by country
Spellings of people, organisations and places recorded in each country. Click a country for its languages and scripts, or pick a second one to compare.
Click any country to break it down.
Languages & writing systems
Across the whole corpus.
Top languages
- English91.1M
- Hindi37.6M
- Marathi36.8M
- Tamil36.4M
- Urdu35M
- Bangla34.8M
- Malayalam34.5M
- Telugu34.5M
- Gujarati34.5M
- Kannada34.5M
- Chinese21.1M
- French18.9M
- Spanish13.2M
- Russian12.8M
- Japanese9M
Show as table
| Name | Distinct spellings |
|---|---|
| English | 91,108,264 |
| Hindi | 37,644,381 |
| Marathi | 36,843,180 |
| Tamil | 36,404,267 |
| Urdu | 35,046,062 |
| Bangla | 34,800,900 |
| Malayalam | 34,532,765 |
| Telugu | 34,519,387 |
| Gujarati | 34,505,380 |
| Kannada | 34,504,171 |
| Chinese | 21,123,024 |
| French | 18,869,394 |
| Spanish | 13,193,433 |
| Russian | 12,842,447 |
| Japanese | 8,985,977 |
A language here means a source tagged with that language, and the lists overlap. A name recorded in a country with several official languages is tagged with each of them. That is why Hindi, Bengali, Telugu and the other Indian languages sit so close together: they mostly share one inventory of names.
Top scripts
- Latin137.2M
- Han23.3M
- Cyrillic14.9M
- Thai3.3M
- Arabic3.2M
- Hangul2.8M
- Devanagari2.6M
- Hebrew821.4K
- Kana550.7K
- Japanese444.3K
- Greek436.6K
- Georgian398K
Show as table
| Name | Distinct spellings |
|---|---|
| Latin | 137,200,796 |
| Han | 23,316,077 |
| Cyrillic | 14,906,490 |
| Thai | 3,327,147 |
| Arabic | 3,208,935 |
| Hangul | 2,843,759 |
| Devanagari | 2,554,944 |
| Hebrew | 821,363 |
| Kana | 550,716 |
| Japanese | 444,265 |
| Greek | 436,555 |
| Georgian | 397,969 |
Whole corpus from here down: the filter above does not apply
Where the names come from
Most of the corpus comes from mondoDB, our own collection of public lists. Reference works, gazetteers and official name statistics add the rest.
Our own collection of public lists: registers, phone and business directories, company registers and web sources, each mapped to a common schema and checked before import.
National statistics offices and population registers. Every bearer figure on a name page comes from one of these.
GeoNames and the place data in Wikipedia and Wikidata.
Wikidata, Wiktionary, name dictionaries and curated lists of titles and name words.
A spelling found in two families counts in both, so the bars overlap and do not add up to the total.
How many ways can one name be spelled in Latin?
A name in a non-Latin script can be written in Latin letters in several ways. The multiplier for each script is our estimate, so the totals are estimates too.
- Latin: 137.2M forms
- CJK: 23.3M forms
- Cyrillic: 14.9M forms
- Thai: 3.3M forms
- Arabic: 3.2M forms
- Hangul: 2.8M forms
- Devanagari: 2.6M forms
- Hebrew: 821.4K forms
- Kana: 550.7K forms
- Jpan: 444.3K forms
- Greek: 436.6K forms
- Georgian: 398K forms
- Ethi: 316.8K forms
- Khmr: 228.5K forms
- Armenian: 194.2K forms
- Hira: 163.1K forms
- Bengali: 152.1K forms
- Mymr: 82.2K forms
- Tamil: 66.4K forms
- Mlym: 50.5K forms
- Telu: 36.9K forms
- Other: 32.5K forms
- Gujr: 23K forms
- Knda: 21.9K forms
- Laoo: 21.4K forms
- Sinh: 17.6K forms
- Guru: 17.2K forms
- Orya: 13.7K forms
- Tibt: 6.3K forms
- Mong: 2.7K forms
- Bopo: 1.2K forms
- Syrc: 1.2K forms
- Tfng: 1.1K forms
- Thaa: 1.1K forms
- Cans: 976 forms
- Cher: 628 forms
- Tglg: 622 forms
- Nkoo: 361 forms
- Yiii: 355 forms
- Goth: 194 forms
- Copt: 175 forms
- Olck: 125 forms
- Mtei: 72 forms
- Java: 49 forms
- Sylo: 47 forms
- Xpeo: 39 forms
- Tavt: 34 forms
- Brah: 28 forms
- Batk: 24 forms
- Adlm: 18 forms
- Aghb: 18 forms
- Bugi: 17 forms
- Phnx: 16 forms
- Kthi: 15 forms
- Tale: 13 forms
- Sund: 13 forms
- Egyp: 13 forms
- Lana: 12 forms
- Talu: 12 forms
- Bali: 9 forms
- Runr: 9 forms
- Phli: 8 forms
- Ogam: 7 forms
- Xsux: 7 forms
- Lisu: 6 forms
- Vaii: 6 forms
- Newa: 5 forms
- Ugar: 5 forms
- Orkh: 4 forms
- Cham: 4 forms
- Merc: 3 forms
- Bamu: 3 forms
- Lepc: 3 forms
- Linb: 3 forms
- Mand: 2 forms
- Tagb: 2 forms
- Kali: 2 forms
- Osge: 2 forms
- Limb: 1 forms
- Khar: 1 forms
- Armi: 1 forms
- Perm: 1 forms
- Ital: 1 forms
- Prti: 1 forms
- Sogd: 1 forms
- Glag: 1 forms
- Wara: 1 forms
- Avst: 1 forms
- Sind: 1 forms
- Tang: 1 forms
Anatomy of a name
Each spelling is labelled with what it names (a person, an organisation or a place) and the part it plays, such as given name, surname or title.
Plus 75.8M ambiguous / dictionary tokens (not counted as named entities).
- People: Given 67782016, Surname 56398042, Patronymic 4442663, Title 1228942, Particle 651595, Full name 594
- Organizations: Org name 40191189, Org word 16558
- Places: City 38885681, Region/Country 32716584
- Other: Other 70966011, Common word 4688766, Profession 159165, Stopword 252
Names across languages
How names relate: the same name in another language, a variant, a feminine form, a name derived from another.
106M
Etymology links
74.2M
Cross-language links
10M
Name roots
How the graph is built
Collect, model, collect again. Each round of models cleans the data the next round learns from.
Collect
Names come from thousands of public sources, and no two share a layout or a level of reliability. Each one is mapped to a common schema, checked and graded before it is used.
- Public registers
- Official name statistics
- Company registers
- Phone directories
- Name dictionaries
- Web sources