What is in the data

The corpus behind every name page on this site, counted on 2026-08-29. Each figure has one definition, and all of them are on the methodology page. When two numbers seem to disagree, they are counting different things.

How big is it? That depends on what you count.

One corpus, four units. Pick a rung to see what it counts and what drops out on the way down to it.

Named-entity spellings

0

Unique spellings labelled person, organization or place.

  1. Deduplicate: billions of rows collapse into distinct (spelling, language) pairs.(−17.6B)

  2. Drop the language axis. The same spelling in twelve languages is one string.(−630.3M)

  3. Keep only spellings labelled person, organization or place; dictionary words and unknowns fall away.(−66.2M)

Bar lengths are logarithmic. The top rung is about 96× the bottom one.

191.2M distinct spellings in scope

Country by country

Spellings of people, organisations and places recorded in each country. Click a country for its languages and scripts, or pick a second one to compare.

fewermore all names

Click any country to break it down.

Languages & writing systems

Across the whole corpus.

Top languages

  • English91.1M
  • Hindi37.6M
  • Marathi36.8M
  • Tamil36.4M
  • Urdu35M
  • Bangla34.8M
  • Malayalam34.5M
  • Telugu34.5M
  • Gujarati34.5M
  • Kannada34.5M
  • Chinese21.1M
  • French18.9M
  • Spanish13.2M
  • Russian12.8M
  • Japanese9M
Show as table
Distinct spellings — distinct spellings
NameDistinct spellings
English91,108,264
Hindi37,644,381
Marathi36,843,180
Tamil36,404,267
Urdu35,046,062
Bangla34,800,900
Malayalam34,532,765
Telugu34,519,387
Gujarati34,505,380
Kannada34,504,171
Chinese21,123,024
French18,869,394
Spanish13,193,433
Russian12,842,447
Japanese8,985,977

A language here means a source tagged with that language, and the lists overlap. A name recorded in a country with several official languages is tagged with each of them. That is why Hindi, Bengali, Telugu and the other Indian languages sit so close together: they mostly share one inventory of names.

Top scripts

  • Latin137.2M
  • Han23.3M
  • Cyrillic14.9M
  • Thai3.3M
  • Arabic3.2M
  • Hangul2.8M
  • Devanagari2.6M
  • Hebrew821.4K
  • Kana550.7K
  • Japanese444.3K
  • Greek436.6K
  • Georgian398K
Show as table
Distinct spellings — distinct spellings
NameDistinct spellings
Latin137,200,796
Han23,316,077
Cyrillic14,906,490
Thai3,327,147
Arabic3,208,935
Hangul2,843,759
Devanagari2,554,944
Hebrew821,363
Kana550,716
Japanese444,265
Greek436,555
Georgian397,969

Whole corpus from here down: the filter above does not apply

Where the names come from

Most of the corpus comes from mondoDB, our own collection of public lists. Reference works, gazetteers and official name statistics add the rest.

mondoDB collection
129.2M+

Our own collection of public lists: registers, phone and business directories, company registers and web sources, each mapped to a common schema and checked before import.

Official name statistics
1.5M+

National statistics offices and population registers. Every bearer figure on a name page comes from one of these.

Gazetteers
31.7M+

GeoNames and the place data in Wikipedia and Wikidata.

Reference works
139.2M+

Wikidata, Wiktionary, name dictionaries and curated lists of titles and name words.

A spelling found in two families counts in both, so the bars overlap and do not add up to the total.

How many ways can one name be spelled in Latin?

A name in a non-Latin script can be written in Latin letters in several ways. The multiplier for each script is our estimate, so the totals are estimates too.

  • Latin: 137.2M forms
  • CJK: 23.3M forms
  • Cyrillic: 14.9M forms
  • Thai: 3.3M forms
  • Arabic: 3.2M forms
  • Hangul: 2.8M forms
  • Devanagari: 2.6M forms
  • Hebrew: 821.4K forms
  • Kana: 550.7K forms
  • Jpan: 444.3K forms
  • Greek: 436.6K forms
  • Georgian: 398K forms
  • Ethi: 316.8K forms
  • Khmr: 228.5K forms
  • Armenian: 194.2K forms
  • Hira: 163.1K forms
  • Bengali: 152.1K forms
  • Mymr: 82.2K forms
  • Tamil: 66.4K forms
  • Mlym: 50.5K forms
  • Telu: 36.9K forms
  • Other: 32.5K forms
  • Gujr: 23K forms
  • Knda: 21.9K forms
  • Laoo: 21.4K forms
  • Sinh: 17.6K forms
  • Guru: 17.2K forms
  • Orya: 13.7K forms
  • Tibt: 6.3K forms
  • Mong: 2.7K forms
  • Bopo: 1.2K forms
  • Syrc: 1.2K forms
  • Tfng: 1.1K forms
  • Thaa: 1.1K forms
  • Cans: 976 forms
  • Cher: 628 forms
  • Tglg: 622 forms
  • Nkoo: 361 forms
  • Yiii: 355 forms
  • Goth: 194 forms
  • Copt: 175 forms
  • Olck: 125 forms
  • Mtei: 72 forms
  • Java: 49 forms
  • Sylo: 47 forms
  • Xpeo: 39 forms
  • Tavt: 34 forms
  • Brah: 28 forms
  • Batk: 24 forms
  • Adlm: 18 forms
  • Aghb: 18 forms
  • Bugi: 17 forms
  • Phnx: 16 forms
  • Kthi: 15 forms
  • Tale: 13 forms
  • Sund: 13 forms
  • Egyp: 13 forms
  • Lana: 12 forms
  • Talu: 12 forms
  • Bali: 9 forms
  • Runr: 9 forms
  • Phli: 8 forms
  • Ogam: 7 forms
  • Xsux: 7 forms
  • Lisu: 6 forms
  • Vaii: 6 forms
  • Newa: 5 forms
  • Ugar: 5 forms
  • Orkh: 4 forms
  • Cham: 4 forms
  • Merc: 3 forms
  • Bamu: 3 forms
  • Lepc: 3 forms
  • Linb: 3 forms
  • Mand: 2 forms
  • Tagb: 2 forms
  • Kali: 2 forms
  • Osge: 2 forms
  • Limb: 1 forms
  • Khar: 1 forms
  • Armi: 1 forms
  • Perm: 1 forms
  • Ital: 1 forms
  • Prti: 1 forms
  • Sogd: 1 forms
  • Glag: 1 forms
  • Wara: 1 forms
  • Avst: 1 forms
  • Sind: 1 forms
  • Tang: 1 forms

Anatomy of a name

Each spelling is labelled with what it names (a person, an organisation or a place) and the part it plays, such as given name, surname or title.

Plus 75.8M ambiguous / dictionary tokens (not counted as named entities).

  • People: Given 67782016, Surname 56398042, Patronymic 4442663, Title 1228942, Particle 651595, Full name 594
  • Organizations: Org name 40191189, Org word 16558
  • Places: City 38885681, Region/Country 32716584
  • Other: Other 70966011, Common word 4688766, Profession 159165, Stopword 252

Names across languages

How names relate: the same name in another language, a variant, a feminine form, a name derived from another.

106M

Etymology links

74.2M

Cross-language links

10M

Name roots

interlingual-equivalent-of
73.3M
named-after
18.8M
variant-of
8.5M
derived-from
2.9M
transliteration-of
859.5K
cognate-of
684.2K
compound-of
320.1K
feminine-of
137.9K
hypocorism-of
137K
pseudonym-of
133.8K
borrowed-from
130K
masculine-of
120.3K
opposite-gender-of
4.3K
contracted-from
869
nickname-of
4

How the graph is built

Collect, model, collect again. Each round of models cleans the data the next round learns from.

again & again

Collect

Names come from thousands of public sources, and no two share a layout or a level of reliability. Each one is mapped to a common schema, checked and graded before it is used.

  • Public registers
  • Official name statistics
  • Company registers
  • Phone directories
  • Name dictionaries
  • Web sources