Our Data
An open evaluation of name-to-nationality inference
We evaluate Nationalize against public datasets in which each person's nationality is already recorded, reporting how often the true country lands in our top 1, 3 and 5 for every set. Each individual prediction is listed below, so the aggregate figures can be checked and reproduced.
Results
Per-dataset accuracy, and every prediction behind it
Every prediction sits next to the person's real country, so you can check the numbers yourself.
Public datasets with a recorded nationality for every person.
Where the true country landed in our ranked shortlist — MRR@5 0.6991
- Rank 1 · 60.34%
- Rank 2–3 · 18.25%
- Rank 4–5 · 6.65%
- Not in top 5 · 14.75%
More datasets added over time.
Showing 88,151–88,153 of 88,153
| Name | True country | Predicted top 5 | Result |
|---|---|---|---|
|
zlem Kaya
|
Türkiye
|
TR · 88.0%
DE · 5.3%
AT · 1.2%
NL · 0.8%
FR · 0.6%
|
Rank 1
|
|
zzet Safer
|
Türkiye
|
TR · 73.0%
FR · 3.1%
DZ · 2.5%
AT · 2.3%
BE · 1.5%
|
Rank 1
|
|
zzet nce
|
Türkiye
|
TR · 76.8%
DZ · 7.6%
FR · 1.3%
MA · 1.0%
BE · 0.8%
|
Rank 1
|
Method
How we measured this
Population. We score the "alive-today" set: people with a known birth year, at most 80 years old.
Top 1 / 3 / 5 is how often the true country lands at rank 1, within the first 3, or within the first 5 of our ranked shortlist. MRR@5 rewards ranking the right country higher, not just somewhere in the list.
Why politicians score higher. Distinctive, less-Anglicised names are often tied to a single country, so a globally diverse set can be easier to place than a Western-skewed one.
Rounding. Figures are the exact measured value to two decimals — we never round accuracy up.