Dit artikel is nog niet vertaald naar het Nederlands — je leest het origineel in het English. Ook beschikbaar in:Deutsch, English, Українська
How Spain's Surname Registry Accidentally Measures Integration
Spain's national statistics institute publishes a surname frequency table. It is a mundane object: a rank, a surname, and how many people carry it. Every country with a civil registry publishes something similar, and the file is normally used for exactly what it looks like — checking whether García really is number one.
But the Spanish table has a column almost nobody uses, and that column turns a surname list into something closer to a clock. It records, for each surname, how many generations it has been in the country. Not by estimating. By counting.
This is not a feature INE designed. It falls out of the shape of Spanish naming law, and it appears to be an accident.
The mechanism: two surnames, published separately
A Spanish person carries two surnames: the first (apellido 1º) inherited from the father, the second (apellido 2º) from the mother. Both are legal, permanent parts of the name, and both are inherited — the child of María López Ruiz and Juan García Soto is García López. Each parent passes down exactly one of their two.
INE's file, Frecuencias de apellidos (Censo, 1 January 2025), publishes three separate columns: how many residents carry the surname as first surname, how many carry it as second surname, and how many carry it in both positions. That third column exists to solve double-counting.
The first two columns do something else entirely.
Consider what has to be true for a surname to appear in the ap2 column. ap2 is the mother's surname. A surname reaches the second-surname column only when a woman carrying it has had a child registered in Spain — and that child is now themselves a resident being counted. A surname that arrived last decade can be extremely common as ap1 and near-absent as ap2, because the carriers' children are not born yet, are not numerous yet, or were born abroad.
So the ratio
ap2 / ap1
is not a demographic estimate. It is a count of a structural fact: how far a surname has propagated through the maternal line inside the Spanish registry.
What the numbers look like
Verified against the INE file directly, Censo as of 1 January 2025:
| Surname | ap1 | ap2 | ap2/ap1 | Reading |
|---|---|---|---|---|
García | 1,446,937 | 1,468,824 | 1.015 | fully established |
Mohamed | 20,645 | 19,252 | 0.933 | same — Ceuta and Melilla |
Kaur | 11,587 | 7,608 | 0.657 | second generation present |
Ivanov | 3,926 | 2,004 | 0.510 | — |
Ndiaye | 8,560 | 2,004 | 0.234 | — |
Traoré | 5,212 | 916 | 0.176 | recent arrival |
Singh | 26,813 | 1,876 | 0.070 | first generation |
The scale runs cleanly from ~1.0 down to ~0.07, and it runs in the order you would predict from the history of each community's arrival — without anyone having supplied that history.
Why the baseline is trustworthy
A ratio is only useful if you know what "settled" looks like. Here the answer is unusually clean, because the top of the Spanish registry is a tight cluster:
| Surname | ap1 | ap2 | ap2/ap1 |
|---|---|---|---|
García | 1,446,937 | 1,468,824 | 1.015 |
Rodríguez | 939,214 | 950,067 | 1.012 |
González | 930,137 | 939,512 | 1.010 |
Fernández | 896,725 | 908,388 | 1.013 |
López | 871,380 | 880,428 | 1.010 |
Martínez | 832,396 | 840,399 | 1.010 |
Sánchez | 818,287 | 828,609 | 1.013 |
Pérez | 779,313 | 794,277 | 1.019 |
Gómez | 496,272 | 500,028 | 1.008 |
Martín | 475,139 | 473,220 | 0.996 |
Jiménez | 399,394 | 400,335 | 1.002 |
Romero | 227,391 | 227,881 | 1.002 |
Twelve of the largest surnames in the country, spanning a 6× range in size, all land between 0.996 and 1.019. That is a remarkably narrow band, and it is what you would expect from a surname that has been circulating through both lines for centuries: the small excess above 1.0 is simply that slightly more of its carriers are women who became mothers than men who became fathers.
That tight band is what makes the low end interpretable. Singh at 0.070 is not a little below the baseline. It is fourteen times below it, and the baseline itself barely moves.
The Ceuta and Melilla result
Mohamed sits at 0.933 — statistically almost indistinguishable from Martín at 0.996 and far from every other surname of North African or South Asian origin in the file.
This is not a paradox. Ceuta and Melilla are Spanish cities on the African coast, and their populations have been Spanish residents for generations. A surname that has been in the registry that long has propagated through the maternal line exactly as García has. The neighbouring evidence is consistent: Abdeselam = 1.001 (ap1 2,089 / ap2 2,091), Amar = 1.079 (2,131 / 2,300), Hamed = 0.843 (2,129 / 1,794).
Meanwhile Ahmed, a name of overlapping linguistic origin but a very different residency history, sits at 0.432 (15,176 / 6,558), and Hussain at 0.087 (6,549 / 568).
The ratio is not reading the surname's language. It is reading the registry.
The second finding: a common assumption, killed by a counter
There is a widespread convention in name datasets that Kaur and Begum are female surnames. In Sikh naming this is broadly the case — Singh is carried by men, Kaur by women — and libraries that split surname lists by gender routinely encode it.
In Spain this is simply invalid, and the registry says so with a single number.
| Surname | ap1 | ap2 | ap2/ap1 |
|---|---|---|---|
Begum | 1,763 | 5,599 | 3.176 |
Bibi | 2,072 | 5,086 | 2.455 |
Devi | 351 | 397 | 1.131 |
Begum has more than three times as many second-surname carriers as first-surname carriers — the highest ratio of any substantial surname in the file, with Bibi right behind it. These are the three names most reliably tagged "female" in the wild, and they are the three names whose ratios most conspicuously exceed 1.0.
The reason is mechanical. The Spanish civil registry inherits apellido without regard to the sex of the bearer. A man whose mother is a Begum legally carries Begum as his second surname. The transmission rule does not consult the surname's original gender semantics, so within a generation the "female-only" surname is being carried by men — and because the community's mothers outnumber its registered fathers of that surname, it accumulates on the ap2 side faster than it entered on the ap1 side.
We had this wrong, and we would have kept having it wrong. The name looked female; the classification felt obvious; the linguistic reasoning behind it is correct in its country of origin. What corrected it was a counter, not an inspection.
The third finding: the sum measures something too
Add the columns up across all 86,512 surnames in the file:
| Quantity | Value |
|---|---|
Sum of ap1 | 46,541,135 |
Sum of ap2 | 43,595,570 |
| Difference | 2,945,565 |
Everyone in the registry has a first surname. Not everyone has a second one — a resident who was not born under a two-surname system has only what their foreign documents carried. The gap between the columns is therefore, to a first approximation, residents with no second surname at all: close to three million people.
The registry is not asked to count that. It counts it as a side effect of the columns being published separately.
Two honest deductions apply. This file lists surnames with ap1 ≥ 20, so it is not the entire population; and for 7,757 surnames the ap2 cell is suppressed rather than zero, which can account for at most ~150,000 of the gap. Neither is close to explaining away three million, but both mean the number is an order-of-magnitude reading, not a certified total.
The honest limits
This has to be said plainly, because the measure is easy to overclaim.
This is a proxy, not an ethnicity census. Spain does not record ethnicity, by law. The ratio does not know who anyone is, what language they speak, what passport they hold, or how they identify. It knows one thing: whether a surname has reached the mother's column.
At least four different situations produce a low ratio, and the number cannot distinguish them:
- Recent arrival — the carriers' children are not born or not counted yet.
- A community with few registered births in Spain — for demographic reasons that have nothing to do with when it arrived.
- A surname also common among long-settled Spaniards — which drags the ratio toward the baseline and hides everything.
- An origin country that already uses a two-surname system — and this one is not hypothetical.
That fourth confounder is visible in the same file and it is severe:
| Surname | ap1 | ap2 | ap2/ap1 |
|---|---|---|---|
Mamani | 3,341 | 4,042 | 1.210 |
de Oliveira | 5,142 | 5,989 | 1.165 |
da Silva | 17,518 | 20,306 | 1.159 |
Quispe | 6,392 | 6,840 | 1.070 |
By ratio alone, these look as settled as García — more settled, in fact. They are not. Brazil, Bolivia and Peru already give their citizens two surnames, so their residents arrive in Spain with an ap2 already attached. The mechanism that makes the ratio meaningful for Singh — that a second surname can only be acquired by being born here — does not hold for them at all. They import the second column wholesale.
Which means the measure is not "generations in Spain." It is generations in the Spanish registry, for surnames that entered it without a second surname. That is a narrower and duller claim, and it is the correct one.
Use it as what it is: a structural reading of a registry, not a statement about people. It cannot tell you anything about any individual, it should not be aggregated into claims about communities, and the countries that share Spain's naming law are invisible to it by construction.
What remains after all those deductions is still unusual. A civil registry, designed only to record who is called what, ends up publishing a column that counts something nobody asked it to count — and does so with a baseline so tight that the signal is legible without any modelling at all.
Data as of 2026-07-17
All surname figures were read directly from the INE file and recomputed, not quoted from secondary sources.
- Source: INE (Instituto Nacional de Estadística), Frecuencias de apellidos, Censo de población as of 01/01/2025. Columns used:
Apellido 1º / Total,Apellido 2º / Total,Ambos apellidos / Total. <ine.es/daco/daco42/nombyapel/…; - File scope: 86,512 surnames. The published caption states a cut-off of 100 for the first surname, but the file in fact runs down to
ap1= 20; the caption appears to be stale. Surnames below that cut-off are not listed. - Suppression: 7,757 rows carry
..in theap2column instead of a count. Treated as unknown, bounded above by the cut-off. - Sums:
ap1= 46,541,135;ap2= 43,595,570; difference = 2,945,565. Recomputed across all listed rows. - Accents: the INE file stores surnames unaccented (
TRAORE,GARCIA). Accented forms are used in this text for readability. - Ratios: computed as
ap2 / ap1per row, rounded to three decimals.Ivanov(2,004) andNdiaye(2,004) sharing anap2value is a coincidence in the source, not a transcription error.
Spain does not collect ethnicity data, and nothing in this article is derived from any such source. Every figure above is a count of surname positions in a civil registry.