ABSTRACT
This work documents multigenerational persistence in economic status, showing that not only do parents influence children’s economic outcomes, but so too do grandparents and great-grandparents. Economic persistence is measured using direct grandfather–father–son links, including up to five generations, in administrative data from Norway spanning nearly 150 years (1865–2011). The findings are robust to alternative ways of measuring the characteristics of the parent generation, as well as to alternative indicators of economic status. High persistence is observed also in subsamples where grandchildren had less chance to interact directly with grandparents, suggesting an important role of unexpressed family characteristics in intergenerational transmission. The results indicate a slower occupational convergence across families over time than what is implied by parent–child associations.
I. Introduction
How persistent are economic outcomes across generations? To what extent are children’s outcomes affected not only by their parents, but also by their grandparents or great-grandparents? An understanding of such persistence is important for understanding the extent of economic mobility over time and can provide information about the relative roles of direct parental involvement and more abstract family human capital in shaping individuals’ economic opportunities. This work demonstrates persistence in economic outcomes over a period of 146 years, using a novel linked data set based on six Norwegian full-count censuses between 1865 and 2011, including 167,411 lineages with occupational data on grandfathers, fathers, and sons. This represents the largest data set on such multigenerational processes to date.
The results consistently find that grandfathers’ economic outcomes predict their grandsons’ outcomes. Even after controlling for father’s occupation, there are sizeable, statistically significant associations between grandfather and grandson across all time periods studied. The findings are robust to alternative measures, including occupational status and income rank. Moreover, measuring the social status of the parent generation in greater detail does not remove the association between grandfather and grandson, ruling out measurement error as the sole factor.
Further, rates of persistence have changed over time in Norway. Persistence is highest in white-collar occupations and among farmers. For white-collar occupations, persistence has decreased over time; the occupational association between a white-collar grandfather and his grandson in 1910 is comparable to that of a white-collar father–son pair in 2011. Despite the decline in white-collar grandfathers’ influence over time, it is always the case that having a white-collar grandfather increases the likelihood that one has a white-collar job. Multigenerational persistence among manual skilled workers has also decreased over time, while persistence has increased for farmers and unskilled workers.
The results demonstrate that the existence of multigenerational persistence does not depend on any specific set of economic institutions, as Norwegian society changed dramatically over the time period studied. In 1865, a majority of the population made their living from farming-related activity, and GDP per capita is estimated to have been only around half that of leading European countries (Bolt and van Zanden 2013). There was no state income tax, and for most of the population, only basic elementary education was available. By contrast, at the end of the study period, there is a comprehensive welfare system, education at all levels is free, and less than 1 percent of the population is engaged in farming. Yet across all time periods, grandfathers’ occupations have a strong predictive influence on grandsons’ occupations.
An exploratory analysis suggests that the influence of grandfathers is not primarily driven by interpersonal interactions. There is strong multigenerational persistence even when grandfather and grandson did not reside in the same region, or when the grandfather died at an earlier age. This runs counter to previous results from China (Zeng and Xie 2014) that attribute multigenerational persistence to direct grandparent–grandchild contact. Hence, the results presented here can be reasonably attributed to some characteristics of the family that are not manifested in observable economic outcomes.
This paper, with unique administrative data spanning nearly 150 years, complements and extends the existing literature on long-run economic persistence. Existing work is generally limited by years of data availability, small sample sizes, potential recall error, and other measurement challenges. For example, while Long and Ferrie (2013a) document changes in intergenerational mobility (father–son associations) over time, less is known about whether persistence across multiple generations has changed in a similar manner. Few studies so far have covered a long time period with consistent measurement of multigenerational transmission of economic characteristics, with some notable exceptions that use administrative data across several generations. Lindahl et al. (2015) combine a sample from the city of Malmö, Sweden in the 1930s with later administrative data and find evidence of persistence in education and income across generations. Dribe and Helgertz (2016) use data from five rural parishes in southern Sweden and observe persistence in occupational status, but not income, with no substantial changes over time. Ferrie, Massey, and Rothbaum (2016) find some evidence of multigenerational educational persistence using U.S. census data from 1910 onward, but suggest that this could be spurious due to challenges in measuring completed education precisely. Knigge (2016) uses marriage registers from five Dutch provinces and finds evidence of a moderate influence by grandfathers on occupational status that is constant over the period studied (the 19th and early 20th century).1 The present study also contributes to a broader literature on multigenerational persistence based on survey data or pseudo-linked panels.2 The results presented here indicate that there is more intergenerational persistence in economic outcomes when grandfathers’ occupations are included in addition to fathers’ occupations.
The paper is structured in the following way: Section II describes a simple theoretical framework. Section III presents the data and discusses linkage and the economic development that took place in the period covered by this study. The main analysis using four occupational categories is conducted in Section IV. Section V demonstrates that multigenerational persistence is also found when more detailed measures of occupational status are used. Section VI repeats the analysis on subsamples in which grandfather and grandson had less opportunity to interact directly and finds that there is high persistence also in this case. Finally, Section VII provides a conclusion.
II. Theoretical Framework
To fix ideas of how to interpret multigenerational persistence, a common starting point is the standard latent factor model. The latent factors are heritable across generations and influence economic outcomes. Either relationship—the heritability across generations or the relationship between the latent factor and outcomes—could change over time. This model is operationalized by Braun and Stuhler (2018) in two equations:
(1a)
(1b)
where y is the observed outcome for individual i in generation t, e is the latent factor and u,v are noise terms uncorrelated with each other and with observables. The “true” heritability of the unobserved (latent) trait is captured by λ̱, while correlations of observed traits across generations are also influenced by the relationship between the latent factor and economic outcomes, modeled by the transferability coefficient ρ.
It follows that there is a nontrivial relationship between a two-generation correlation and how differences in economic outcomes are transmitted across more than two generations. A strong two-generation correlation could reflect a strong transferability of the latent factor (high λ), a strong association between the latent ability factor and economic outcomes (high ρ) or both.
If the transmission processes did not change over time, one could in principle attempt to identify the parameters of Equation 1 separately by comparing regression coefficients from two- and three-generation regressions (see, for example, Braun and Stuhler 2018, p. 583–585). Given the large economic and societal changes throughout the 19th and 20th century, it is no surprise that previous studies have found substantial changes over time in the way transmission processes operate, making it challenging to infer parameters directly by comparing regression results.
An alternative interpretation of multigenerational persistence is that individuals are directly influenced by the presence of or social interaction with grandparents. A simple formulation of this mechanism can be given as
2
It is evident that a significantly nonzero grandparental coefficient—a positive association between grandfather and grandchild even when parental characteristics is controlled for—can indicate both a latent transferability of ability (Equations 1a–1b) or a direct transmission of characteristics from grandparents to grandchildren through personal interaction (Equation 2). This “duality”—noted both by Braun and Stuhler (2018) and earlier work such as Mare (2011)—shows that multigenerational persistence does not in itself point to one explanatory model being more “correct” than the other.
As has been the convention in the economics literature, the models specified by Equations 1 and 2 are formulated for continuous outcomes, such as income. Most existing studies on multigenerational persistence that use administrative data (for example, Lindahl et al. 2015; Ferrie, Massey, and Rothbaum 2016) use continuous status variables. However, in the present study, occupational data are primarily treated as discrete and unordered rather than continuous and/or ordered.3 Ordered rankings of occupations become harder to interpret when comparisons are made across a long period, as it is difficult to take into account changes in relative status and/or payoff over time. For this reason, I use a measure of multigenerational occupational persistence that is not dependent on any particular ordering. While approaches to multigenerational processes in unordered outcomes have been discussed in the theoretical literature (see, for example, Hodge 1966), these issues have previously not been taken into account in empirical studies of multigenerational persistence.4 Before presenting the analysis, we now turn to a description of the data and institutional context of the study.
III. Data and Economic Context
A. Sample Construction
All the data used in this study are obtained from official statistical sources. Full-count census data with information on all individuals in Norway, including occupation and location of residence, are available for the years 1865, 1900, 1910, 1960, 1970, 1980, and 2011. From 1960, all information on individuals can be linked using the Norwegian national ID number. Individual records from before 1960 are linked using name, birth time, and birth place.5 As linking women whose last names change at marriage presents major challenges, and there is little occupational information on women prior to 1960, this study will focus mainly on men and paternal lineages.
We construct samples by selecting a set of birth cohorts for the “son” population to achieve links that are as comprehensive as possible. For this to be achieved, three conditions must be fulfilled. First, there must be good father–son links within each source to connect generations together. Second, for the time periods before 1960, individuals must be linked between two different sources (census records). Third, the children, fathers, and grandfathers must be of working age in the years when economic characteristics can be observed.
The unit of observation in this study is a dynasty of three generations—son, father, and grandfather. The construction starts with the “son” observation, observed as an adult. Each son is identified when young, using either the population registry links (1960 and onward) or name and place and date/year of birth (before 1960). Then the father of that young individual is located, and his occupation used as the “father” observation. Finally, this father is again observed as a child in a third source, and his father’s occupation used as the “grandfather” observation.
The spacing of the censuses is used to group the observations into four distinct samples from four periods. Table 1 provides an overview of the size and data coverage of these samples. The earliest sample (Sample A) observes grandfathers in 1865, fathers in 1900, and sons in 1910, while the final sample (Sample D) observes grandfathers in 1960, fathers in 1980, and sons in 2011. The sizes of the samples range from 2,086 lineages in Sample A to 131,194 lineages in Sample D. For the two final samples it is also possible to add more ancestors; we return to this in Section IV.D below.
Overview of Data Set
Occupation is the only variable that is recorded throughout the period. Occupations are grouped into four major categories that are frequently used in the analysis of intergenerational mobility (Long and Ferrie 2013a; Boberg-Fazlic and Sharp 2018; Azam 2015): white-collar, farmer, manual skilled, and manual unskilled. To avoid life-cycle bias, only occupational information on individuals between age 30 and 60 is used. Because of the long time period covered with associated changes in the relative status of occupations, the baseline specification does not use imputation of status or income level by occupation.6 We return to imputation of economic status to occupations in Section V.A.
For each of the four samples, we group individuals in three generations (grandfather, father, son) into four occupational categories (white-collar, farmer, manual skilled, manual unskilled), giving a 4 × 4 × 4 matrix of occupational attainment. The number of individuals in each of the 64 cells is shown in Online Appendix Table A1.
Figure 1 gives a preview of the shape of persistence for white-collar occupations over time in Norway. The lowermost line shows the overall probability of entering a white-collar occupation for men, measured as the share of individuals (subject to some sample restrictions described below) with such occupations at each of four censuses. In 1910, 13 percent of the men in the sample held a white-collar occupation, while the share in 2011 is 59 percent. The middle line shows similar probabilities for men whose fathers also held white-collar occupations. Despite the low rates in the overall population in 1910, a majority of sons of white-collar men are able to enter white-collar occupations, reflecting limited intergenerational mobility. The uppermost line, however, shows that the importance of family background extends beyond father–son associations. If the grandfather also held a white-collar occupation, the probability in 1910 jumps from 58 percent to 77 percent and in 2011 from 74 percent to 80 percent.
Probability of Entering a White-Collar Occupation, by Family Background (Father’s Occupation)—Men in Norway, Selected Cohorts
B. Representativity of the Sample
The linkage of individuals on name, time of birth, and place of birth means that selection into the sample used in this study is not completely random. While great care has been taken to link individuals by means of time-invariant characteristics only—name, birth time, and birth place—the structure of the censuses used as base data means that only observations on families with parent–child age differences that match the observation periods can be used for linkage. Moreover, individuals whose characteristics are less unique (for example, common names, being born in large municipalities) are harder to link in a reliable manner than those with less common names from small places. This section gives an overview of how some of these challenges are handled.7
When links are made across three periods (grandfather–father and father–son) the match rates compound. In Sample B, for example, in the population for which we know the son’s occupation in 1960, in 36.1 percent of cases we have a 1910 father’s observation that satisfies all the criteria: that the son is identified in 1910, that there is a link between son and father in 1910, that the father is in the correct age interval (30–60), and that the father has a recorded occupation. Among these father–son pairs, we can identify 8.5 percent of the grandfathers in 1865 according to the same criteria, giving a final sample size of 3.5 percent of men born between 1900 and 1910. The corresponding gross match rate—the share of individuals whose occupations are observed at t for whom both fathers and grandfathers are known, with an observed occupation, and between 30 and 60 years old at t − 1 and t − 2, respectively—for Sample A is 1.4 percent, for Sample C 4.4 percent, and for Sample D 24.9 percent.
Because the selection process may change with time, we should as a rule be careful in interpreting small differences between periods as time trends. However, as the main purpose of this work is to document the robustness of the economic impact of grandfathers over and above that of fathers in different periods and for different economic variables, biases that vary in magnitude over time do not necessarily weaken the analysis if the effects of these biases occur in all of the samples.
Any biases in the samples are handled in the analysis in three ways. First, the estimates of persistence are based on odds ratios, which are invariant to changes in the marginal distribution. This means that overrepresentation of a given occupational group does not directly drive the estimation results. Second, age controls for all generations are added to all occupational regressions. Third, individuals (in the final generation) born in 1881–1899, 1911–1919, and 1951–1959 are excluded from the analysis (as shown in Table 1) because their ancestors’ year of birth fits poorly with the years in which occupations and family links can be observed.
C. Economic Development and Intergenerational Mobility in Norway
In 1865, Norway was a predominantly rural society; 40 percent of the adult male population were farmers (owners, tenants, or managers), while an additional 20 percent were cottagers with limited property rights. The oldest grandfathers in this study were born in 1815, immediately after the end of the Napoleonic wars and contemporaneous hunger and economic crisis in Norway. From 1860 to 1913 there was substantial emigration to the United States, with more than 800,000 individuals emigrating (Norway’s total population in 1865 was 1.7 million).
Norway industrialized relatively late compared with core European countries, but around the turn of the century many industrial ventures were started, often in locations dictated by the availability of hydroelectric power. In 1910, 32 percent of the working-age male population were farmers, and 31 percent list a manual skilled occupation. During this period, Norway was heavily dependent on the world economy, in terms of large-scale emigration, food imports, and raw material exports, even though many people still lived on small farms in remote areas and had to travel substantial distances to even the closest urban center.
After 1910, in which the final generation in Sample A is observed, the economic development of Norway shared several characteristics with the rest of Western Europe. While Norway was neutral during World War I, the economy was still affected, with increasing prices causing hardship for the poor and high shipping rates profiting a small group of shipowners. The first decades of the 20th century represented a period of increasing power for the labor unions, with the first stable Labor Party government being formed in 1935. The country was under German occupation during 1940–1945, though material destruction was limited except in the far North. The 1950s and 1960s saw rapid economic growth, and the number of workers in manufacturing peaked in this period. This is widely regarded as an era of equalization of opportunities, with the quality of elementary education improving. Aaberge, Atkinson, and Modalsli (2020) find that income inequality in Norway was relatively high until the late 1930s, but fell to lower levels by the early 1950s.
Semmingsen (1954) ties the emergence of the Norwegian industrial and middle classes from the 1860s onwards to the large population movements in the second half of the 19th century. The grandfathers in Samples A and B are hence observed as adults in 1865 in a predominantly agricultural society with relatively low social fluidity. While the father generation had more economic opportunities in term of industrial employment, Modalsli (2017) documents that father–son occupational mobility in Norway in 1865–1900 period was low compared to contemporary United States and 20th century Norway.
Pekkarinen, Salvanes, and Sarvimaki (2017) find that intergenerational mobility (measured by brother as well as father–son income rank correlations) increased from the 1950s onwards, with lower correlations for children born after the Second World War. This is a period of increased spending on primary education, as well as several expansions of social insurance and other social programs.
Since the start of North Sea oil production in the 1970s, economic growth in Norway has continued at a fast pace, with Norwegian GDP per capita ranked as one of the highest in the world. The labor force is increasingly concentrated in white-collar occupations. While Modalsli (2017) finds an increase in father–son occupational mobility until the 1980–2011 period, estimates based on intergenerational income elasticities (for example, Bratberg, Nilsen, and Vaage 2005; Nilsen et al. 2012) find some evidence of decreasing parent–child income mobility around the turn of the 21st century.
IV. Multigenerational Occupational Persistence
With the data structured as described in the previous section, where each observation consists of information on economic status from three generations, we now turn to the measurement of persistence using occupational categories.
A. Odds Ratios across Occupations
Transmission of discrete characteristics such as occupation groups across generations are best understood using transition matrices. Consider the parent–child transition probability matrix
3
where the parent’s occupation is held constant across rows and the child’s occupation across columns. For example, if we denote white-collar occupations by i and non-white-collar occupations by j, the variable pii states the probability that the child of a white-collar parent obtains a white-collar occupation, while pjj is the probability that the child of a non-white-collar parent obtains a non-white-collar occupation.
In this simple case of two generations with two occupational groups for each generation, we can measure intergenerational mobility using a two-way odds ratio (Agresti 2002, p. 44). Denoting the probability of the son of a father with occupation i entering an occupation j as pij, the odds ratio for father–son mobility is
4
A high value of Θ corresponds to high intergenerational persistence (that is, low intergenerational mobility), while a value of one can be interpreted as no association between father’s and son’s occupations (no intergenerational persistence and very high intergenerational mobility).8 One advantage of using odds ratios rather than simple transition probabilities is that we abstract from changes in the marginal distribution of occupations. For example, consider i as a white-collar occupation category with only 10 percent of the population. In this case, the probability that the son of a white-collar father obtains a white-collar job is 17 percent (and 83 percent that he does not, giving an odds of 0.17/0.83 = 0.20). The son of a non-white-collar father has a 9 percent chance of getting a white-collar job (and 91 percent that he does not, odds of 0.09/0.91 = 0.10). Θ is the ratio of these odds. In this case, Θ= 2(0.20/0.10), a relatively high persistence rate.
For each occupational category and time period, we can create 2 × 2 tables indicating whether fathers and sons hold the relevant occupation or not and calculate odds (probability ratios) and odds ratios (Θ). As shown in Table 2, for white-collar occupations, we obtain very high odds ratios Θ in the early period: 15.5 for Sample A and 9.7 for Sample B. An odds ratio of 9.7 means that the probability ratio (odds) of the son of a white-collar father entering a white-collar occupation is 9.7 times higher than the corresponding ratio for a son of a non-white-collar father. For Sample C (1910–1980) the odds ratio is 6.3, and for Sample D (1960–2011) it is 3.1, reflecting increased mobility. For farmers, there is no such trend towards mobility. Skilled workers have initially higher persistence but follow a similar trend to that of white-collar workers, while the trend for unskilled workers is less clear.
Odds Ratios Calculated from 2 × 2 Tables, for Four Occupational Classifications
One can also calculate odds ratios for grandfather–grandson tables in a similar manner, as shown in Table 2. These persistence measures are slightly lower than the father–son ratios.9 However, such associations do not incorporate the information from the father generation. For this reason, we now move to a framework where we can utilize information from all three generations.
B. Binary Outcomes across Three Generations
The simplest way to construct a three-generation analogy of two-generation odds ratios is to use a logit model. We choose an occupational category and set the outcome variable to one if this occupation is entered by the final generation and zero if it is not entered. The baseline approach in the present paper is then to regress this outcome against father’s characteristics Xf and grandfathers’ characteristics Xg as background variables (ι indexes the dynasty):
5
In the case where there is no grandparental information (Xg is empty), fathers’ characteristics are represented by a simple 0–1 dummy for occupational category, and there are no age controls for the son generation, the estimates of β from Equation 5 are equivalent to the log of the odds ratios from the 2 × 2 table.
The setup in Equation 5 will form the basis of the analysis of occupational persistence in this paper. Isolating persistence in this way is important because most countries have experienced major changes in both the distribution and the income rank of occupations over time. In particular, the number of farmers, a heterogeneous group that cannot always be ranked reliably relative to nonfarming occupations, was very high in many countries in the late 19th and early 20th century, and measures based exclusively on incomes or ranked occupations could hence give misleading results when applied across such a long time range.
The simplest joint model of fathers and grandfathers uses a dummy variable D for each generation that is equal to one if that generation holds the occupation in question. This will be the baseline specification. Age controls will also be used throughout. Hence, we have the following expression for the covariate vector for generation q ∈(f,g) in Equation 5:
6
that is, the model
7
is estimated four times, with the indicator variable D as white-collar, farmer, manual skilled, and manual unskilled, respectively. Table 3 gives the resulting parameter estimates. In line with the common terminology of changes in economic characteristics across generations, persistence and mobility will be taken as opposites: high mobility equals low persistence and vice versa.10
Odds Ratio Coefficients for Binary Occupational Regressions on Father’s and Grandfather’s Occupations
We start with the outcome of the son entering a white-collar occupation, as opposed to entering an occupation in one of the other three categories. The corresponding right-hand-side variables are a dummy variable for whether the father had a white-collar occupation, a dummy for whether the grandfather had a white-collar occupation, and second-degree polynomials controlling for the age (at the time of observation) of each of the three generations.11 The top panel of Table 3 shows the exponentiated coefficients for father’s and grandfather’s occupations and can be interpreted as odds ratios.
The coefficient on father’s occupation—11.8—is similar in magnitude to that in the two-generation case reported in the previous section. We now focus on the coefficient on grandfather’s occupation. For an individual observed in 1910 (Sample A) with a given father’s occupation, having a grandfather with a white-collar occupation increases the odds of entering a white-collar occupation by 2.8. In other words, the grandfather effect in the 1865–1910 sample is comparable in size to the father effect in the 1960–2011 sample. The grandparental effect in Sample B (sons observed in 1960) is also large, while the grandparental coefficients in Samples C and D (sons observed 1980 and 2011) are lower, in accordance with the generally higher mobility into and out of white-collar occupations. However, all coefficients are significant and substantial.
The second panel of Table 3 reports father and grandfather coefficients for farmers, where again the variables of interest are dummies for whether the father and grandfather belonged to the same occupational category. Once again, there are large and highly significant coefficients. For both white-collar occupations and farmers, the grandparental coefficient in all four samples is significant at the 1 percent level. The grandfather odds ratio for farmers is moderate in the two earliest periods (Samples A and B) but substantially higher for the later samples (Samples C and D), with a high of 3.9 for the final period.
While white-collar persistence is higher in early periods, and farmer persistence is higher in later periods, we have a qualitatively similar finding—that grandfathers matter. Their influence declines over time but is still apparent even in the most recent generation. Knowing that an individual had a father in the given occupation group substantially increases the probability that the individual himself holds the occupation. Controlling for father’s occupation, knowing that the individual had a grandfather in this category also increases the probability substantially, but less so than for the father. This is strong evidence that the multigenerational process in these occupations is more persistent than one would expect from a simple father–son regression.12
The two remaining occupation groups, manual skilled and manual unskilled occupations, exhibit larger differences in regression results across samples. For manual skilled occupations, the picture is similar to that of white-collar occupations in Samples A and B, with statistically significant coefficients of substantial magnitude. In Sample C, the grandparental coefficient is negative; this means that when we compare two individuals with the same father’s occupation, the one whose grandfather did not have a manual skilled occupation would have a higher probability of entering such an occupation (though this difference is not statistically significant). Examining the full matrix of occupations shows that this negative coefficient is a result of the relative probabilities for grandsons of manual unskilled individuals, who are more likely to enter manual skilled occupations than any other group. For manual unskilled occupations, the coefficient on grandfathers is significant in Samples B–D, while it is close to zero and insignificant in the initial period.13
Just as the magnitude of the observed associations differs depending on which occupational category one uses to split the population, the evolution of multigenerational persistence over time also differs across occupation categories. While there is a clear evolution towards lower persistence for white-collar and manual skilled occupations, persistence has increased for farmers and manual unskilled occupations. One way of interpreting these differences is that decreasing father–son persistence (increasing mobility) occurs together with lower persistence also for earlier ancestor generations.
The significance of grandfather’s occupation does not depend on the specific grouping of occupations used here. Online Appendix Table A2 shows coefficients from estimations on smaller occupational groups. In the case of specific occupations with a limited number of individuals, some of the cells in a 2 × 2 × 2 transition matrix will frequently not be fully occupied—in these cases, coefficients cannot be estimated. However, where there are a sufficient number of observations, the pattern for the detailed sample occupations is similar to that in Table 3, though slightly higher on average (as would be expected from more precise categories).
The way persistence is modeled here—considering one occupational category at a time rather than all four categories jointly—could potentially hide some characteristics of multigenerational persistence that only emerge when the full 4 × 4 × 4 matrix is considered. For example, the probability of a son entering a white-collar occupation could differ depending on whether the father had a manual skilled or manual unskilled occupation. However, careful consideration of the odds ratios from such interactions reveals no substantial interactions that change the interpretations above. Including them in a measure of average odds ratios (the Altham statistic) confirms a picture of reduced grandparental influence over time. Online Appendix Section IX.A presents this in detail.
Finally, the categories used here are based on the characteristics of occupations and may therefore change in social status over time. This could be a concern for white-collar occupations, which today encompass a broader segment of society than 150 years ago, as the size of the occupation group has increased. For this reason, in Section V we examine status-based measures of multigenerational persistence and show that splitting the sample by high- and low-status occupations yields a similar result to that found for white-collar occupations. On the other hand, the changing role of persistence in farming over time is possibly better understood with respect to structural change and the decreasing number of farms. Though the first-order decrease in occupation size is implicitly controlled for in the construction of the odds ratio, it could still be the case that the remaining population of farmers is in some way “selected,” so that persistence is higher among the families that have remained in this occupational category.
C. Record Matching and the Measurement of Persistence
To assess to what extent the results obtained here can be regarded as valid for the entire population, we can compare father–son intergenerational mobility for the subsample whose grandfather’s identity is unknown with those for whom a father–grandfather link is successfully obtained. The difference between odds ratios when the same controls for age are imposed on both samples is given in Table 4. The table reports the regression results for a dummy variable for whether grandparent’s occupation is available, interacted with father’s occupation in a regression where the outcome variable is son’s occupation. This could be interpreted as the ratio of odds ratios calculated on the matched versus the unmatched sample.
Interaction Effects between Father’s Occupation and Whether Father–Son Pair Can Be Linked to a Grandfather
The coefficients are exponentiated; a number larger than one means that father–son pairs for which the grandfather is known (“matched dynasties”) have higher measured persistence than father–son pairs with an unknown grandfather (“unmatched dynasties”). For example, in Sample A, the matched dynasties have 0.2 percent higher persistence than unmatched dynasties, while in Sample B, the matched dynasties have 22.9 percent higher persistence.
In general, these differences are small, indicating that the findings are not driven by artifacts of the matching process. However, in particular for farmers and manual unskilled workers in the first sample there are larger differences, with matched dynasties experiencing lower persistence than unmatched dynasties. The results for these groups in these time periods should therefore be interpreted with some caution.
An alternative approach to this direct comparison of samples is to weight observations by the match rate of subgroups, as suggested by Bailey, Cole, and Massey (2020). Here, we follow this approach, as implemented in Ferrie, Massey, and Rothbaum (2020), and regress the probability of being in the sample on a set of covariates, using this to predict a matching probability at the individual level. This information is then used in an inverse probability weighting framework to rerun the persistence regressions on the weighted sample. Online Appendix Table A7 gives the full table of weighted and unweighted regression results. For three of the samples (Samples A, C, and D), there are only minor differences between the weighted and unweighted analysis. For Sample B, the differences are larger, for both father and grandfather coefficients. However, there does not appear to be a systematic bias, as coefficients are in some case higher (for example, for father and grandfather farmers) and in other cases lower (for example, for manual unskilled grandfathers). The large deviations for Sample B occur because a few cells (combinations of occupations across generations) have very few observations, and these cells are given very high weight and undue influence in the weighted regressions.14
D. What about the Great-Grandfathers?
So far, we have examined the influence of two generations of ancestors on the outcome of the final generation. The data allow examination of the influence of an even larger set of generations.
Two caveats must be kept in mind when conducting such an analysis. First, as match rates are imperfect, the reduction of the sample size is compounded when more generations are considered and the observation years are irregularly spaced. Second, the number of ancestors increases geometrically with the number of generations, and we only consider paternal ancestors here.
Table 5 shows the results of logit regression with son’s occupation as the outcome for models with further ancestor generations for each of the four occupational groups. In Sample C, information on the great-grandfather is available; in Sample D, we have information on both great-grandfathers and great-great-grandfathers.
Odds Ratio Coefficients for Binary Occupational Regressions on Four and Five Generations
The first column of Table 5 shows the results of the four-generation models, with the paternal lineage observed in 1865, 1910, 1960, and 1980, respectively. For white-collar occupations, great-grandfathers have substantial predictive power, with estimated 53 percent higher odds of entering a white-collar occupation for the great-grandson of a white-collar worker conditional on father’s and grandfather’s occupations. There are substantial coefficient values also for farmers and unskilled manual workers, though these are not statistically significant. Online Appendix Table A6 compares three-generation regressions for the baseline sample and for the subsample where the fourth generation is available. When the number of observations is reduced from 28,091 (with three generations) to 2,422 (with four generations), several of the grandparental coefficients lose significance.
The second column of Table 5 shows the same model as the first, with estimated coefficients for four generations, but in a later time period (measured in 1910, 1960, 1980, and 2011). We again observe a statistically significant positive coefficient for the great-grandfather for white-collar occupations, and now also for farmers. The third column introduces a fifth generation (great-great-grandfathers). This reduces the sample size substantially, and no great-great-grandfather coefficients are significantly different from zero. For manual skilled occupations, we observe negative coefficients for great-grandparents in all three models (statistically significant in the second model). For given father’s and grandfather’s occupations, an individual would have a lower probability of entering a manual skilled occupation if his grandfather was a manual skilled worker. This is again driven by the higher probability of descendants of manual unskilled workers entering manual skilled occupations.
V. Persistence and the Measurement of Economic Characteristics
While the analysis above shows positive and significant coefficients for grandfathers in all time periods and for most occupational categories, one might fear that this effect was driven by changes in the status of occupations, or by otherwise coarse measurement of the father characteristics. This section shows that the results are robust to alternative specifications. Data on incomes are used to impute status to occupations and create categories that hold status constant over time. We also establish that the grandparent coefficient remains significant and of a similar magnitude if more detailed occupational information on the parent generation is added.
A. Persistence in High- and Low-Ranked Occupations
The analysis presented in the previous section maintained a four-way separation of occupation groups (white-collar, farmer, manual skilled, and unskilled). While the use of odds ratios allows study of changing persistence that is not mechanically driven by changes in the share of the population in each occupation category, a comparison between time periods is still complicated by these changes. Entry into white-collar occupations is much more open in the 21st than in the 19th century, and 21st century white-collar occupations also employ a much larger share of the population.
By shifting the focus from occupation to income, one can get more information about whether to think of the grandparental influence as horizontal movements across fields, or as vertical movements between different levels of economic status. While there is no individual-level income information available prior to the 1960s, occupational income averages can be used to impute the status of each occupation.
Average occupational incomes are obtained from official statistics. For 1980 and 2011, direct individual-level links between occupations and tax records are used to calculate occupation mean incomes. For 1960, a similar method is used, but using incomes from 1967 (the earliest available year). There are 255 distinct occupational categories in 1960, 379 in 1980 and 202 in 2011. For 1910 and 1900, data from Statistics Norway (1915) was used (based on contemporary data links between the 1910 census and tax records from the same year), giving 107 and 103 occupational categories with distinct mean incomes. For 1865, income tabulations from Norwegian Department of Justice (1871) provide the basis for a comparable estimate with 73 categories, though with a slightly different population definition (incomes calculated on the basis of men aged 25 and above, compared to men aged 30–60 for other years).
In each year, occupations are ranked by mean income. Some occupation categories (in particular, farmers) are very large in some time periods. For this reason, it is not feasible to study arbitrary percentile cutoffs. However, for three levels there is sufficient granularity to construct reliable indicator variables for social status. At the top of the occupational status distribution, an indicator variable for being in the top 11 percent separates the highest-status occupations from the rest. Adding some other high-status occupation groups allows for an indicator variable for the top 22 percent. At the lower end of the distribution, an indicator variable is constructed for the 22 percent of the population in the lowest-paid occupations.15
With these constant status groupings, we maintain the setup of Equation 7 and run logit regressions for each time period. The outcome variable is an indicator variable for the son’s occupational status (top 11, top 22, or bottom 22 percent), while the right-hand-side variables are similar indicator variables for fathers and grandfathers, plus age controls. Figure 2 plots the coefficient on grandfather’s occupational status for the four time periods. The full results are reported in Online Appendix Table A9.
Changes in Persistence over Time for High- and Low-Status Occupations
Notes: Figure shows exponentiated coefficient on grandfathers, with separate regressions for each time period and occupation classification. (95 percent confidence intervals.)
The left panel compares the grandfather coefficients on the high-ranked occupations to the previously reported coefficient on white-collar occupations. There is a clear downward trend (towards lower persistence) over time. Having a grandparent in an occupation category that pays at the 11th percentile or higher increases the odds of ending up in a similar category by 2.4 in 1910, 2.1 in 1960, 1.9 in 1980, and 1.5 in 2011. If we instead consider the top 22 percent, the decrease is larger in the first period, but the developments in the three different ways of partitioning the population are otherwise remarkably similar.
The right panel of Figure 2 shows results for lower-status occupations. There is an increase in the grandparental coefficient for the lowest-paid occupations (the bottom 22 percent) over time, consistent with increased persistence for movement into and out of the category of manual unskilled occupations in general. These results show that the estimated development in multigenerational persistence over time described in Section IV are not artifacts of changing sizes of occupational categories.
For the two final time periods the analysis of occupational status can be supplemented with results for individual incomes obtained from tax data. This analysis is presented in the Online Appendix (Section IX.B). Linear regressions of income ranks of three generations provide comparable results to what we learn from occupational status.
B. More Detailed Observation of the Parent Generation
A relatively coarse occupation specification, with few categories, leaves open the possibility that including grandparent information simply gives us more precise information on the father’s background. This could be the case, for example, because a high-status grandfather makes it more likely for the father to hold a high-prestige occupation within his occupational category. The reliance on paternal lines (because of limitations in the historical data) also means that the grandfather coefficient could capture the results of nonrandom marital matching, where the grandfather association is confounded with human capital transfers from the mother. This subsection shows that this is not likely to be the case.
If the results seen so far are mainly driven by imperfect measurement of the father’s generation, we would expect the grandfather coefficient to be substantially reduced when more information on the middle generation is added. To investigate whether this is the case, we replace the single dummy in Equation 5 with a more detailed specification of father’s occupation (two-digit HISCO for 1900–1910, two-digit NYK for 1960–1980). That is, while we maintain the grandfather specification
from Equation 6, we replace that of the father with
8
Table 6 reports coefficient estimates using a range of models that incorporate different amounts of information on the parent generation. All columns report exponentiated coefficients from a logit regression where the outcome is whether the son has a white-collar occupation, similar to the first panel of Table 3. Model 1 includes only a dummy for whether the grandfather has a white-collar occupation, while Model 2 is simply a repetition of the baseline results from Section IV, where there is also a dummy for whether the father has a white-collar occupation. Comparing the two first columns, it is clear that not controlling for father’s characteristics at all (Model 1) clearly ascribes too much of the family background influence to the grandfather.
Coefficients from Regressions with More Detailed Information on the Parent Generation
In the third column, the characterization of father’s occupation is extended to the full set of occupational categories used in the census data (covariates from Equation 8). This additional information increases the predictive power of the model as described by the χ2 likelihood ratio tests and the pseudo-R2.16 There is only a moderate change in the coefficient for grandfather’s occupation, which in all models remains a binary variable indicating whether the grandfather has a white-collar occupation. This provides one indication that the grandfather coefficient is not driven by imprecise measurement of the father, but we can go further.
In Model 4, we return to the binary occupational variable (Equation 6) but include a dummy variable for mother’s occupation. Model 5 reports estimates using the full set of dummies (Equation 8) for both mother and father.17 In all four time periods, the grandparent coefficient remains significant and robust to the improved measurement of parent characteristics. While there is a slight decrease in magnitude, it is small compared to the overall effect. For this reason, we conclude that the grandfather effect is indeed a reflection of latent family characteristics rather than mismeasurement of the parent’s occupation.
For the third and fourth time period, we can also use data on individual incomes (from the tax records; see Online Appendix Section IX.B). Column 6 shows the results of an estimation using individual income ranks for fathers and mothers. Again, we observe that the addition of information on the parent generation does not substantially alter the estimated grandfather coefficient.
The final column reports, for reasons of comparison, results from a regression where no controls for grandfather’s occupation are included. We observe that while the coefficient of father’s occupation is somewhat mediated (from 14.96 to 11.79 in the first period and less in later periods), the magnitude remains similar, giving another indication that grandfather’s effects do not work solely through father’s observable characteristics.
In order to compare the explanatory power of the model with and without grandfathers, a starting point would be to consider the pseudo-R2 of Models 7 (fathers only) and 2 (fathers and grandfathers) in Table 6. The explanatory power of the grandfather–father model is in all cases greater than that of the father-only model. This carries over to a larger set of measures of fit typically used in evaluating logit models; for nine different ways of measuring explained variability and/or fit improvement, the model with grandfathers and fathers exhibits a higher score than the model with fathers only.18
A similar exercise can be performed for the other three occupation categories. These results are reported in Online Appendix Section VIII.B. In general, any significant coefficients in Table 3 remain significant when these more complex specifications are used, and the explanatory power of the model increases.19 Taken together, these results suggest that the grandfather effects described in the previous sections are not driven primarily by imperfect measurement of the characteristics of the parent generation.
We now turn to a further distinction between possible mechanisms driving the influence of grandfathers by using information on whether there was likely to have been direct contact between the generations.
VI. The Role of Personal Contact
The analysis so far has established that there is an association between grandfathers’ and grandsons’ social status, as measured by both occupation and income, when controlling for father’s status. However, we do not know whether this reflects underlying family characteristics, which even with perfect measurement would only be partly reflected in the observed status of the father, or the direct influence of the grandfather’s presence during the grandson’s childhood.
Zeng and Xie (2014) find that the relationship between grandparents and grandchildren’s economic outcomes in rural China reflects direct interaction between generations. In their study, the effects of grandparental characteristics on grandchildren’s education are strong for coresident grandparents but nearly nonexistent for noncoresident grandparents. Does the same hold in Norway? A supplementary approach to using geographical moves to infer the direct influence of grandparents on their grandchildren is to use information on the grandparents’ time of death.
This section explores the mechanisms behind the grandparental association by comparing multigenerational persistence between lineages with varying geographical distance between grandparents and grandchildren and between lineages with different times of death of the grandfather. These factors are not completely orthogonal to the process of multigenerational transmission of occupations or income. Geographical mobility is correlated with occupational mobility, and mortality is higher in groups that are less economically advantageous. We therefore have to keep potential confounding factors in mind while performing this analysis. However, using two separate indicators of grandparental presence provides a richer picture of how multigenerational persistence mechanisms operate.
We start by comparing the grandparental persistence parameters previously reported in Table 3 in cases where individuals grew up in the same municipality as their grandparents with cases where they did not.20 The relevant variables are the residential location of the grandfather when he was observed as an adult (time t) and the residential location of the grandson when he was observed as a child (time t + 1). The assumption is that a geographical move sometime between these two observations will reduce the direct interaction between grandfather and grandson. For example, for a son in Sample C who resided in Oslo municipality in 1960, we have a “nonmover dynasty” if the grandfather resided in Oslo or Aker municipalities in 1910 (the two municipalities were merged in the intervening period) and a “mover dynasty” otherwise. We then rerun the regressions from Section IV.B on the two subsamples and compare the grandparental coefficients.
We expect less influence between generations for mover dynasties, as there is less scope for interpersonal contact. However, this loss of contact is much more likely to take place between grandfather and grandson than between father and son, as movements of young sons would likely have taken place together with the father, but not necessarily the grandfather. Hence, the coefficient on grandfather’s occupation is expected to be higher (stronger persistence) for the nonmover sample than for the mover sample.
Table 7 reports these coefficients on grandfather’s occupation for the mover and nonmover subsamples, with one column for each time period. The coefficients are obtained by interacting a dummy variable with the full set of controls in Equation 5. The first panel reports coefficients using white-collar occupation for son as outcome variable. We observe that grandparent’s occupation is statistically significant for both subgroups in all time periods except for the first. The coefficient is lower for movers than for nonmovers, but the difference is only of moderate size. In Samples C and D, the difference between the subgroups is statistically significant. For nonmovers in Sample C (grandfathers in 1910, fathers in 1960, sons in 1980), we have a grandfather odds ratio (holding father’s occupation constant) of 1.87. For movers, the odds ratio is 1.49. The ratio of these numbers is 1.26. Since the odds ratios correspond to exponentiated coefficients, the ratio is equivalent to the difference between the nonexponentiated coefficients, and the t-statistic of 2.6 can be interpreted as a t test, here rejecting a null hypothesis that persistence is equal across the two subgroups.21
Comparison of Grandparental Persistence for Movers and Nonmovers
We observe a slightly different picture for the other occupation groups. For manual occupations, the difference between movers and nonmovers is not statistically significant. For farmers, on the other hand, there is higher persistence for nonmovers than for movers. This is not surprising, as the farmer occupation is usually connected to a specific farm location. If the father had already moved away from the area of the ancestral farm when the child was born, the tie to the farming occupation is likely to have been substantially weakened.22
The extent to which grandfathers participate in the upbringing of their grandchildren may also depend on how far apart they live during the grandchild’s formative years, even if they do not reside in the exact same municipality. An analysis of the relationship of the geographical distance between grandchild (growing up) and grandfather (measured in kilometers) shows no systematic relationship across occupations beyond that depicted in Table 7.
Next, we consider time of grandfather’s death and how it relates to occupational persistence across generations. Digitized death records are only available from 1960 onwards, limiting this analysis to Sample D (1960–1980–2011). We do a subsample analysis similar to the one in the previous subsection. The sample is split on whether the grandfather (whose occupation was recorded in 1960) is still alive in 1980, when the father’s occupation is recorded. In no cases are the grandparent coefficients statistically significant, indicating no difference in persistence by whether the grandfather was alive to directly influence the grandson. The standard errors for white-collar, manual skilled, and manual unskilled occupations are low, while the zero for farmers is less exact. This result holds up if the sample is instead split according to whether the grandfather survived until 2011 or if the full set of dummies for mother’s and father’s occupations is used.23 Similarly, a regression using the number of years in childhood when the grandfather was alive (see Online Appendix Section XI.C) does not give any indication of a strong relationship.
Braun and Stuhler (2018) and Dribe and Helgertz (2016) previously investigated the relationship between grandfathers’ death dates and multigenerational persistence. In these cases, because of small sample sizes, the authors did not attach much weight to the absence of significant effects. However, the relatively exact zeros observed in the present case can be seen as weakening the hypothesis that social interaction with grandfathers contributes substantially to their grandchildren’s occupational choice.
VII. Concluding Comments
This work has shown that grandfathers’ characteristics matter for their grandsons’ occupational outcomes. These results hold across a wide range of historical and institutional settings. The association is robust to several different ways of measuring economic characteristics, remains significant when we measure the status of parents in a more detailed way, and is found across subsamples split by geographical movement or grandparental mortality.
Multigenerational occupational persistence is observed not only for white-collar occupations, but also for farmers and for skilled and unskilled manual workers. The results are clear for white-collar occupations. Persistence is high across generations, for all generations studied spanning grandfathers born between 1805 and 1930 and grandsons born between 1870 and 1981. The relationship holds when controlling for father’s occupation. There is also evidence that great-grandfathers have an influence. The magnitude of the coefficient on grandfather’s occupation is up to one-third of the coefficient on father’s occupation. This is high, given that only one grandfather is observed, as data on the maternal grandfather are not available in this study.
High persistence is confirmed by the use of income rank data in the two final samples. For farmers and for manual occupations (skilled and unskilled), the results are more nuanced. There are some periods where a significant grandfather coefficient is not obtained. Skilled workers appear to show high persistence in the early period in particular, while results for unskilled workers are more pronounced in later samples.
There are differences across occupational groups. White-collar persistence is evident in all periods, but it is particularly strong in the late 19th century. It appears to decrease throughout the 20th century. The high association parameters observed for white-collar occupations hint at some sort of “elite persistence.” This is also reflected in analyses where a narrower definition of high-status occupations is used (the best-paid 11 percent or 22 percent of occupations in any given time period). On the other hand, unskilled and low-status occupations show a tendency towards increasing persistence over time.
Farmers always exhibit strong persistence. While this is not surprising, it should be kept in mind in any study of historical intergenerational mobility. In most countries outside the most industrialized part of Western Europe, a substantial proportion of the population were farmers well into the 20th century, and studies of mobility that rely on imputed status or income for this period are likely to be affected strongly by trends in mobility into and out of farming.
Families staying in the same place across generations show slightly higher multigenerational persistence. However, there is substantial persistence also in dynasties that move between observations. The difference allows for some effect due to direct personal interaction between grandfather and grandson but may also reflect shared personal networks or region-specific competencies. The timing of grandfather’s death does not influence the degree of multigenerational persistence, suggesting that the role of personal interaction is limited.
This work documents that there are substantial differences in the strength of persistence over time, from the 19th through to the 21st century. During this time, Norway grew wealthier, and education, health, and social insurance policies were substantially expanded. At the same time, persistence in white-collar and manual skilled occupations fell substantially. Hence, while persistence is found in all time periods, there still appears to be a role for institutions in the transmission of social status across generations, as we observe substantial changes in the level of measured persistence over time. Moreover, the differences across occupational groups suggest that both vertical (status) and horizontal (sector, notably agricultural/nonagricultural) must be taken into account. These explanations do not rule out unobservable inherited characteristics as an important channel of influence. The changes over time do not, however, support the argument by Clark (2014) that such latent characteristics are strong and virtually unchanged over time, regardless of the institutional framework.
Nonmanifested genetic traits (inheritance) may explain some of the dynasty persistence. However, traits may well be manifested in the father and yet not be reflected in his choice of occupation. In most cases, people cannot find a job in which all their skills are useful, and the son of a carpenter may well find joy in carpentry (and pass this on to his son through social interaction) even though he ends up working in another occupation. The social network of the family may also reflect the ancestors’ economic life and affect the occupational choice and success of the grandson. For some occupations, even outside farming, direct inheritance and family wealth may play a role as well. What we certainly learn is that persistence in economic outcomes spans more than two generations.
Footnotes
The author thanks Rolf Aaberge, Lars Kirkebøen, Andreas Kotsadam, Mikael Lindahl, Kjetil Telle, anonymous referees, as well as participants at workshops and conferences for helpful comments and discussions. Support from the Norwegian Research Council is acknowledged. The data files of individual records from the census files, tax registry, and population registry collected less than 100 years ago can only be obtained by authorized researchers according to the guidelines described in http://www.ssb.no/en/omssb/tjenester-ogverktoy/data-til-forskning. There are legal limitations (the Norwegian Statistics Act) on making these data publicly available. However, summary tables of occupational associations are presented in the Online Appendix, and programs and do-files used in writing the paper can be made available upon request. The use of individual administrative records in this project has been approved by Statistics Norway’s Data Protection Officer (personvernombudet) in accordance with ethical and legal regulations.
Supplementary materials are freely available online at: http://uwpress.wisc.edu/journals/journals/jhr-supplementary.html
↵1. When data are drawn from a limited geographical region, results could be biased, as those who migrate into (in the first and third case above) and out of the region (second and third case) are not covered. Moreover, smaller regions could have particular characteristics in terms of industrial structure or demography that limits the way results can be generalized into countries as a whole. This is less of a problem in the U.S. study, which is presumably representative of the nation as a whole (but has low match rates). In any case, the present paper is the first to use data drawn from an entire country and covering three centuries (a measurement span of 100 years for each generation).
↵2. Solon (2018) gives a partial overview of the literature. Examples of survey-based studies are Chan and Boliver (2013) and Hertel and Groh-Samberg (2014), who find some grandparental effects on social class; Warren and Hauser (1997), who use several composite outcomes but find no evidence; and Zeng and Xie (2014), Braun and Stuhler (2018), and Kroeger and Thompson (2015), who find some evidence of persistence in education. Only the latter two have some coverage of the time dimension in that they use more than one survey from the same country. Lindahl et al. (2015) is also partly based on survey data, augmented with administrative registers. Online Appendix Section VIII. A provides a brief discussion on the relationship between survey-based and register-based studies of multigenerational persistence. Studies based on pseudo-links use administrative data, but without direct linkage between individuals. Rather, information on the joint distribution of names and economic outcomes is utilized (Clark and Cummins 2015; Clark 2014; Olivetti, Paserman, and Salisbury 2018; Guell, Mora, and Telmer 2015). Interpretation of these results is sensitive to the distributions of surnames or estimation of specific parameters, making comparisons with conventional measures of intergenerational mobility/persistence challenging.
↵3. Long and Ferrie (2018) impute incomes for five occupational groups and use ordinary least squares regressions (some categorical matrix comparison is done, but no interpretation of the magnitudes is provided), while Dribe and Helgertz (2016) use an ordered logit model based on status rankings. Knigge (2016) uses a different approach based on occupational status in multilevel models.
↵4. The study of two-generation (father–son) mobility by Long and Ferrie (2013a) takes into account changes in occupational distribution, but the main results present an aggregation of odds ratios (the Altham statistic, see also Online Appendix Section IX.A) that is not as easy to interpret. The discussion of how to take changing marginal distributions into account was pursued further in published comments to Long and Ferrie (2013a) (Xie and Killewald 2013; Hout and Guest 2013; Long and Ferrie 2013b).
↵5. The use of linked historical microdata in economic studies has become more widespread in the last decade; two recent working papers (Bailey et al. 2019; Abramitzky et al. 2020) give overviews. The present approach has most in common with what Bailey et al. denote “Method D,” where potential multiple matches are disambiguated based on a scoring system that weighs differences in characteristics against each other. Moreover, the method used here takes into account possible conflicting matches in both data sets, in principle evaluating dissimilarity scores for all combinations of observations from Period 1 and Period 2 (cutoffs in the algorithm limit the number of calculations that have to be done in practice). In their review of recent practices, Abramitzky et al. (2020) express a high level of confidence in the use of automatic linkage methods to compare economic mobility across different countries. In some ways, the Norwegian data differ from the U.S. data that is typically discussed in overviews of linking methods (such as the two papers referred above). Most importantly, birthplaces in Norway are reported on the municipality level, as opposed to the state level as in the United States, giving a much higher resolution of an important characteristic used in matching individuals. Online Appendix Section VIII.A documents the linkage procedure in further detail.
↵6. Censuses before 1960 do not list education, and income data are available electronically only from 1967 onward.
↵7. The relationship between the way records are matched and the results on multigenerational persistence will be further discussed in Section IV.C.
↵8. Values of less than one indicate that families within the occupation group are less likely to transition into the group than individuals outside. Values substantially below one are not often seen empirically, and such outcomes (which are analogous to negative intergenerational income elasticities) are rarely discussed in the intergenerational mobility literature.
↵9. Associations with standard errors and controls for age are given in Columns 1 (grandfathers) and 7 (fathers) in Tables 6 and Online Appendix Tables A3–A5.
↵10. In some cases, interpreting logit coefficients across samples can be misleading (Mood 2010). However, because the marginal distributions of occupations change over time, some normalization of marginal distributions is necessary (see the works cited in Footnote 4). Moreover, as argued by Buis (2016), there are some applications (such as the importance of social backgrounds) where direct interpretation is appropriate. For this reason, the logit model will be used as a baseline specification. Results using linear probability models are available upon request.
↵11. Age controls do not matter greatly for the results.
↵12. Section V.B presents an explicit comparison of the models, including measures of model fit.
↵13. One could also add an interaction term between grandfather’s and father’s occupation. The coefficient on such an interaction term is in most cases close to one (no effect), insignificant, and does not substantially alter the coefficient reported for grandfather’s occupation here.
↵14. For the final sample, results have also been replicated on subgroups differentiated by age. This does not substantially alter the results; see Online Appendix Section VIII.E.
↵15. Because of the limited number of occupation categories, the size of the highest-ranked category ranges from 10.1 percent (in 1960) to 12.3 percent (in 1865); the wider highest-ranked category from 20.1 percent (in 2011) to 23.9 percent (in 1910); and the lowest-ranked category from 21.9 percent (in 1900) to 23.0 percent (in 1980). The population shares are calculated on the full sample of men 30–60 with observed occupations.
↵16. The increase in predictive power is substantial in absolute terms; in relative terms, the increase is only moderate, as some of the explanatory power in all Models 1–6 is attributable to the age controls.
↵17. The number of observations is lower for Models 3–6 for two reasons. First, some of the detailed occupational categories have very few members in the parent generation, and observations with dummy variables that perfectly predict outcomes are dropped from the regression. Second, when including data on the mothers, we impose the same age requirements (30–60), and there are also some individuals whose mother’s identity is unknown. A majority of mothers in the early samples do not have a stated occupation; these are kept in the sample and assigned to an additional “homemaker” occupational category. Models with mothers also incorporate a second-degree polynomial in mother’s age at time of observation. Results of regressions on balanced samples give similar results and are available upon request.
↵18. This is the full set of measures reported by the Stata fitstat command. In addition to the unadjusted McFadden R2 reported in Table 6, these are: adjusted McFadden, McKelvey and Zavoina, Cox-Snell/ML, Cragg-Uhler/Nagelkerke, Efron, Tjur’s D, and unadjusted and adjusted count R2.
↵19. If a linear probability model is used, differences in coefficient magnitudes between the reference model and the models with more details on parents’ occupations is larger. However, statistically significant grandparent estimates remain so with additional controls in the linear probability case as well. Results are available upon request.
↵20. We do not use the Zeng and Xie (2014) coresidence approach directly, as grandparental coresidence is very rare in Norway throughout the study period. In 1865, only 7 percent of households with children also had a resident grandparent; in 1910, this was down to 3 percent.
↵21. All results are qualitatively similar if we instead use the model with a full set of controls in the parent generation (see Online Appendix Table A16) or if a linear probability model is used (results available upon request).
↵22. Farmers are also less likely to move. Compared to white-collar occupations, the difference in movement propensity is 27 percentage points in Sample A (14 versus 41 percent) decreasing to 18 percentage points in Sample D (20 versus 38 percent).
↵23. The regression coefficients are shown in Online Appendix Tables A17–A18.
- Received October 2018.
- Accepted February 2021.








