Abstract
Does holding schools accountable for student performance cause good teachers to leave low-performing schools? Using data from New York City, which assigns accountability grades to schools on the basis of student achievement, I perform a regression discontinuity analysis and find evidence of the opposite effect. At the bottom end of the school grade distribution, lower accountability grades decrease teacher turnover and increase joining teachers’ quality. A likely channel is that accountability pressures increase principal effort at lower-graded schools, which teachers value. In contrast, at the top end of the school grade distribution, low accountability grades may negatively impact joining teachers’ quality.
I. Introduction
Since the mid-1990s, school accountability systems have become a central focus of education reform in the United States. Even before the No Child Left Behind Act (NCLB) made accountability mandatory across the United States in 2001, many states and districts had already instituted some form of accountability.
Policymakers and observers often worry that these systems, which attempt to hold schools accountable for student performance, could make it difficult for low-performing schools to attract and retain good teachers. Evidence from surveys of teachers suggests that good teachers may want to avoid the stress, restricted autonomy, and emphasis on “teaching to the test” that they think accountability brings to low-performing schools (Jones et al. 1999; Kirtley 2012). Since high-quality teachers improve students’ long-run educational attainment and earnings more than low-quality teachers (Chetty, Friedman, and Rockoff 2014; Kane, Rockoff and Staiger 2008), this means that accountability systems could have lasting, negative implications for students at low-performing schools. Indeed, some have suggested that poor accountability ratings could start low-performing schools down a negative quality spiral where high-quality teachers leave, causing the best students to leave, thus causing more teachers to leave, and so forth.
However, from a theoretical perspective, the effect of low school accountability ratings on teacher quality is actually ambiguous. For example, teachers could prefer to teach in lower-rated schools. Although that may sound counterintuitive, one potential reason is school improvement: low accountability ratings are designed to improve school performance, and a large literature shows that they do (Carnoy and Loeb 2002; Chiang 2009; Figlio and Rouse 2006; Hanushek and Raymond 2005; West and Peterson 2006; Rockoff and Turner 2010). Teachers may prefer to teach in schools where achievement is improving, perhaps because they value achievement or because the process of improvement is satisfying. Thus, it is ultimately an empirical question whether the impact of accountability pressures on teachers is positive or negative.
This paper exploits the introduction of an accountability system in the New York City Department of Education (NYCDOE) to provide new evidence on this issue. In November 2007, the NYCDOE launched a comprehensive accountability system that assigned schools letter grades for school performance. The grades were based on continuous performance metrics, with determination of the actual grade based on strict thresholds. This allows for the use of a regression discontinuity (RD) analysis to estimate the effect of the reform, as previously shown by Rockoff and Turner (2010). While these authors focused on the within-year impacts of receiving a low grade on student performance (finding large positive improvements), here I focus on the impacts on the teacher labor market.
Using data from the first two years of NYCDOE’s accountability system, I find evidence that accountability pressures can, in some cases, improve teacher retention and recruitment. At the bottom end of the school grade distribution (that is, the C/D and D/F thresholds), where the NYCDOE accountability sanctions have more “bite,” receipt of a lower school accountability grade, which occurs early in the school year, decreases teacher turnover at the end of the year by three percentage points. This is a large effect, representing roughly 20 percent of baseline turnover. This pattern is robust across several different specifications: I consistently reject the null of no effect and, a fortiori, the conventional wisdom that low accountability scores might lead to a sizeable increase in teacher turnover. The decrease in turnover likely benefits students in low-graded schools, as turnover has been shown to decrease student achievement (Ronfeldt, Loeb, and Wyckoff 2013).
I next examine the sorting implications of accountability grades and again find evidence that lower accountability grades help the low-performing schools at the bottom end of the grade distribution. The evidence suggests that joiners to lower-graded schools have higher value-added than the joiners to higher-graded schools, but there is no difference in the value-added of leavers between lower-graded and higher-graded schools.
There are two main hypotheses that could explain these effects. The first hypothesis is that receiving a lower school grade increases the attractiveness of the jobs at the school (the job desirability hypothesis). One potential reason is school improvement: Rockoff and Turner (2010) show that the lower-rated schools in the NYCDOE improved their performance within the same year the grade was received. Teachers may prefer to teach in schools where achievement is improving because they value achievement per se, appreciate the changes that enable the performance improvements, or expect that it will lead to higher accountability grades in the future. Another potential channel for job desirability is that, induced by accountability pressure, principals at lower-graded schools put more effort into making the schools better places for teachers to work or into attracting and retaining high-quality teachers. The second hypothesis is that receiving a lower school grade attaches a negative stigma to the teachers at the school, reducing their perceived value to potential employers, who view low grades as a signal of low unobservable teacher quality (the stigma hypothesis). Note that this is a rational hypothesis.1
I argue that the job desirability hypothesis matches the data better than the stigma hypothesis. The fact that lower-graded schools have higher-quality joiners than higher-graded schools is more consistent with increased job desirability. Also suggestive is the fact that the effect of lower grades on turnover is primarily driven by a decrease in out-of-district departures. Data on the transfer applications submitted by teachers also show that teachers who applied to transfer out of lower-graded schools were no less likely to be able to transfer out and that the number of teachers who applied to transfer to lower-graded schools did not fall.
The results of surveys conducted with teachers provide additional evidence for the potential mechanisms underlying job desirability. At the bottom end of the grade distribution, teachers at schools that received lower accountability grades at the beginning of the year gave their principals higher leadership ratings at the end of the year than teachers at higher-graded schools, agreeing more strongly with statements such as “the principal is an effective manager who makes the school run smoothly” and “I feel strongly supported by my principal” (note that the results are significant at the 10 percent level). Since principal leadership ratings also correlate with lower turnover, these effects suggest that one channel for the positive impact of lower grades may be that principals at low-graded schools respond to accountability pressures by making management changes that appeal to teachers.
Given the important role for principals suggested by these results, I also look at heterogeneity in the RD impacts of low grades by baseline measures of principal quality. I find that the decreases in turnover at lower-graded schools are driven primarily by schools with higher-quality principals (as measured by baseline principal leadership ratings). This suggests that principal capacity to transform accountability pressures into positive changes for the school may be a key ingredient that enables low accountability grades to have positive labor market effects.
Thus, the results suggest that the accountability system benefited low-performing schools at the bottom end of the grade distribution through two labor market channels: decreased turnover and increased teacher quality. These effects appear to be driven by the low-graded schools becoming more attractive to teachers. However, at the top end of the school grade distribution (the A/B and B/C thresholds), the results differ. Here, I find no evidence of positive effects of receiving a lower grade, and there is some suggestive evidence of negative impacts on the quality of the joiners relative to the leavers and on teacher survey responses about their principals’ leadership.
The difference in the results at the top and bottom ends of the grade distribution (that is, the fact that accountability seems to benefit lower-rated schools at the C/D and D/F thresholds while hurting them at the A/B and B/C thresholds) likely reflects the fact that accountability pressures are higher at the C/D and D/F thresholds, and so only motivated positive changes there (Rockoff and Turner 2010).2 Are there other ingredients besides high stakes that help us predict when accountability systems will positively impact the teacher labor market? The results on mechanisms suggest that positive impacts may be more likely in settings where principals are good leaders and have the latitude to implement positive changes, but of course such discussion is still speculative at this point.
This paper contributes to the literature on determinants of teacher mobility and sorting (Scafidi, Sjoquist, and Stinebrickner 2007; Falch and Rønning 2007; Imazeki 2005; Hanushek, Kain, and Rivkin 2004; Hendricks 2014; Dolton and von der Klaauw 1995; Hensvik 2012; Venhorst, Van Dijk, and Van Wissen 2011; Jackson 2009, 2012). Within the more specific literature examining how accountability affects teachers, the paper addresses two of the primary challenges. The first challenge is identification: accountability reforms are often instituted simultaneously with many other reforms that also affect teachers, making it difficult to identify accountability’s effects cleanly. As with all RD designs, the analysis used in this paper focuses on schools right next to the grade thresholds and thus holds fixed the effects of concurrent reforms. This builds on some earlier papers in the literature that used difference-in-differences strategies that could be more subject to these identification concerns (Clotfelter et al. 2004; Boyd et al. 2008b), with Feng, Figlio, and Sass (2010) and Gjefsen and Gunnes (2016) more recent exceptions. Feng, Figlio, and Sass (2010) use unexpected changes to the school accountability grading system in Florida that exogenously “shocked” some schools’ grades, while Gjefsen and Gunnes (2016) use triple difference specifications.3
The second challenge is finding good data on teacher quality and, specifically, on teachers’ contributions to student learning, or their “value-added.” Having a good measure of teacher quality is important for understanding the implications of accountability: high turnover could either reflect high-quality teachers leaving or low-quality teachers being pushed out. Value-added is widely regarded as unmatched as a measure of teacher quality, with, for example, important predictive power over students’ long-run outcomes (Chetty, Friedman, and Rockoff 2014). Unfortunately, value-added estimation has extensive data requirements, so most of the existing literature (Boyd et al. 2008b; Clotfelter et al. 2004; Gjefsen and Gunnes 2016) could only use other teacher characteristics, which often do not proxy well for value-added (see, for example, Hanushek and Rivkin 2010; Rivkin, Hanushek, and Kain 2005).4
This paper makes several new contributions that go beyond the results contributed by Feng, Figlio, and Sass (2010), which also make use of value-added data to examine a similar question. First, my analysis speaks to how accountability affects teacher quality in a more comprehensive manner, by incorporating value-added data on joining teachers, not just leaving teachers. This is important because, as the results here show, there can be asymmetric effects for the two groups, and it is the comparison of the two effects that determines the overall impact. Second, this is the first paper to unpack the mechanisms behind accountability’s effects on teachers by trying to disentangle the two main factors that affect teacher mobility: changes in choice sets vs. changes in job desirability. This helps us learn not just about accountability’s channels, but also about teacher preferences and mobility decisions more broadly.
In contrast to my findings, previous literature looking at similar phenomena largely concludes that accountability pressures hurt low-performing schools by accelerating turnover (Feng, Figlio, and Sass 2010; Clotfelter et al. 2004; Gjefsen and Gunnes 2016), especially at the bottom of the grade distribution in Feng, Figlio, and Sass (2010). Section VII discusses potential reasons for the different results.
The remainder of the paper proceeds as follows: Section II describes the institutional background. Section III describes the data and empirical strategy. Section IV presents the main results, while Section V discusses potential mechanisms for the results. Section VI examines robustness. In Section VII, I discuss my results in the context of the overall literature. Section VIII concludes.
II. Background
A. The NYCDOE Accountability System
I now review the key features of the NYCDOE accountability system, much of which was previously described in Rockoff and Turner (2010). The NYCDOE launched its current accountability system in November of 2007. Under the system, schools receive progress reports with letter grades meant to capture school performance relative to peer schools. The progress report also contains the school’s NCLB status, and the score from a school’s quality review, a two-to three-day qualitative evaluation. The NYCDOE links the letter grades with rewards and sanctions and makes the reports publicly available in an effort to incentivize low-performing schools to improve their performance.
The letter grade is based on a numeric score. For elementary and middle schools (the focus of this study), the score reflects three measures: student progress, student performance, and school environment. Student progress represents 60 percent of the score and measures year-to-year changes in student scores on the New York State standardized tests in mathematics and English language arts (ELA). Student performance (25 percent of the score) captures the level of test scores. School environment (15 percent of the score) reflects attendance and parent, student, and teacher surveys results.
School scores are calculated as a weighted average of the school’s “city horizon score” (one-third weight), which compares the school to all others of the same school type (that is, that serve the same grades), and its “peer horizon score” (two-thirds), which compares it to a peer group of up to 40 similar schools.5 The overall pre-additional-credit score, which ranges from 0 to 100, is then calculated as the weighted average of the scores for each grading measure. Schools can also earn additional credit if their “high-need” students make “exemplary gains” (that is, improve their performance by at least one-half of a proficiency level in ELA or math). The credit is added to the school’s pre-additional-credit score to determine the final score.
Thresholds for letter grade assignment are determined based on the distribution within school type of pre-additional-credit scores. For example, in the first year of the program, the NYCDOE set the threshold for receipt of an A, B, C, and D at the 85th, 45th, 15th, and 5th percentiles of pre-additional-credit scores, respectively. Grades are then determined by comparing each school’s score to the thresholds, with the thresholds strict (see Online Appendix Figure A1).
The NYCDOE links the letter grades with rewards and sanctions. Quoting the guidelines, “schools that are given an overall grade of A receive financial rewards, unless they score poorly on the Quality Review. Schools that receive an overall grade of D or F are subject to school improvement measures and target setting and, if no progress is made over time, possible leadership change, restructuring, or closure. The same is true for schools receiving a C for three years in a row. Over time, school organizations receiving an overall grade of F are likely to be closed. Ultimately, schools are accountable for making progress and receiving an overall grade of A, B, or C” (NYCDOE website, 2010).6 The sanctions associated with receiving low accountability grades (that is, D or F) are significant. For example, after receiving the first report cards in November 2007, the NYCDOE told five F schools in December that they would be closed immediately or phased out at the end of the school year. At the top end of the grade distribution, higher letter grades are associated with modest rewards. Principals of schools that had a score among the top 20 percent of schools and that received a “Well Developed” or “Proficient” quality review rating were eligible for bonuses of $7,000–$25,000. Schools receiving an A and a “Well Developed” quality review rating received roughly $33 per student in extra funds, to be used at the principal’s discretion. Finally, schools receiving an A or B grade and a “Well Developed” or “Proficient” quality review rating received $1,500–$3,000 per student that transferred in from an F school or a school not in good standing under NCLB.
Besides these consequences, there were no other financial ramifications of school grades; for example, D and F schools did not receive any additional funding, resources, or staff for school improvement.
B. The NYCDOE Teacher Transfer System
Since 2005, NYCDOE’s staffing system has been an open market system built around the principle of “mutual consent”: schools post openings, teachers apply, principals choose which teachers to hire, and then teachers decide whether to accept the job offers. This means that the effects estimated in this paper are the effects on equilibrium matches within a market-based system.
C. Differences between NYCDOE and Other U.S. Contexts
The NYCDOE system has lower paperwork requirements for failing schools than some other accountability systems, such as Florida, where failing schools are required to complete regular, extensive reports (Rouse et al. 2013). The base NYCDOE system also does not link progress reports with teacher-level incentives, whereas North Carolina’s and Florida’s systems include teacher performance bonuses.7 Teachers unions are much stronger in New York than in Florida or North Carolina, which could partially explain some of the differences in the accountability program designs. Finally, some accountability systems give failing schools additional funding for school improvement, but that was not the case in the NYCDOE.
III. Data and Empirical Strategy
A. Data
I use data from several sources within the NYCDOE. The accountability data come from publicly available files downloaded from the NYCDOE website. The data contain each school’s accountability score and components, as well as NCLB status, quality review rating, and school identifiers. The data are available for the 2007–2008 through 2011–2012 school years (where the school year given is the year in which the accountability grade was released; report cards are released in fall of the school year and depend on performance results from the previous year.)
The second data source is demographic and exam performance data at the student level, provided by the NYCDOE and covering the 1998–1999 through 2008–2009 school years. The demographic data include gender, ethnicity, free-lunch, and specialeducation status. The exam performance data include student scores on mathematics and ELA tests administered statewide in fourth and eighth grade and citywide in the third, fifth, sixth, and seventh grades.
The majority of the teacher data come from the NYCDOE HR/payroll system and contain teacher experience and education, and school- and grade-level identifiers, from the 1999–2000 through the 2009–2010 school years. This is my primary source of data on teacher turnover and mobility. Based on guidance from the NYCDOE, to calculate turnover, I define a teacher as having left a school if she leaves between May of one school year and November of the subsequent school year, since the (rare) midyear departures tend to reflect emergencies (for example, sickness, birth) and would increase noise, but I show robustness to this definition.
I combine these data with annual data from the open market transfer system showing which teachers submitted transfer applications, with a tag for which applicants successfully transferred. Data are available from the 2006–2007 through 2009–2010 school years. Although the HR/payroll system allows me to calculate transfers, the open market data are useful for understanding (i) which teachers want to transfer even if not successful and (ii) how many transfer applications different schools receive. However, the data have some weaknesses. First, although I see every application submitted by the 60 percent of teachers who did not ultimately transfer, for the 40 percent of applicants who do transfer, I only see the application to the school to which they ultimately transferred. This means that I can create a proxy for how many transfer applications a given school received, but it is incomplete. Second, the open market transfer data do not always align with the payroll system data. Although 97 percent of the successful transfers in the open market transfer data are associated with transfers in the payroll data, only 80 percent of the transfers in the payroll data are associated with a transfer application, and only 63 percent with a successful transfer application. NYCDOE has not been able to identify the source of discrepancies or provide me with updated data, but has advised me that the administrative data from the payroll system are the more reliable source for actual teacher transfers. My primary results thus use the payroll system data (the payroll data also contain additional outcomes such as leaving NYCDOE entirely instead of transferring); the open market transfer data are used as suggestive supplementary evidence.
I also use data from surveys conducted with teachers by the NYCDOE as part of the accountability system. Surveys were conducted near the end of the school year, after high-stakes tests were conducted but before the results were received. I use survey data from the 2006–2007 through 2009–2010 school years. The schools asked questions on a broad range of issues, from contact with parents to school safety to interaction with principals.
I study schools in the first two years that accountability grades were released: the 2007–2008 and 2008–2009 school years.8
B. Value-Added Estimation
To estimate teacher value-added, I created a matched panel of student and teacher data.9 I use the approach that has been experimentally validated in the economics of education literature (Kane and Staiger 2008). Online Appendix A describes the estimation in detail; I follow the recent literature and use empirical Bayes shrinkage estimates (Jackson 2012; Koedel, Mihaly, and Rockoff 2015). (Recent literature has highlighted the potential biases of the value-added approach; see Online Appendix 1 for discussion of why these biases should not be problematic here.) A primary strength of NYCDOE data is that the matched panel exists for eight years prior to the institution of accountability. This allows me to estimate value-added using data from the preaccountability period and to avoid conflating teacher quality with responses to accountability. As a result, value-added is only available for teachers who taught in tested grades before 2008.
Estimated value-added data are unfortunately only available for a subset of teachers (27 percent of all teachers in the sample, 21 percent of leavers, and 9 percent of joiners). I thus follow Jackson (2012) and calculate a predicted value-added measure based on observable characteristics that are available for all teachers in the data set. See Online Appendix A.1 for more details. Note that because I have a more limited set of observable characteristics than used in some of the existing literature (for example, Jackson 2012), the correlation between my predicted value-added measure and estimated value-added is relatively small (correlation of 0.06), although highly statistically significant. I thus focus on the estimated value-added results as my main measure in the exposition, but use the predicted value-added results to provide suggestive evidence of whether the results in the subsample with estimated value-added data might extend more broadly.
C. Empirical Strategy
The RD approach adopted in this paper is similar to much of the literature studying the effects of accountability grades (for example, Rouse et al. 2013; Chiang 2009) and most closely follows Rockoff and Turner (2010), who use a similar specification to estimate how the NYCDOE accountability reforms affected short-run achievement. I estimate equations of the following form:
(1)
where j indexes teachers, t indexes time, g indexes accountability grades, Yjt is the outcome variable of interest (for example, an indicator that the teacher left the school), Igjt is an indicator for the grade received by a school, Sjt is the school’s accountability score, h( ) represents a flexible control function allowed to differ on either side of the grade threshold, and εjt is a mean 0 error term. I follow Rockoff and Turner (2010) in interacting the control function with an indicator for school type,
, as well as for the year, It, since the grade thresholds are specific for all school types and years.10 My base specification for h() follows Hahn et al. (2001) and much of the recent literature (for example, Malamud and Pop-Eleches 2011) in using a locally linear control function and a rectangular kernel.11 I also explore robustness to parametric regression functions. All standard errors are clustered at the school level.
To increase statistical power, I follow Rockoff and Turner (2010),12 and group the schools from the bottom thresholds together (C/D and D/F) and from the top thresholds together (A/B and B/C). The accountability pressures placed on schools (and the marginal increases in those pressures across grade thresholds) are much stronger at the bottom of the grade distribution than at the top, and indeed, Rockoff and Turner (2010) find that accountability induced improvements in achievement at both bottom grade thresholds but not at the top.13
Online Appendix B discusses selection of the base bandwidth used for the analysis. I also explore the sensitivity of the results to a range of bandwidths.
As with all RD analyses, the treatment effect under estimation is local to schools adjacent to the grade thresholds and does not capture any universal effects of accountability. Since the analysis pools estimates from multiple cutoffs, the treatment effect represents the weighted average of the local effects at each individual cutoff, where cutoff values that are more likely to occur and that have more observations are given higher weight (Cattaneo et al. 2016). The identification assumption is that, conditional on the continuous metric underlying the grade, the grade itself is exogenous. Below I discuss the two primary potential violations of this assumption: manipulation of the running variable and “ex post gaming” (that is, selection into the analysis sample based on the grade received).
1. Density evidence
A primary potential violation of the RD identification assumption is manipulation of the running variable. Manipulation is unlikely to be problematic in this setting because of the use of fixed grade thresholds and the fact that the underlying components of the score are all publicly verifiable and difficult to manipulate precisely (like test scores).14 To investigate this further, I perform the density test suggested in McCrary (2008) and do not find evidence of bunching. Specifically, I use the number of schools within a 0.5-point or 0.1-point bandwidth as the dependent variable in Equation 1 and the base bandwidth used in the analysis (five score points), and none of the coefficients on the dummy for being on the lower-graded side of the grade threshold are statistically significant. See Online Appendix Table A1 for results.15
2. Sample construction and “ex post gaming”
My sample consists of all noncharter elementary, K–8, and middle schools that received accountability grades in the 2007–2008 or 2008–2009 school years (1,987 school–year observations).16 I exclude the school–year observations where the school closed in the following year: this represents five school–year observations, all from the 2007–2008 school year. As a result, unlike in most RD settings, one could be concerned about “ex post gaming” here—if administrators took accountability grades into account when selecting schools for closure, this would violate the identification assumption. However, most of the school closures took place far from the thresholds. Only one of the schools that was closed during the sample period fell within a five-point bandwidth of any grade threshold (the base bandwidth used in the paper), and there is no significant RD relationship between that closure and accountability grades (see Online Appendix Table A2, Column 1). Since just one school–year observation is excluded, the results are also robust to bounding exercises assuming that outcomes at that school fall at the extreme ends of the distribution (see Section VI). Thus, ex post selection is not driving the results. I also exclude six school–year observations that do not appear in the teacher data (two within a five-point bandwidth of any grade threshold). These observations fall across the grade distribution, and there is no significant RD relationship between accountability grades and missing data (Online Appendix Table A2, Column 2). The final sample has 1,976 school–year observations, 1,243 within five points of any grade threshold.
In addition to schools that were immediately closed, there were also schools that were phased out over time. According to Rockoff and Turner (2010), at the end of the 2007–2008 school year, the NYCDOE announced that seven schools would be either closed or phased out. The administrative data on school closures show that only five schools were closed immediately, leaving two schools that began phase-out at the end of the first year of the program, and potentially more at the end of the second year. These phase-out schools are included in the base analysis sample. Although they do not represent a threat to identification, from an interpretation perspective, it is important to check whether different dynamics at these schools affect the results. Using a proxy for phase-outs,17 I show that they do not affect the results: First, there is no significant RD relationship between phase-outs and accountability grades (Online Appendix Table A2, Column 3), and, second, the results are robust to dropping phase-out schools (see Section VI).
My base analysis sample includes all schools. However, I also show my main results for a restricted sample that excludes the schools that had begun restructuring prior to the institution of accountability (that is, schools in year 2+ of restructuring in the 2007–2008 school year). Restructuring often entails significant staffing shifts that are predetermined to accountability and could dampen their ability to respond to the schoolbased accountability system. Thus, while the full sample has broader external validity, the restricted sample is interesting to look at for a more pure examination of how schools respond to one specific accountability system.18
D. Summary Statistics
Descriptive statistics for the base analysis sample (that is, within five points of the grade thresholds) are presented in Table 1. (Statistics are very similar when excluding the restructuring schools, shown in Online Appendix Table A10.) Panel A shows descriptive statistics about the sample of teachers teaching in the base analysis sample schools in the 2007–2008 and 2008–2009 school years. The two-year panel contains 50,616 unique teachers and 71,677 teacher–year observations. Roughly 27 percent of the teachers have math value-added data.19 Baseline teacher value-added generally increases with the accountability grade.
Panel A of Table 1 shows that there is 11 percent teacher turnover across the sample period, with turnover increasing across accountability grades. Eight percent of the turnover is teacher retirements, 33 percent is transfers made between teaching positions in the NYCDOE, and 57 percent reflects departures from NYCDOE.20 Turnover is generally higher among less-experienced and less-educated teachers.21
Table 1 shows that 15 percent of schools in the sample were part of the New York School Bonus Program (NYSBP), a pilot program of teacher incentive pay started by the NYCDOE in the fall of 2007 in a small subset of schools. The program had limited impact, with no evidence of effects on overall teacher or student behavior (Fryer 2013; Goodman and Turner 2010; Springer 2011). Since program participation was randomly assigned and determined both prior to and independently of the accountability system, it should be unrelated to accountability grades and should not affect the results; indeed, Column 5 of Online Appendix Table A2 shows that there is no RD relationship between grades and the program. I also include a NYSBP control in the vector of school-level controls in the regressions (described more below), but the results are invariant to exclusion of the control.
E. Balance Tests
Online Appendix Table A3 tests for balance in baseline characteristics by estimating Equation 1 using baseline characteristics as the outcome variables. Each column contains two cells displaying the results from separate regressions, with the upper cell showing the results from the bottom end of the grade distribution (C/D and D/F thresholds pooled) and the lower cell showing the results from the top end (A/B and B/C pooled). The coefficient on the indicator for receiving the lower grade at the threshold is shown. I find roughly the number of significant coefficients that would be expected due to chance (5.8 percent of coefficients are significant at the 5 percent level). I thus follow Lee and Lemieux (2010) and perform a joint test of the significance of all coefficients and, reassuringly, fail to reject the null that all of the coefficients are equal to 0 (the p-values are 0.30 for the top grade thresholds and 0.16 for the bottom grade thresholds). At the bottom of the grade distribution (the primary focus of the paper), the specific variables that are unbalanced are related to the racial composition of the school: the lower-graded schools have fewer black students and teachers and more Hispanic students. These three variables are highly correlated with each other, with the correlation between percent black students and percent black teachers (percent Hispanic students) equal to 0.82 (−0.39). All regressions control for a vector of school-level variables that includes any variables unbalanced at the 5 percent level, in addition to other controls designed to improve precision.22
IV. Results
A. Turnover Results
I begin with graphical evidence. The left column of Figure 1 plots average residual turnover against accountability scores. The top graph shows the results at the bottom of the grade distribution (C/D and D/F pooled), and the lower graph shows the top pooled (A/B and B/C). To create residual turnover, I regress an indicator for whether a teacher left a school on a vector of covariates.23 I then group schools according to their accountability scores relative to the grade threshold, with each dot representing the schools within a one-point bandwidth and plot the average residual turnover at the end of the school year on the average accountability score received at the beginning of the year.24 For ease of interpretation, I also plot a locally linear regression line in the figures, fitted to the raw teacher-level data and fitted separately on either side of the grade threshold. Shaded areas represent 95 percent confidence intervals around the line.
Residual Turnover, by Accountability Score
Notes: The left column plots the actual turnover results. The x-axes show schools’ average accountability scores relative to the closest grade threshold (so the grade threshold is always displayed at zero). Each dot represents the average for all schools within a one-point bandwidth on the x-axis. The y-axes show average residual turnover in the summer after the schools received their accountability grades. The right column shows placebo turnover results: there, the y-axes show residual turnover in the year before schools received grades. Residual turnover is calculated by regressing an indicator for leaving a school on a vector of covariates (see Table 2 notes for a list of covariates). The lines correspond to local polynomial smooth plots, and the shaded areas represent their 95 percent confidence intervals.
At the bottom of the grade distribution (C/D and D/F), lower-graded schools appear to have locally lower turnover. In contrast, at the top of the grade distribution, there is virtually no discontinuity at the threshold.
To test for the significance of these results, Columns 1 and 2 of Table 2 present the regression results, calculated from estimation of Equation 1 using an indicator for whether a teacher left the school at the end of the school year as the dependent variable. Each cell contains the results from a separate regression, with the results from the bottom end of the grade distribution shown first. All regressions control for the vector of controls outlined in Section III.E. Column 1 shows the full-sample results, whereas Column 2 shows the results in the nonrestructuring sample.
Consistent with the graphical evidence, the regressions show that at the bottom end of the grade distribution lower accountability grades decrease turnover by roughly three percentage points, significant at the 5 percent level. This reduction in turnover should provide a direct benefit to low-graded schools, since turnover has been shown to decrease student achievement (Ronfeldt, Loeb, and Wyckoff 2013). The magnitude of the effect is economically meaningful, equal to roughly 20 percent of average turnover among those schools. To put the magnitude in the context of the literature, studies on the elasticity of turnover to wage changes find elasticities ranging from roughly −1 to −1.5 (Hanushek, Kain, and Rivkin 2004; Hendricks 2014; Dolton and von der Klaauw 1995) to roughly −3 to −4 for women and −3 to −8 for men (Imazeki 2005; Clotfelter et al. 2008), implying that the 20 percent decrease in turnover seen here would be equivalent to an increase in salary of anywhere from 2.5–7 percent to 13–20 percent. The magnitude of the point estimates is substantial, but seems reasonable in the context of the literature on mobility responses to other reforms, including accountability, for which estimates of the magnitude of the effects on teacher turnover range from 25 percent to 65 percent.25 While the point estimates are large in magnitude, since turnover is a rare event, the confidence intervals are relatively wide.
In contrast, at the top end of the grade distribution, where accountability pressures have less bite, grades do not affect turnover: the estimate is small in magnitude (0.5 percentage points, Column 1) and not statistically distinguishable from zero. I examine the robustness of the turnover results in Section VI.A.
The credibility of the RD design rests on the assumption that schools are as if randomly assigned at the grade thresholds. To examine the validity of this assumption, I also perform a placebo test, checking whether there are any baseline differences in turnover between schools on either side of grade thresholds. Thus, I estimate the exact same regression model, but using turnover from the previous year, that is, from the year before the accountability grade was received, as the dependent variable.26 Column 3 of Table 2 shows the placebo tests using both years of data, while Column 4 shows the placebo from the first year of the program only in order to limit to preaccountability data.27 Reassuringly, none of the placebo coefficients are statistically significant, and the magnitudes are relatively small. The right column in Figure 1 presents the placebo results graphically, with consistent results.
B. Teacher Sorting Results
I now look at how accountability grades affect the quality of the teachers who leave schools (leavers) and the teachers who join in the next year (joiners). Note that the results in this section should be interpreted as suggestive because the samples of teachers are small, especially the sample of joiners with estimated value-added data. The placebo tests presented in the next subsection are also not dispositive.
Panel A of Table 3 presents the RD estimates of the effect of receiving a lower grade on the average quality of leavers and joiners, calculated by estimating Equation 1 using teacher value-added as the dependent variable and using either the leavers or the joiners as the sample. I use mathematics value-added as my value-added measure since the literature has shown that teacher fixed effects in mathematics tend to have more predictive power over student outcomes than ELA fixed effects (for example, Jacob and Lefgren 2008; Jackson and Bruegmann 2009). This probably reflects the fact that, while most students learn language skills from many sources (for example, their parents, the television), the primary source of math knowledge for many students is their teachers.28 Figures 2 and 3 depict the estimated value-added results graphically, with consistent results.
There are two ways to examine leaver quality: (i) by looking at how leaver quality varies on either side of the grade threshold or (ii) by estimating the turnover regressions from the previous section separately for teachers with different value-added. The former is shown in Table 3, while the latter approach is shown in Table 4.29 I focus on the former specification for the exposition since it enables an easier comparison with the joiner quality results, but the results are consistent.
Beginning with the bottom end of the grade distribution, where we saw that lower accountability grades decreased average turnover, there is not strong evidence that turnover varies by quality (Columns 1–4). The point estimates are negative using estimated value-added (implying that higher value-added teachers were more likely to stay), but positive using predicted value-added. None of the coefficients using this specification are statistically significant, with the estimated value-added result significant at the 10 percent level in the Table 4 specification. However, there is suggestive evidence of an impact on joiner quality. The point estimates in Columns 5–6 of Table 3 suggest that lower-graded schools attract joiners with 0.9 SD higher estimated value-added, statistically significant at the 5 percent level. Since estimated value-added is only available for a small subsample of the joiners, it is reassuring that the results using predicted value-added, which is available for the full sample, are qualitatively consistent, if weaker, potentially reflecting the low power of the predicted value-added measure in predicting value-added (Columns 7–8). It is also reassuring that there is no significant selection into having estimated value-added data by accountability grade (Panel B, Column 4). Overall the results provide suggestive evidence of positive sorting implications for lower-graded schools, as lower-graded schools may attract higher-quality joiners relative to leavers.
Average Math Value-Added of Leavers, by Accountability Score
Notes: The left column plots the actual leaver quality results. The x-axes show the schools’ average accountability scores relative to the grade threshold (so the grade threshold is always displayed at zero). The y-axes show the average value-added of leavers (that is, of the teachers who left their schools in the summer after their schools received the accountability score and grade). The right column shows the placebo results: there, the y-axes show the average value-added of the teachers who left their schools the year before their schools received the accountability score and grade. The lines correspond to local polynomial smooth plots, and the shaded areas represent their 95 percent confidence intervals.
In contrast, at the top end of the grade distribution (that is, A/B and B/C), there is no evidence of a positive effect on quality, and some evidence that lower accountability grades may, if anything, decrease the quality of joiners relative to leavers. However, the results are only suggestive, as the leaver result only shows up in the predicted value-added data and the joiner only in the estimated value-added.30 Table 4 and Panel B of Table 3 also show results for other teacher characteristics (experience and education); none of the results are statistically significant. In general, no single characteristic proxies well for value-added (Rivkin, Hanushek, and Kain 2005; Hanushek and Rivkin 2010), so I do not see these results as inconsistent with the others.31
Online Appendix Table A4 presents placebo results regressing the average previous-year characteristics of leavers and joiners on accountability grades, shown for both years of data (Panel A) and for the first year only (Panel B). There are roughly the number of significant coefficients that would be expected due to chance in Panel A (none significant at the 5 percent level and 2/16 significant at the 10 percent level), and a somewhat higher percentage for Panel B (2/16 significant at the 5 percent level). See the right column of Figures 2 and 3 for placebo graphs. The robustness of the joiner results is examined in Section VI.B.
Average Math Value-Added of Joiners, by Accountability Score
Notes: The left column plots the actual joiner quality results. The x-axes show schools’ average accountability scores relative to the closest grade threshold (so the grade threshold is always displayed at zero). Each dot represents the average for all schools within a one-point bandwidth on the x-axis. The y-axes show the average value-added of joiners (that is, of the teachers who joined schools in the summer after their schools received the accountability score and grade). The right column shows the placebo results: there, the y-axes show the average value-added of the teachers who joined their schools the year before their schools received the accountability score and grade. The lines correspond to local polynomial smooth plots, and the shaded areas represent their 95 percent confidence intervals.
C. Summary of Results
At the bottom end of the grade distribution, receiving a lower accountability grade causes teacher turnover to fall, and improves the quality of the joiners hired relative to the leavers. These results thus provide a more hopeful story for accountability than suggested by policymakers, implying that, through their labor market effects, accountability systems can in some cases benefit, not harm, the most disadvantaged schools. In contrast, at the top end of the grade distribution, the teacher labor market effects of receiving a lower grade are not positive, and, if anything, are negative. In the next section, I examine which mechanisms could explain these findings.
V. Mechanisms
The finding that receiving a lower accountability grade causes teacher turnover to decrease is somewhat unexpected. Turnover decisions can in general either reflect supply-side factors, demand-side factors, or some combination of the two. Here, supply-side factors would mean that the desirability of working at these schools has increased, which I refer to as the job desirability hypothesis. Demand-side factors would mean that external employers do not want to hire teachers from lower-graded schools, which I refer to as the stigma hypothesis.
The job desirability hypothesis is somewhat counterintuitive, since many of the factors one might first associate with lower grades (such as lower prestige, higher pressure, higher risk of future closure, etc.) are negative. Thus, the job desirability hypothesis requires that between the beginning of the year (when schools receive grades) and the end of the year (when teachers make turnover decisions), lower-graded schools respond to accountability pressures by making changes that make the school a more desirable place to work. One potential channel would be performance improvements: the previous literature has shown that schools respond to accountability pressures by improving their performance (Carnoy and Loeb 2002; Hanushek and Raymond 2005; Chiang 2009; Rockoff and Turner 2010). Accountability pressures could also motivate principals to do things that are attractive to teachers, such as focusing more on teacher development, offering more opportunities for teachers to collaborate, or providing more autonomy. Relatedly, accountability pressures could cause principals to focus more on hiring and retaining the best teachers (Shipps and White 2009). Alternatively, lower grades could foster collaboration as teachers work together to improve. Teachers could even like the challenge of the lower grade, perhaps because it makes them feel that they are making a difference for students.32
The evidence presented so far seems more aligned with the job desirability hypothesis, although is far from definitive. The fact that lower-graded schools attract higher-quality joiners provides some evidence for the job desirability hypothesis by suggesting that the lower-graded schools or principals made some positive changes or increased their recruitment efforts. The fact that the decrease in turnover is only seen at the bottom of the grade distribution may also support the job desirability hypothesis, as the change in accountability pressures when crossing a grade threshold were larger at the bottom end of the grade distribution, and thus any pressure-induced increases in job desirability should have been larger there too. In contrast, although possible, it is not clear why one would expect ex ante that stigma would change more across thresholds at the bottom end of the grade distribution than at the top.
I now provide further evidence on mechanisms. I first provide evidence for the plausibility and potential channels for the job desirability hypothesis. I cannot distinguish between all of the potential channels outlined above since I do not have data on them all and they all are likely highly correlated, but I do provide evidence that supports the hypothesis that there are some ways that the lower-graded schools were improving as well as suggestive evidence that these changes are attractive to teachers. I then use data on transfers and transfer applications to provide additional tests of the stigma hypothesis. I do not find evidence supporting the stigma hypothesis.
A. Mechanisms for the Supply-Side (Job Desirability) Hypothesis
Rockoff and Turner (2010) show that accountability pressures spurred achievement improvements at lower-graded schools at the bottom of the grade distribution in the NYCDOE even within the same year, so that is a plausible mechanism for improved job desirability. Results of teacher surveys conducted by the NYCDOE at the end of the school year can provide further evidence on mechanisms. The NYCDOE grouped questions in the survey by topical area (for example, principal leadership, campus environment); I created indexes for each topical area equal to the average standardized responses across all the questions in each section.33 Panel A of Table 5 presents the results. At the lower grade thresholds, the strongest impacts are on the principal leadership index, which is roughly 0.3 SD more positive at lower-graded schools than higher-graded schools, although only significant at the 10 percent level (Column 1). Panel B shows the detailed questions that make up the principal leadership index. Teachers were more likely to think that their principals supported them and that their principals were effective managers who made the school run smoothly. Since the differences in perceived principal leadership were not present at baseline, this suggests that low grades encouraged principals to work harder at their jobs and that teachers appreciated these changes.34
We do not see similar results at the top end of the grade distribution: accountability pressures may not have been large enough to induce principal responses, perhaps explaining why there also were no turnover impacts at those thresholds. In fact, if anything, there is suggestive evidence that the effect on survey responses were negative, consistent with the general trend that, unlike at the bottom of the grade distribution, at the top, low accountability grades seem to have negatively impacted schools, presumably because the accountability pressures were not strong enough to spur positive changes by principals and teachers.
Online Appendix Table A5 tests for other changes at lower-graded schools, testing for changes to the number of joiners, staff size, enrollment, class size, and principal turnover. (The change in number of leavers, that is, turnover, is included as Column 1 to compare with the joiner effects.) There are no significant impacts on any of these variables.35 Of course, other changes could be happening for which I do not have data.36
In order for changes such as higher student performance or better principal leadership to increase job desirability, teachers must value them. Table 6 provides suggestive evidence that teachers do value these types of changes by showing that, in general, when schools have higher performance or principal leadership survey scores (conditional on previous-year performance or principal leadership scores), those schools also experience lower teacher turnover.37
Further evidence on mechanisms can be provided by looking at whether the RD effects are larger at schools that we would predict would respond more positively to accountability pressures. The teacher survey results suggest an important role for principals; high-quality principals may be able to better channel accountability pressures into positive changes for their schools. Consistent with this hypothesis, Table 7 shows that the decreases in turnover are six percentage points larger among schools with above-median baseline principal leadership (as measured in the teacher surveys). This difference is substantial in magnitude and significant at the 5 percent level. This suggests that schools with good leadership may have been better able to translate accountability pressures into positive changes that the teachers appreciate. The results for heterogeneity in the effects on survey responses about principal leadership and joiner quality are not statistically significant, although the point estimate for the survey responses is large, suggesting three times as large an effect for schools with better principals at baseline. For joiner quality, the point estimate is actually negative, but confidence intervals are very wide, so I cannot reject that the improvements in joiner quality were substantially larger at schools with higher-quality principals.
B. Additional Tests: Destinations and Transfer Applications
The stigma and job desirability hypotheses have different implications for how the results will vary by teachers’ destinations. Since external (non-NYCDOE) employers are unlikely to look up school accountability grades, stigma should primarily affect intradistrict transfers, whereas the job desirability hypothesis could affect all types of turnover. Table 8 shows the turnover results by destination, with Column 1 replicating the overall result from Table 2 and Columns 2–4 showing the results separately for retirement, transfers between NYCDOE schools, or leaving the NYCDOE.38 The turnover result is driven almost entirely by fewer teachers leaving NYCDOE (Column 2), which accounts for more than 75 percent of the decrease in turnover at lower-graded schools, somewhat larger than the share of turnover driven by departures from the NYCDOE (50 percent). In contrast, within-district transfers (Column 4) represent a smaller percentage of the effect than of overall turnover. (Note that in neither case can we reject equality.)
Columns 5–8 bring in transfer application data from the open market transfer system to shed further light on the mechanisms (note that I consider these analyses suggestive due to the caveats with the open market data raised in Section III.A). Under the stigma hypothesis, transfer applicants from lower-graded schools should be less likely to receive job offers and thus to ultimately transfer. The number of transfer applications received by lower-graded schools should likely fall as well. This is not what the data suggest: using either the open market transfer data (Column 6) or the payroll data (Column 7) to define which applicants ultimately transferred, lower grades do not decrease the transfer rate among applicants, and lower-graded schools do not receive any fewer applications to transfer, with both point estimates in fact positive (that is, running against the stigma hypothesis).
Note that at the top of the grade distribution, consistent with the general trend that, if anything, lower accountability grades seem to make schools worse off through the labor market impacts, we see more transfer applications submitted (significant at the 10 percent level), although this is hard to interpret due to the discrepancies between the open market transfer and payroll data; in the payroll data, transfers themselves remain flat.
VI. Robustness of the RD Results
A. Robustness of the Turnover Results
Table 9 shows robustness of the turnover results to the RD specification used. Column 1 shows the results estimated without including baseline covariates; the coefficient remains negative, with the magnitude still large but smaller than the base specification (shown in Column 2), and standard errors much larger. Columns 3–6 show the results using linear specifications with a range of bandwidths (specifically, 50 percent and 200 percent of the base bandwidth, and the base bandwidth ±1). Columns 7–8 show that the results are qualitatively similar if one uses a parametric regression function (either quadratic or cubic in the accountability score, estimated separately by grade using a bandwidth of 200 percent the base bandwidth). Column 9 shows the specification linearly for all of the components of the accountability score separately instead of the composite score. The coefficient is smaller in magnitude but maintains its sign, and I cannot reject equality with my base specification.39 Column 10 shows the results after collapsing the data at the school level, weighted by the number of teachers at the school.
Given the noise in the graphs, one might be concerned that there are random breaks in the regression function. Per Lee and Lemieux (2010), I perform a specification test, testing for discontinuities at points other than the grade thresholds, and present p-values in Table 9.40 Reassuringly, the test statistic is not rejected in any of the specifications.
Online Appendix Table A6 shows that the results are not driven by different dynamics at schools that were being phased out, as the results are robust to excluding those schools (Column 2), nor are the results driven by sample selection due to the one school within five points of a grade threshold that was closed and thus excluded from the analysis, since the results are robust to a bounding exercise where we include the closing school and assign it turnover at the extremes of the distribution (the 99th percentile or first percentile, Columns 3–4). The results are also robust to counting midyear departures as turnover (Column 5).
Given the density and placebo tests presented earlier, I do not think that gaming is driving the results. However, looking at the results for 2007–2008 and 2008–2009 separately can also provide more insight: 2007–2008 was the first year of the accountability system, and so it is especially unlikely that schools could have manipulated their scores around the cutoffs in that year.41 Columns 6–7 show that, reassuringly, the results are similar when one estimates Equation 1 separately for the 2007–2008 and 2008–2009 school years. Online Appendix Table A7 also shows that the estimates are robust to excluding schools that fall directly on a grade threshold.42
One may also wonder if the results vary at the different thresholds that are grouped in the analysis (for example, C/D vs. D/F). Columns 1 and 2 of Online Appendix Table A8 show that the main turnover results are qualitatively consistent across thresholds. At the bottom end of the grade distribution, the magnitude of the coefficient is larger at the D/F threshold than the C/D, but I cannot reject equality.
B. Robustness of the Joiner Quality Results
Table 10 repeats the analyses from Table 9 to examine the robustness of the joiner quality results. The results at both the top and bottom ends of the grade distribution are relatively robust across specifications, although the coefficient magnitude and significance vary somewhat. The positive result at the bottom end of the grade distribution is not robust to the cubic specification. However, the parametric specifications (quadratic and cubic) may allow too much flexibility given the noise in the data, as suggested by the fact that the p-value for the specification test is rejected in the quadratic specification (this test can also be seen as a test for whether the regression function is well approximated by the control function within the bandwidth). Online Appendix Table A9 shows that the results are robust to including the schools undergoing restructuring or excluding the schools that were being phased out and that the results were qualitatively similar in 2008 and 2009.43 One potential concern with the results would be if they were driven primarily by teachers coming from the schools that were closed for having low accountability grades. Since there were no restrictions on where these teachers could be hired, this would not be an internal validity concern, but could be an external validity concern if there were more teachers from closed schools in the NYCDOE than there typically are in other settings. However, teachers from closing schools do not seem to drive the results: these teachers represent less than 2 percent of the overall population of joiners and less than 5 percent of the joiners with value-added data, and the results are robust to omitting these teachers from the analysis (Column 5).44
VII. Discussion
In this section, I discuss my findings in the context of the literature. This paper suggests that accountability pressures help low-performing schools by decreasing turnover and potentially improving teacher quality. The findings are inconsistent with the majority of the previous literature, which has largely found that accountability pressures hurt low-performing schools by accelerating turnover (Feng, Figlio, and Sass 2010; Clotfelter et al. 2004; Gjefsen and Gunnes 2016), and more so at the bottom of the grade distribution in Feng, Figlio, and Sass (2010).45
What might explain the differences between my results and some of the other findings in the literature? Part of the explanation is likely the stakes: Gjefsen and Gunnes (2016) study an accountability regime that does not have high-stakes rewards and sanctions, which likely explains the differences, as their results are more analogous to my results at the top end of the grade distribution than the bottom. However, other papers, such as Feng, Figlio, and Sass (2010), also study high-stakes settings and find different results. My results on heterogeneity by principal leadership capacity suggests that variation in principal capacity and control across settings may contribute to the difference in results. For example, in the Florida accountability system studied by Feng, Figlio, and Sass (2010), administrators do not seem to directly close low-graded schools as a disciplining mechanism as they do in the NYCDOE (Chiang 2009; Feng, Figlio, and Sass 2010). So, it may be the case that the Florida system does not place as much pressure on principals as the NYCDOE system. However, this is speculative, as I do not have enough information on the institutional differences regarding principals to say for sure, and there are of course many other differences between the settings.46 Thus, I cannot definitively say why the results differ; with just a few settings, the problem is not identified.
VIII. Conclusion
In this paper, I present evidence that accountability pressures impact the teacher labor market. At the bottom end of the grade distribution (the C/D and D/F thresholds), I find that accountability positively impacts lower-graded schools by decreasing turnover and by increasing the quality of joiners. These results echo the results from the earlier literature showing that lower accountability grades also improve academic performance (for example, Rockoff and Turner 2010). A plausible explanation here is that teachers actively choose to stay in the lower-graded schools because job desirability increases at those schools, perhaps because principals—especially high-quality principals—put more effort into leading their schools and into transforming the accountability pressures into positive change. In contrast, at the top end of the grade distribution (the A/B and B/C thresholds), where the accountability pressures are more mild, there is no evidence of positive impacts, and instead evidence that, if anything, receiving a lower accountability grade hurts schools through the labor market impacts, with suggestive evidence of a decrease in the quality of joiners relative to leavers.
These results raise an important question: When will accountability have positive impacts through its labor market effects, and when negative? The results suggest two ingredients for when accountability might lead to positive effects: in cases where (i) accountability is high-stakes enough to motivate schools to change, and (ii) principals are good leaders and have the latitude to implement changes. These hypotheses are motivated by the fact that (i) I only see positive impacts at the bottom end of the grade distribution, where the stakes were higher, and (ii) the positive turnover effects at the bottom end of the grade distribution are primarily driven by schools led by high-quality principals. Of course, these hypotheses are still speculative, and the question remains why other papers in the literature looking at high-stakes systems find dissimilar results. I cannot definitively say why the results differ, primarily because, with just a few settings, the problem is not identified. One area for further research is to investigate the reasons for the differences, and in particular, the extent to which they reflect the context and design features of the accountability system. A more thorough understanding of these features would enable policymakers to design accountability systems that better improve the performance of disadvantaged schools.
Footnotes
The author is grateful to Ran Abramitzky, Pascaline Dupas, Caroline Hoxby, and Seema Jayachandran for help and guidance and to Marinho Bertanha, Elise Dizon-Ross, Anil Jain, Jonah Rockoff, Fabiana Silva, Jenny Ying, and numerous participants at the Stanford Applied Lunch for helpful comments. Christine Cai provided excellent research assistance. All errors are those of the author. Dizon-Ross acknowledges the generous support of the Shultz Graduate Student Fellowship in Economic Policy, the endowment in memory of B.F. Haley and E.S. Shaw, the Spencer Foundation, and the National Science Foundation. Contact: University of Chicago Booth School of Business. This paper uses confidential data from the New York City Department of Education. The data can be obtained by filing a request with the NYCDOE (outlined at: http://schools.nyc.gov/Accountability/data/DataRequests).
Supplementary materials are freely available online at: http://uwpress.wisc.edu/journals/journals/jhr-supplementary.html
↵1. This type of stigma is distinct from the social stigma that the existing accountability literature has often deemed responsible for spurring test score improvements in response to low accountability grades (Figlio and Rouse 2006; West and Peterson 2006). In this paper’s categorization, social stigma would in fact impact job desirability. The direct effect would be negative, but the indirect (and thus net) effect could be positive if it spurs test score improvements that teachers value; this is thus a potential channel for job desirability to explain this paper’s findings.
↵2. Specifically, the explanation is that teachers prefer schools that (i) have improved in some way (for example, environment, performance, principal effort) in response to accountability pressures and (ii) have a higher nominal accountability grade. At the top end of the grade distribution, accountability does not bind, and so the first plays no role while the latter dominates. At the bottom end of the grade distribution, low-performing schools do improve in response to accountability, and so the first explanation—that teachers prefer schools that have improved—dominates.
↵3. On a related topic, Shirrell (2018) uses RD to look at the effect of “subgroup-specific” accountability (that is, whether a school was accountable for the performance of specific racial/ethnic student subgroups or not) on teacher turnover.
↵4. In related work, Li (2011) uses data on principal value-added to examine how accountability affects principals.
↵5. To calculate the peer horizon score, the NYCDOE assigns each school a peer index based on student demographics (elementary and K–8 schools) or past test scores of current students (middle schools). They then sort schools by peer indexes within school types to form peer groups, which consist of the 20 schools above and below a given school.
↵6. The NYCDOE was unable to provide specifics about exactly what the school improvement measures entailed, except to confirm that the measures did not involve any additional resources being given to the schools, and that there is no record of any formal guidelines for the process.
↵7. There was a pilot schoolwide bonus program instituted in NYCDOE that did tie progress reports to teacher bonuses, but it was a randomized pilot conducted with a subsample of schools and was not part of the core program instituted for all schools. As discussed in Section III.D, this program does not affect the study results. In North Carolina, teacher bonuses were guaranteed, whereas in Florida, schools received payments that could be used either for bonuses or other school improvement, at the school’s discretion (Peterson 2006).
↵8. I do not use the data from later years of the accountability program for two main reasons: first, changes were made to the program in the later years that made the thresholds less strict, including that outcomes began to depend not just on current grades but also on past performance; second, data from those years are not included in the data files I was able to obtain from the NYCDOE.
↵9. For the 2004–2005 through 2006–2007 school years, I matched teachers with classrooms based on a file maintained by the NYCDOE with student-level math and ELA teacher linkage data that have been verified by the schools. Following guidance from the NYCDOE, for school years previous to 2004–2005, I matched elementary school students to teachers based on their homeroom identifiers and middle school students to teachers based on course section identifiers.
↵10. The qualitative findings are the same if I omit the interaction term, but the estimates are less precise.
↵11. Cheng, Fan, and Marron (1997) show that the triangular kernel has boundary optimal properties, but, in practice, the results are not very sensitive to choice of kernel. Imbens and Lemieux (2008) and Lee and Lemieux (2010) recommend using a rectangular kernel and checking sensitivity to small bandwidths as an arguably more transparent method of putting more weight on observations close to the cutoff.
↵12. Rockoff and Turner (2010) only adopt this specification when using locally linear control functions and small bandwidths, which is the specification adopted in this paper.
↵13. Specifically, I assign all D (B) schools to the C/D or D/F (A/B or B/C) thresholds on the basis of whether they were above or below the median score for other D (B) schools of the same year and school type. I then estimate Equation 1, controlling for a linear trend in the accountability score for each school type–year combination and including a dummy for whether a school received the lower grade at the grade threshold it was assigned to.
↵14. The largest potential threat to identification is thus not in gaming of the score itself, but in gaming of the cutoffs by accountability officials. Although this concern is mitigated by the fact that the accountability program used round-number percentile cutoffs for thresholds and so would have had to manipulate the thresholds by large amounts to accommodate individual schools, it could still be the case that accountability officials changed the threshold to accommodate certain schools that they felt should receive given grades. However, since there are multiple schools receiving any given score, I rely on the fact that, even in the scenario in which accountability officials changed thresholds to accommodate schools, that could only be for one to two schools, while the empirical results are driven by the many other schools at the threshold.
↵15. Each column contains two cells with the results from separate regressions: the results at the top and bottom ends of the grade distribution. See Online Appendix Figure A2 for figures showing accountability scores on the x-axis and the number of elementary schools on the y-axis. The figures are shown separately for each school type and year to show the full distribution, since the thresholds are specific to the school type and year.
↵16. Charter schools did not receive accountability grades in 2007–2008; in 2008–2009, 40 charter schools received accountability grades, but are excluded because accountability may affect them differently and because I do not have other data for them.
↵17. Although the NYCDOE has administrative data on school closure dates, the NYCDOE does not track data on when school phase-outs began. However, I can proxy for the beginning of a phase-out by tagging schools that received accountability grades in one year but not in any subsequent year. This method tags 1 percent of schools, 75 percent of which closed in the ensuing five-year period, implying that this proxy method seems reasonable. Results are also robust to using eventual closing as the proxy.
↵18. Six percent of all schools and 10 percent of schools within six points of a grade threshold had begun restructuring. Since the restructuring designation was predetermined relative to accountability, it should not be related to accountability grades, and, reassuringly, Column 4 of Online Appendix Table A2 shows that there is no RD relationship between the two.
↵19. Roughly 32 percent of them have either ELA or math value-added. The reason that so many teachers do not have value-added data is that, in grades K–5, only Grades 4–5 have usable value-added data (because only Grades 3–5 are tested and one year of lagged test score is necessary for construction of value-added estimates), and in Grades 6–8, only one of a student’s approximately five teachers will be the subject teacher for math or for ELA. Thus roughly one-third of teachers should be eligible to have ELA value-added data, and one-third for math.
↵20. I cannot follow these teachers in the data: they could take teaching positions in other districts, take other nonteaching positions, stop working, or take nonteaching roles within the NYCDOE.
↵21. See Online Appendix Table A11, which uses data from the preaccountability era to understand the correlates of turnover. The negative correlation between experience and turnover could reflect the fact that there is no technical role for seniority in NYCDOE’s open market system.
↵22. School controls include: average student achievement from the previous year; the previous year’s accountability score (second year only); the percent of students who receive free and reduced price lunch, who are immigrants, and who are black, Hispanic, or Asian; the percent of teachers who are black, Hispanic, or Asian; fixed effects for school size; five-year average school turnover before the institution of the accountability system; and whether the school was part of the NYSBP. For the turnover regressions, I also include a vector for teacher-level covariates that the literature has shown to influence teacher turnover, including fixed effects for teacher experience and age, teacher education level, teacher race, and teacher gender.
↵23. The covariates used for creating the residuals are the same used in the regressions, described below.
↵24. Note that, because density varies along the x-axis, the number of schools in each dot varies; Online Appendix Figure A3 shows versions of the graphs where the number of schools per dot is fixed instead of the bandwidth per dot (so each dot represents the average score and the average turnover across ten schools).
↵25. Most of the other literature on school accountability finds effects in the other direction, where the differences in the sign of the effects likely comes from reasons that will be discussed in Section VII, but the magnitudes are substantial: Feng, Figlio, and Sass (2010) estimate that being shocked to receive an F grade in Florida increases turnover by 42 percent, Clotfelter et al. (2004) estimate that being labeled a low-performing school increases turnover by 25 percent, and Gjefsen and Gunnes (2016) estimate that the institution of school accountability increased teacher turnover by roughly 45–65 percent. One can also benchmark against other types of policies and reforms, and again the estimates do not seem unreasonable. For example, Smith and Ingersoll (2004) find that giving new teachers a mentor in their field or encouraging collaboration decreases turnover of first-year teachers by 30 percent and 43 percent, respectively.
↵26. Since regressions are run at the teacher level, using previous-year school turnover as the dependent variable means using teachers who were teaching in the school in the previous year as the sample. Note that the vector of control variables are taken from the previous year as well, so that the control variables are predetermined, not measured after the (placebo) outcome.
↵27. Using both years has the advantages of testing for baseline differences across the full sample that identifies the actual results and of having the same power for rejecting the null as the actual results. However, some of the previous papers in the accountability literature use data from the first year of the accountability system only, so that the falsification exercise is performed purely with preaccountability data. I perform that test as well. Placebo results for the second year only are the same as well. Placebo results are shown using the full sample, but the results in the nonrestructuring sample are nearly quantitatively identical.
↵28. The results using ELA value-added are statistically weaker and less robust (available upon request). Note that roughly 70 percent of teachers with any value-added data have both math and ELA value-added.
↵29. Specifically, the table presents results from estimating Equation 1 where the RD control function and the indicator for receiving the lower grade are both interacted with a teacher characteristic.
↵30. Note that, unlike Table 3, Table 4 does not seem to suggest that teachers with higher predicted value-added are more likely to leave. The discrepancy reflects the functional form, that is, the fact that I need to convert to a binary variable for the Table 4 specification and so use “above-median” for the split. If I split at other places in the distribution (for example, 25th percentile or 90th percentile), then, consistent with Table 3, teachers with higher predicted VA do appear to have higher turnover, although the results are never significant.
↵31. For experience, an indicator for having at least four years of experience is shown since the quality gains to experience generally taper after four years (Boyd et al. 2008a; Kane, Rockoff and Staiger 2008; Rivkin, Hanushek, and Kain 2005; Rockoff 2004), and, in the NYCDOE data, there is no significant relationship between experience and value-added after four years, but the results look similar with other experience measures. The correlation between teacher value-added and master’s degrees is not statistically significant in the NYC-DOE data and is generally tenuous, sometimes even negative (Rivkin, Hanushek, and Kain 2005).
↵32. Indeed, in a related context (subgroup-specific accountability policies), Shirrell (2018) finds that, when a school is held accountable for the performance of black students, black teachers’ turnover falls, with one potential mechanism being that the teachers’ motivation increases when they see that the school is trying to improve the performance of black students.
↵33. Sometimes two sections covered the same topic, so I grouped those sections together. All responses were normalized so that more positive was better.
↵34. Across the ten variables shown, no placebo results are significant at the 5 percent level, and one (the course offerings index in Panel A) is significant at the 10 percent level. Note that one potential concern with these results is that teacher survey results are an input into future accountability grades. Although they only represent roughly 3 percent of grades (teacher, student, and parent surveys together represent two-thirds of the school environment score, which represents 15 percent of the overall score), some teachers may still have answered questions strategically to try to affect their future grades, and the likelihood of answering strategically could vary with accountability grades. Since all survey questions affect the accountability grade equally, this could cause a bias in the average responses across all questions, but is less likely to cause more positive responses to some questions than others. Thus, the fact that there are much more positive responses for principal leadership than for, say, safety, likely reflects true effects, not strategic responses. The fact that there are some negative coefficients is also suggestive that teachers were not responding strategically.
↵35. The number of joiners falls by 1 percent at lower-graded schools, which is not enough to offset the fall in leavers, and thus the number of teachers increases by 2 percent (not statistically significant). Enrollment increases by 1 percent, and class size falls by 1 percent. There is no effect on principal turnover, and the turnover results are also robust to excluding principals who left (coefficient of −3.2 percent at the bottom of the grade distribution when leaving principals are excluded). All regressions are weighted by staff size.
↵36. One thing that could change is teacher expectations of future school closures. When schools close in the NYCDOE, teachers are not fired. If they cannot find permanent positions, they are given work as substitute teachers. If teachers prefer substitute work, they could stay at lower-graded schools hoping that the schools will be closed in the future. I do not view this hypothesis as very plausible because (i) all of the closures had already been announced before teachers made turnover decisions, so they would have needed to anticipate closures a full year in the future, and (ii) anecdotal evidence suggests that teachers dislike being substitutes.
↵37. Note that the regressions pool across all schools and teachers in my sample, not just those at the bottom or top of the grade distribution, but the results are different if one limits to schools in one area of the grade distribution. Principal leadership results are also similar if one does not condition on previous-year leadership. The correlations with joiner value-added are more mixed than the turnover correlations, but joiner value-added is also much less precise.
↵38. Teachers who leave the NYCDOE could be changing professions, taking a short stint away from teaching, transferring to a different district, or taking nonteaching roles within the NYCDOE.
↵39. My base specification adopts the more standard approach in the literature of using the running variable itself as opposed to its underlying components.
↵40. Specifically, I test for whether the discontinuities at all one-point intervals from the grade threshold are all equal to zero. Results are robust to different interval widths.
↵41. See Rockoff and Turner (2010) for a complete timeline of events. It is unlikely that schools knew what their 2007 accountability grades would be in advance. In April 2007, the NYCDOE informed principals of the progress report methodology and gave principals pilot progress reports based on 2005 and 2006 results. These reports did not contain letter grades, only numeric scores, and did not inform principals about how the numeric scores would be mapped to grades. The pilot reports also omitted other key information (for example, peer groups, environmental scores) that would ultimately affect the schools score. Anecdotal newspaper evidence indicates that some principals were surprised to receive low grades.
↵42. Fort, Ichino, and Zanella (2016) outline potential bias that can arise in multithreshold RD settings when all thresholds have exactly one observation located precisely on the threshold (that is, exactly one observation for which the running variable takes the value zero). Here, that concern likely does not apply since only 17 percent of the thresholds I analyze (4/24) have exactly one school at the threshold, but, to be conservative, Online Appendix Table A7 shows that the estimates are nearly identical when estimated excluding the observations at the cutoff, which is the recommended strategy to address the potential bias.
↵43. The 2009 results are shown without controls only since the versions with controls are not identified due to small sample sizes and the large number of controls.
↵44. Columns 3–8 of Online Appendix Table A8 show the main joiner and leaver results separately by grade threshold. The results are qualitatively consistent across grade thresholds, although the results, especially at the lower end of the grade distribution, should be treated as suggestive at best due to the small samples involved. Note that, when using estimated VA as the outcome variable (Columns 7–8), the D/F threshold sample has a particularly small sample. The result is shown for transparency, but the coefficient should be interpreted with great caution due to the small sample and potential for overfitting. Indeed, the result is sensitive to the covariates included: if we limit the school covariate vector to only the covariates that were unbalanced in the balance tests, the coefficient falls from 2.1 to 1.7 (p-value of .06), and when we remove all school covariates, it falls to 0.83 (p-value of 0.20).
↵45. In a related setting, Boyd et al. (2008b) exploit within-school variation to find that, when New York State introduced high-stakes testing for fourth-grade teachers, turnover among fourth-grade teachers fell.
↵46. My findings also contrast with those of Clotfelter et al. (2004), who use a difference-in-differences approach to estimate the effect of the institution of accountability in North Carolina and find that accountability accelerated teacher turnover at low-performing schools. Again, institutional differences could have played a role. For example, North Carolina linked teacher-level incentives with school accountability ratings, whereas the NYCDOE system only used school- and principal-level incentives. It is also possible that their results are partially explained by other reforms instituted concurrently with accountability, such as streamlining the process of teacher dismissals and dramatically changing salary structures and tenure requirements.
- Received October 2015.
- Accepted March 2018.









