When peer review fails…
…and accountability falls short
I recently had the worst review experience of my 25-year publishing career. I’m writing about it because I think the details are instructive — not as a story about one bad reviewer, but as a case study in what happens when a journal’s quality control fails at every stage and the available accountability mechanisms don’t provide a meaningful remedy. This story should matter to YOU.
My co-authors — four of whom are students — and I submitted a quantitative paper to SN Social Sciences, a Springer Nature journal. I received one review, which contained factual errors about cited works, fundamentally mischaracterized our stated research question, and showed no evidence of engagement with our methods, results, or appendix. When I raised these concerns, the journal offered no substantive response for nearly a month. When it finally responded on the substance, it offered a remedy that penalized us for the journal’s own failure. And when I reported the matter to the Committee on Publication Ethics (COPE), they declined to pursue it, reasoning that the journal had procedurally considered my complaint and offered a de novo resubmission route.
I want to walk through exactly what happened, because every step of this process deserves scrutiny.
The paper
The details of our paper matter only to the extent that they reveal how badly the review went wrong, so I’ll keep this brief. We study teacher–student race matching in Texas public schools. Using administrative data covering 8,691 campuses and 5.4 million students, we ask whether schools are structured such that students have a reasonable opportunity to encounter a same-race teacher — what we define as “race-match sufficiency.” This is explicitly a campus-level structural opportunity measure. It is not a claim about what happens in individual classrooms. We state this in the abstract, the introduction, the methods, the results, and the limitations. It is the central framing of the entire paper.
The review
[Addendum: I think there is a strong possibility the review was written with generative AI. The journal denies this, but then, how could they possibly know for sure? In any case, the review, along with my point-by-point rebuttal, is at the end of this post.]
The paper was rejected based on a single set of reviewer comments, labeled “Reviewer 1.” Here are a few examples of how problematic this review was.
A factual error used to justify rejection. The reviewer claims that Gershenson et al. (2022) — a key paper we cite — simply uses data from the 1986 STAR project, and that the situation is therefore “completely different” from ours. This is wrong. Gershenson et al. links STAR participants to long-run administrative records to track outcomes over decades, and it replicates its main findings using public school data from another state — an entirely separate, non-experimental dataset. We cite this paper for a specific, clearly stated reason: the empirical finding that even one same-race teacher is associated with improved long-term outcomes like graduation and college enrollment. We don’t claim our study replicates their design. The reviewer appears to have either not read the cited paper or not read our discussion of it, and then used this mischaracterization as a central plank of the case for rejection.
A fundamental misread of the research question. The reviewer writes: “The only adequate level is what happens in a class. Everything else is too distal, unreliable and invalid.” But our paper is not trying to measure what happens inside classrooms. It is measuring whether schools are structured to make same-race teacher contact possible — a different question entirely, and one we articulate repeatedly throughout the manuscript. We write, for example, that our metric “should be understood as a campus-level measure of structural opportunity rather than a direct observation of classroom assignment.” Structural and institutional analyses are a well-established complement to individual-level studies across the social sciences. The reviewer is not identifying a flaw in our work; they are rejecting our paper for not being a different paper.
Vague, unsubstantiated criticism with no actionable content. Several comments amount to a word or a sentence fragment. One critique reads, in its entirety: “Specify.” Another dismisses an extensive, evidence-based appendix with “This is not really convincing” — no indication of what is unconvincing, what evidence would satisfy the concern, or what specific assumption is in question. These are not substantive critiques. They give no clear sense of what evidence, argument, or revision would address the concern.
The review also contains no engagement with our mathematical framework, our model selection procedure, our diagnostic tests, or our appendix. It raises concerns about within-school teacher assignment — which we address in our limitations section. It dismisses our quantitative findings because the directional ranking of racial groups is “not unexpected; you don’t need sophisticated analyses to show this.” The purpose of empirical social science is not only to discover surprises but to quantify magnitudes and clarify patterns that intuition alone cannot resolve, and the reviewer’s own conclusion — that the paper’s value “only lies in the method, but absolutely not in the content” — is internally contradictory for a paper whose content is the application of that method to a substantive policy question.
What happened next
I sent a detailed, point-by-point rebuttal to the journal on the day of rejection, March 6, 2026, copying the publisher and the journal’s executive editors. I did this because I believed the problems extended beyond the handling of a single manuscript.
For ten days, no one replied. I followed up. Over the next two weeks, there were acknowledgments and forwarding — a Springer Nature project manager passed my complaint to the journal, an assistant editor said she would escalate to the managing editor, my message was forwarded to a senior editor — but no one addressed the substance of my concerns.
On April 3 — nearly a month after my initial complaint — the senior editor responded on the substance. The response acknowledged that the Associate Editor (the handling editor) had conducted the review, confirmed this was permitted under journal policy, and offered one remedy: I could resubmit the manuscript as an entirely new submission, going through the full editorial workflow from scratch.
I want to be clear about what this means. The manuscript had not changed. What failed was the review. But the proposed resolution required my team to absorb the full cost of that failure — another round of queue time, another review cycle, potentially months of additional delay, on top of the months already lost. This is not a resolution. It is a reset that treats the authors as though they bear responsibility for a process that broke down on the journal’s side.
Timeline
March 6: Rejection received; rebuttal sent same day to journal, publisher, and executive editors. March 6–16: No response. Follow-up sent. Shortly after rejection: Complaint filed with COPE. March 17: Springer Nature project manager forwards complaint to journal. March 19: Assistant editor promises escalation to managing editor. March 25: Follow-up; message forwarded to senior editor. April 3: Senior editor responds: resubmit as new manuscript. April 7: COPE declines to pursue. April 8: Journal claims two reviews existed; reiterates individual-level critique.
I wrote back asking the editorial team to do something more basic: evaluate the rebuttal I had already provided, assess whether the review actually supports rejection, and make an expedited decision on the manuscript. If they determined after that evaluation that additional review was warranted, I would accept that — but the starting point should be what had already happened, not a blank slate that erases the record of what went wrong.
The story shifts
I was only ever sent one set of reviewer comments, labeled “Reviewer 1.” On April 8, the journal wrote again, this time stating that the paper had been reviewed by two people: one external reviewer and the Associate Editor. This was the first time an external reviewer had been mentioned. The preceding month of correspondence had discussed only the Associate Editor’s role.
I cannot determine from the materials I received whether a second review existed but was not shared with us, or whether the account changed after the fact. Either way, the situation is problematic. If two reviews existed, the journal’s failure to provide one of them to the authors is itself a serious editorial lapse. If the account changed, that raises a different set of concerns. What I can say is that the facts as presented to me shifted over the course of these exchanges in ways that did not inspire confidence.
The April 8 response also offered a new substantive rationale for the rejection: that our paper uses aggregate data to estimate mismatch probabilities at the individual level and does not demonstrate that these estimates correspond to actual individual-level mismatches. This is the same mischaracterization the reviewer made — and that we address repeatedly in the paper. We do not claim to measure individual-level mismatch. We define and measure a structural property of schools. This concern is stated and addressed in our manuscript. That the journal, after weeks of reviewing my complaint, would offer as its definitive reason for rejection a critique that the paper explicitly anticipates and discusses is difficult to reconcile with the manuscript’s repeated statement that the measure is campus-level structural opportunity rather than individual-level mismatch.
COPE
I reported the matter to the Committee on Publication Ethics shortly after the rejection. A COPE officer acknowledged my complaint and asked for additional details, which I provided. I continued to update COPE as the situation with the journal developed. On April 7, the Facilitation and Integrity subcommittee informed me that they could not pursue the case. Their reasoning was that COPE’s scope is limited to editorial process rather than editorial judgment, and that the journal had “taken steps” and “offered an avenue” by allowing a de novo resubmission.
There are two things to say about this. The narrow factual point is that COPE reviewed the matter only for procedural handling and declined to intervene in editorial decision-making. That is a defensible institutional boundary. But the normative consequence is that it leaves authors entirely exposed when a journal can satisfy procedural expectations with a restart-from-scratch option that imposes all costs on the authors. A journal can reject a paper on the basis of a demonstrably flawed review, decline to evaluate its own failure, and offer a “remedy” that costs the authors months of additional delay — and this is treated as the process working. That is a very low bar if the aim is to provide meaningful recourse when peer review quality breaks down.
Why this matters
Every working academic has received bad reviews. It comes with the territory, and most of the time you absorb it and move on. I am not writing this because my paper was rejected. Papers get rejected.
I am writing this because the system failed at every level — the review, the editorial handling, the complaint process, and the external accountability mechanism — and at no point did anyone with authority evaluate whether the review actually supported the decision. That is the part I cannot let go of.
When a review contains factual errors, ignores the paper’s stated research question, and engages with none of the methods or results, the appropriate response from an editorial team is to get a new review. It is not to ask the authors to start over. It is not to offer shifting accounts of who conducted the review. And it is not to rearticulate the reviewer’s flawed critique as though it were the journal’s own considered judgment.
This matters especially for early-career researchers and students. Four of my co-authors on this paper are students. They invested months in this work. They are watching how the profession handles quality control, and what they are learning is that there is none — or at least, none that functions when it’s most needed. I have 25 years of experience and the professional standing to push back publicly. Most people in this situation don’t, and the system relies on that asymmetry.
I do not expect peer review to be perfect. I expect it to be minimally competent, and I expect journals to take responsibility when it isn’t. SN Social Sciences failed on both counts. I won’t be submitting there again, and I’d encourage colleagues to weigh this experience when deciding where to send their work.
Your Neighbor,
Chad
p.s. The review, and my point-by-point response is below. The journal never engaged with any of the substance of my response. Proceed at your own peril.
Dear SN Social Sciences:
I am writing in response to the review (below) of our manuscript “Sufficiency in Teacher-Student Race Matching.” I am copying the publisher and executive editors on this correspondence because I believe this situation raises issues that extend beyond the handling of a single manuscript. I will address the reviewer’s comments individually below, but I must first address the review process itself.
To be direct: this review is not acceptable. It does not meet the minimum standard of scholarly engagement that authors are entitled to expect when they submit work to your journal. The reviewer makes factual errors about cited works, repeatedly mischaracterizes my paper’s explicitly stated research questions, offers vague criticisms with no actionable specificity, and recommends rejection on grounds that reveal a fundamental failure to engage with the manuscript’s actual content. Several comments suggest the reviewer either did not read the paper in its entirety or did not read it with care. This is not a matter of scholarly disagreement — it is a breakdown of your review process. Authors invest months of rigorous work in manuscripts and are entitled to reviews that reflect a commensurate level of care. What we received instead is a review that could have been written by someone who skimmed the abstract and flipped through a few pages. It’s unscionable that this is the review your have provided to my coauthors and me.
I also note that the review exhibits features consistent with AI-generated text: superficial engagement across many points without depth on any, confident assertions unsupported by evidence or citation, and a pattern of gesturing at critique without performing the underlying intellectual work. I believe that, separate from handling of my manuscript, your editorial team has an obligation to investigate. If this review was produced or substantially assisted by AI, that is a serious breach of integrity.
POINT-BY-POINT RESPONSE
Reviewer: “Research consistently shows” is “far too assertive.”
This claim is overstated by the reviewer. Our manuscript acknowledges mixed results for Hispanic students in the very next sentence (citing Egalite et al., 2015, finding no significant effects in Florida). The overall weight of evidence, synthesized in meta-analyses and review articles such as Redding (2019), does consistently support positive effects of race-matched assignment for Black students and generally positive effects across groups. We are willing to soften the language marginally, but the reviewer’s assertion that “there are certainly studies that find the opposite” is offered without a single citation. In a peer review, the burden of evidence applies to reviewers as well.
Reviewer: “Overall, race matched… Specify.”
This comment is too vague to be actionable. The sentence in question summarizes Redding (2019), itself a comprehensive review article. The reviewer does not indicate what additional specificity is sought. We cannot meaningfully respond to a one-word directive.
Reviewer: The only adequate level is what happens in a class. Everything else is “too distal, unreliable and invalid.”
This comment reveals a fundamental misreading of the paper. Our manuscript states — explicitly and repeatedly — that race-match sufficiency is a STRUCTURAL OPPORTUNITY MEASURE AT THE CAMPUS LEVEL, not a measure of realized classroom contact. This is articulated in the abstract, the introduction, the methods section, the results discussion, and the limitations. We write, for example, that our metric “should be understood as a campus-level measure of structural opportunity rather than a direct observation of classroom assignment.” The reviewer’s objection amounts to criticizing the paper for not being a different paper. Structural and institutional analyses are a well-established and essential complement to individual-level studies. Dismissing all non-classroom-level analysis as “invalid” reflects an unjustifiably narrow view of what constitutes legitimate educational research.
Reviewer: Why is Hispanic and Black capitalized and white not?
This follows standard contemporary convention in social science, consistent with APA style guidelines (7th edition) and the practice of major journals in the field.
Reviewer: “As a defensible… This is not really convincing.”
The reviewer offers no explanation of what is unconvincing or why. Our appendix provides extensive justification for the exposure parameter tau, grounded in Texas Administrative Code certification requirements, TEA assignment charts, published district scheduling documents from Fort Worth ISD and Houston ISD, and a federally funded IES study on departmentalization. A reviewer who finds this “not convincing” has an obligation to specify what evidence would be, or what specific assumption is objectionable. Without that, this comment is not a substantive critique.
Reviewer: Gershenson et al. (2022) uses data from the STAR project (1986) and the situation is “completely different.”
The reviewer is correct that the primary identification strategy in Gershenson et al. (2022) leverages the Tennessee STAR experiment. However, the reviewer’s characterization is misleading in ways that matter, and the conclusion drawn from it — that our citation is therefore inappropriate — is wrong. First, Gershenson et al. link STAR participants to long-run administrative records (high school graduation via the National Student Clearinghouse, college enrollment data), so this is not merely a study of 1986 survey responses — it is a longitudinal analysis tracking causal impacts over decades. Second, Gershenson et al. explicitly replicate their main findings using North Carolina public school administrative data, a completely separate non-experimental dataset, demonstrating that the results generalize beyond the STAR context. Third, a footnote in that paper (fn. 3, p. 303) specifically cites a working paper using Texas data that documents long-run race-match effects among Hispanic students — directly relevant to the population we study. The reviewer either missed or ignored these features of the paper they claim to have read. More fundamentally, we cite Gershenson et al. for a specific, clearly stated purpose: the empirical finding that even one same-race teacher is associated with measurable improvements in long-term outcomes (graduation, college enrollment). This motivates our sufficiency threshold of srt >= 1. We do not claim that our study replicates their design, operates at the same level of analysis, or applies to the same population. The reviewer conflates “citing a study for empirical motivation” with “claiming identical conditions hold.” This is a basic misunderstanding of how citations function in academic writing.
Reviewer: Pages 14–17 should be explained in layperson’s language.
Our manuscript is submitted to a social sciences journal whose readership includes quantitative researchers. The methods section uses standard statistical terminology (logistic regression, natural splines, random intercepts, AIC) that is appropriate for this audience. The reviewer does not identify any specific passage that is unclear. That said, we are open to adding additional plain-language signposting if the editor considers this useful.
Reviewer: Are there differences between Black, Hispanic, and white teachers regarding the subject they teach?
This is a fair question, and one we address in our limitations section, where we note that Hispanic teachers may be concentrated in particular subjects or program areas (citing Lindsay et al., 2021 and Redding, 2019) and acknowledge that campus-level data cannot capture within-school assignment processes. The reviewer appears not to have read this discussion.
Reviewer: The ordering of white, Black, Hispanic is “not unexpected” and doesn’t need “sophisticated analyses.”
This comment misidentifies our contribution. The rank ordering is not the finding. Our contribution is the quantification of the nonlinear S-shaped relationship between student concentration and sufficiency, the identification of group-specific threshold effects, the demonstration that identical demographic compositions yield vastly different sufficiency levels depending on racial group, and the estimation of how structural features like school level and enrollment mediate these patterns. None of these results are obvious from intuition, and none are obtainable without the modeling framework we develop. Dismissing quantitative analysis because the direction of an effect is “not unexpected” would invalidate most of empirical social science.
Reviewer: The urban vs. non-urban finding “of course also holds.”
Same response. That urban districts have higher sufficiency than non-urban ones is directionally intuitive. The contribution is in quantifying the magnitude, showing how it varies by race (the gap between white and Black students widens outside of major cities), and identifying which specific district types drive the pattern. The reviewer’s dismissal again confuses direction with magnitude and mechanism.
Reviewer: “All the analyses have not been performed at the right level.”
We refer back to our response above regarding the campus-level structural opportunity framework. This objection has been addressed at length. The reviewer is entitled to prefer a different research question, but that preference does not constitute a valid critique of the paper we wrote.
Reviewer: “It is impossible to draw this conclusion, as the wrong data have been analyzed.” And: to have an impact, a student and teacher should be “together for longer periods of time, months or even years.”
The first assertion repeats the same mischaracterization addressed above. The second is a theoretical claim offered without any supporting evidence. It is directly contradicted by the published literature: Dee (2004) documents effects from single-year, randomly assigned teacher-student pairings, and Gershenson et al. (2022) finds that a single same-race teacher in elementary school is associated with improved high school graduation and college enrollment. The reviewer is asserting a position against the weight of peer-reviewed evidence without citing a single study. Furthermore, we remind you of our basic question in this paper, which is not about the effect of race match, but rather, the extent of structural opportunity for it. This is stated over and over and over, and ignored by the referee.
CONCLUSION
The reviewer’s recommendation to reject rests on two pillars: (1) that campus-level structural analysis is inherently invalid because only classroom-level data matter, and (2) that our use of the supporting literature is inappropriate. The first position reflects a narrow methodological preference, not a flaw in our work; our paper’s structural-opportunity framing is explicitly stated, carefully justified, and well-precedented in educational research. The second is grounded in factual errors about the cited works.
The review contains no engagement with our statistical framework, our model selection procedure, our diagnostic tests, or our appendix. Multiple comments address topics (subject-area concentration of teachers, within-school sorting) that are already discussed in our manuscript. Several criticisms amount to single phrases (“Specify,” “not really convincing”) without explanation. The final verdict — that the paper’s value “only lies in the method, but absolutely not in the content” — is internally contradictory for a paper whose content is the application of that method to a substantive policy question.
Authors who submit to your journal are entitled to feedback that demonstrates the reviewer actually read the manuscript, engaged with its arguments, and understood its stated research questions. This review fails on all three counts. A journal that accepts this standard of review — and bases editorial decisions on it — has a quality control problem, and that is why I am bringing this to the attention of the publisher and executive editors.
My team has now wasted several months on the basis of a review that contains factual errors, misrepresents the paper’s stated aims, and shows no evidence of having engaged with the methods, results, or appendix. That delay is not recoverable. You now have an obligation to evaluate this rebuttal itself, determine that the review does not support rejection, and expedite a decision on the manuscript. I expect you to take responsibility for the failure here and act accordingly.
