Science is supposed to be self-correcting. But what happens when the journals entrusted to uphold this principle become its gatekeepers? We have collaborated for almost a decade in an effort to correct the fraud permeating throughout the most influential depression study ever conducted. And what our efforts reveal is that the mechanisms for accountability in medical research are broken.
While the risk of fraud in pharmaceutical company-funded psychiatric research is now widely acknowledged,1-4 this fraud occurred in a taxpayer-funded study presumed to be trustworthy. This assumption made the fraud harder to challenge—and has allowed its flawed findings to be widely accepted and embedded into everyday medical practice.
We have collaborated on writing this piece. However, the first part of this effort to correct the STAR*D record stretches back two decades, when one of us—Ed—first discovered irregularities in the reported results from the STAR*D study. As such, this blog toggles between an initial “I” part and a subsequent “we” part to reflect this history.
The Sequenced Treatment Alternatives to Relieve Depression (STAR*D) study is the largest, longest, and most expensive antidepressant trial ever conducted. Costing American taxpayers 35 million dollars, STAR*D was designed to identify the best next-step treatment for the many depressed patients who fail to respond adequately to their first and/or subsequent antidepressant medication trial.
In order to mirror real-world clinical practice, STAR*D used an open-label research design (ie, patients knew what drug(s) they were taking), enrolled patients seeking routine medical or psychiatric care, and included individuals with comorbid medical and psychiatric conditions, thereby maximizing the generalizability and clinical relevance of STAR*D’s findings.
STAR*D enrolled 4,041 patients who screened positive for depression and then started on Celexa, a selective serotonin reuptake inhibitor (SSRI). Patients who did not respond to Celexa in step-1 of the trial selected acceptable treatment options for randomization in steps 2 through 4, to “empower patients, strengthen the therapeutic alliance, optimize treatment adherence, and improve outcome.”5,p.483 Cognitive therapy was also offered as an option in step-2; however, too few patients chose it, and it was therefore excluded from the major step-2 articles published in The New England Journal of Medicine (NEJM).6,7
Besides allowing patient choice of treatment options in steps 2-4, STAR*D implemented several other measures designed to optimize clinical outcomes. These included a rigorous visit-by-visit symptom monitoring system (SMS) to guide aggressive medication management to “ensure that the likelihood of achieving remission was maximized and that those who did not reach remission were truly resistant to the medication.”8,p.31
Patients and family members also received education during each visit based on the neurochemical imbalance theory of depression. This included “a glossy visual representation of the brain and neurotransmitters” and “the mechanism of action” of each patient’s current medication and emphasized that “depression is a disease, like diabetes or high blood pressure…The educator should emphasize that depression can be treated as effectively as other illnesses.”9,p.6-8
In 2006, The American Journal of Psychiatry (AJP) and NEJM published STAR*D’s seven primary articles. Two decades later, these high-impact articles continue to profoundly influence antidepressant medication treatment guidelines worldwide. The summary article reported that 67% of the depressed patients achieved a complete remission after up to four sequential medication trials,10 a figure that continues to shape clinical guidelines and expectations among clinicians and the public regarding the likelihood of treatment success.
But, as I (Ed) soon discovered, this was all a lie.
Intrigued after reading a front-page Washington Post story on STAR*D, I downloaded the step-1 Celexa article in the spring of 2006.8 Careful reading revealed tells of research fraud inflating Celexa’s remission rate.
The first tell was the novel research design for assessing antidepressant effectiveness. I coined it the tag-you-are-healed design.11 Using STAR*D’s SMS algorithm, many patients were declared remitted after just 4 to 6 weeks on Celexa and promptly removed from the step-1 Celexa subject pool and shifted into the 1-year follow-up care subject pool—even though the acute phase of Celexa treatment lasted for up to 14 weeks.
This sleight-of-hand ensured that Celexa remitters could never un-remit, thereby exploiting the well-known fact that depression ebbs and flows for the vast majority of sufferers. To discover STAR*D’s true remission rate, I would have to wait for the 12-month follow-up care data to be published.
The second tell was STAR*D excluding from data analysis 234 patients who quit their Celexa therapy during the first 2 weeks, even though the protocol specified that such patients should be counted as treatment failures. The STAR*D investigators never disclosed that these patients were taking Celexa when they dropped out. This sleight-of-hand alone increased Celexa’s reported remission rate from 25.4% to 28%.
We later learned that this trickery served an additional, more troubling purpose. In their step-1 article, the STAR*D investigators asserted that, “There were no suicides in the 2,876 participants in this acute-phase Celexa study.”8,p.33 This assertion though was true only by excluding from analysis the 234 early-dropout Celexa patients when they were most at risk for becoming suicidal.
In 2009, the STAR*D investigators published an analysis of emergent suicidality during Celexa treatment.12 They reported that 7% of patients who were not suicidal at baseline reported emergent suicidality during their first two weeks of Celexa treatment. The investigators then go on to acknowledge that “One of the participants who did not return after baseline was reported as a likely suicide secondary to falling off his apartment complex roof 10 days after Celexa was prescribed.”12,p.69
How clever on the STAR*D PIs’ part. This protocol violation was a twofer. It inflated the remission rate while covering up a suicide.
Then the fog horn blared . . . Troubled, I downloaded STAR*D’s 2004 Research Design article and discovered that the second measure used to report Celexa’s remission rate was specifically excluded from being used as such because it was non-blinded and central to STAR*D’s SMS-algorithm guiding treatment.13 This protocol violation bumped up Celexa’s remission rate further to 33%.
The fog thickened…In March 2006, I downloaded STAR*D’s step‑2 drug‑switch and Celexa‑augmentation articles from NEJM.6,7 These two comparative-effectiveness studies were STAR*D’s raison d’être and the justification for publication in America’s preeminent medical journal. Yet the two NEJM articles obscured the number of STAR*D patients progressing from step‑1 to step‑2 treatment.
STAR*D enrolled 4,041 patients. However, only 3,110 met the depression severity criteria for major depressive disorder (MDD) when assessed by expert raters blind to patients’ study status (entry, exit, or follow-up), and therefore met STAR*D’s inclusion criteria for data analysis.8 Figure 1 The remaining 931 patients stayed in STAR*D because the investigators did not learn patients’ depression severity scores at entry into the study until after the 12-month follow-up care phase ended and the study’s blind broken.
The step-1 Celexa flowchart reports that 1,201 eligible depressed patients advanced to step‑2 treatment.8 Figure 1 But when combining the step‑2 treatment groups reported in NEJM and AJP the total is 1,439. This means that without clear public disclosure, the STAR*D investigators included 238 patients in their analyses who failed to meet inclusion criteria due to insufficient depression severity when first enrolled into the study.
Darkest before dawn…In November 2006, I downloaded STAR*D’s summary article, which reported remission rates for steps 1–4 treatments and the 12-month follow-up care data using survival-analysis.10
The key survival analysis figure was nearly impossible to interpret (see below).
Like most readers, I tried to decipher from the graph how many remitted patients survived follow-up care without relapsing and/or dropping out. After weeks pondering the squiggly lines, I realized the answer was in plain sight—just not in the graph—but the table above it.
By shifting focus, I learned that only 108 (2.7%) of the original 4,041 STAR*D patients had a remission after up to four increasingly aggressive antidepressant medication trials, and survived follow-up care without relapsing and/or dropping out.
The failure of STAR*D’s treatment model—sequential trials of increasingly aggressive antidepressants and their combinations—was hiding in plain sight. The investigators simply never drew attention to the abysmally low stay-well rate. Instead, they promoted a 67% cumulative remission rate—a false claim that collapses under scrutiny yet has been repeated countless times in medical journals, the media, and during patient consultations for the past 20 years—and sustained by the silence of the two journals that gave STAR*D global credibility.
With colleagues in 2008, I submitted our first STAR*D reanalysis to NEJM. This manuscript received a form-letter rejection despite documenting that only 108 of STAR*D’s 4,041 patients had a sustained remission without relapsing and/or dropping out.14
In April 2009, I submitted our deconstruction of STAR*D’s findings to AJP documenting numerous protocol violations, each one inflating further STAR*D’s reported success rates.
This submission included a letter to Dr. Freedman, AJP’s Editor-in-Chief, highlighting how an unbiased presentation of STAR*D’s findings discredits the American Psychiatric Association’s guidelines for treating depression.15 I informed Dr. Freedman that “only 108 of STAR*D’s 4,041 patients (2.7%) who were initially started on Celexa in step-1 had a remission (after up to four treatment trials) and during follow-up neither dropped out nor relapsed.”
I offered to send Dr. Freedman my email exchanges with Dr. Wisniewski, STAR*D’s chief biostatistician, confirming the accuracy of our analyses.16,17 Dr. Freedman never requested the emails and our submission received another form-letter rejection without peer review.
In July 2009, my colleagues and I submitted our manuscript titled “Efficacy and Effectiveness of Antidepressants: Current Status of Research” to Psychotherapy and Psychosomatics (PPS). The first section reviewed prior placebo-controlled antidepressant studies, and the second deconstructed STAR*D’s findings. It was published in July 2010, and became PPS’s most frequently downloaded article for 22 consecutive months.18
Soon after the PPS publication, I started collaborating with University of Connecticut researchers who had obtained STAR*D’s patient-level dataset from NIMH.
In June 2012, we submitted to NEJM a reanalysis of the dataset adhering to the NIMH-approved protocol I obtained through one of three Freedom of Information Act (FOIA) requests.19 The cover letter stated:
“Our paper warrants publication in NEJM because it reanalyzes the STAR*D dataset according to its research protocol. This analysis documents that the reported benefits of antidepressant treatment in the first step have been greatly exaggerated…Furthermore, our analyses found STAR*D’s researchers violated the protocol inclusion and exclusion criteria, resulting in patients who were not depressed at the start of the study being included in subsequent phases.”20
Once again, NEJM rejected our submission without peer review.
In 2018, I teamed up with Dr. Jay Amsterdam, Emeritus Professor of Psychiatry at the University of Pennsylvania, Perelman School of Medicine, who had directed its Depression Research Unit for more than two decades, starting in 1980.
Jay had an illustrious career at Penn in research psychiatry. Despite gradually losing his eyesight, he published more than 285 peer-reviewed articles and 6 textbooks on treatment resistant depression.21-26 Jay’s research includes prospective and retrospective studies consistently showing a 20% to 30% increase in drug tolerance (ie, loss of antidepressant effectiveness) with each increase in the number of prior antidepressant treatments taken by the patient.27-34
Jay’s provocative findings, published over a 25-year period, were mirrored in the progressive loss of effectiveness during the STAR*D study when patients were exposed to repeated—and increasingly aggressive—antidepressant medication trials in steps 1-4. What Jay’s research, among others,35-37 suggests is that the repetitive administration of antidepressant medications over time often results in an iatrogenic (ie, doctor-induced) treatment-resistant depression.
This concept—that a medication once helpful can become iatrogenic—is not foreign to medicine. Antibiotics lose effectiveness with overuse. Benzodiazepines and stimulants lose potency as tolerance kicks in. Antidepressants appear to follow the same trajectory.35-40 What begins as treatment over time becomes part of the problem—transforming episodic depression into a chronic, treatment-resistant illness. And when the response to that loss of effectiveness is simply more medication, the cycle far too often becomes self-perpetuating.
In March 2019, we posted a “Call to action: RIAT reanalysis of the sequenced treatment alternatives to relieve depression (STAR*D) study” in The British Medical Journal (BMJ).41 In addition to documenting the STAR*D investigators’ numerous protocol violations, we found a jaw-dropping new one:
STAR*D included in data analysis 117 patients who were in protocol-defined remission before starting their step-2 treatment.
This shameless deception violated both the protocol—and common sense—to double-count remitted patients. We only discovered this flagrant protocol violation because we now had STAR*D’s patient-level dataset—and bothered to look.
As part of our Restoring Invisible and Abandoned Trials (RIAT) grant application,42 we notified the STAR*D PIs of our Call to Action, documented their protocol violations, and announced our intention to reanalyze the dataset with fidelity to the original NIMH-approved protocol. We invited the PIs to conduct the reanalysis themselves but they declined.
With RIAT funding, we hired two Penn graduate students (now Drs. Colin Xu and Thomas Kim) to analyze the dataset. Along with Dr. Irving Kirsch, Associate Director of the Program in Placebo Studies at Harvard University, we published in BMJ Open the article titled, “What are the treatment remission, response and extent of improvement rates after up to four trials of antidepressant therapies in real-world depressed patients? A reanalysis of the STAR*D study’s patient-level data with fidelity to the original research protocol.”43
Prior to publication, the BMJ Open editor invited the STAR*D PIs to respond publicly to our reanalysis and documentation of their protocol violations. Again, they declined to do so.44 In contrast to the STAR*D investigators’ claim of a 67% cumulative remission rate, we found it was only 35% when adhering to the protocol—and this was before assessing how many of those who remitted actually stayed well.
The importance of our BMJ Open article was quickly recognized by the editors of Nature Mental Health and Psychiatric Times. Two months after its publication, Nature Mental Health’s Editor-in-Chief wrote:
“Some of the bedrock of clinical wisdom in psychiatry has begun to erode…Recent work has challenged long-held and widely accepted statistics about the efficacy of antidepressants. A reanalysis of the patient-level dataset from the STAR*D study demonstrated that, in contrast to the previously reported cumulative remission rate of 67%, after examination of key methodological deviations from the research protocol and related publications, the actual rate of remission was only 35.0% of participants — a precipitous drop in what has become a benchmark estimate for antidepressant efficacy.”45
This was followed by the Editor-in-Chief of Psychiatric Times, writing in a cover story titled, “STAR*D Dethroned?”:
“Recently, a provocative and well-researched publication in the BMJ reanalyzed the original raw data obtained from the NIMH and challenged the conclusions of the STAR*D publications…For us in psychiatry, if the BMJ authors are correct, this is a huge set back, as all of the publications and policy decisions based on the STAR*D findings that became clinical dogma since 2006 will need to be reviewed, revisited, and possibly retracted.”46
Soon thereafter, Drs Nicolas Badre and Jason Compton published an article in Psychiatric Times titled, “STAR*D: It is Time to Atone and Retract.”47 The article was written in response to an exchange in Psychiatric Times detailing our allegations of protocol violations48 and the STAR*D PIs’ response.49 Drs. Badre and Compton’s conclusion echoed their title. If our analyses were correct, the time had come for the American Psychiatric Association’s marquee journal to Atone and Retract STAR*D.
Next, we reanalyzed STAR*D’s drug-switch and Celexa-augmentation articles NEJM published and posted our RIAT re-analyses on the medRxiv preprint server. Drs. Martin Plöderl50 and Kevin Kennedy51 added their expertise to our efforts. Each preprint included the statistical code used to analyze the dataset thereby allowing researchers worldwide to confirm or challenge our findings.
Our February 2025 drug-switch reanalysis, submitted to NEJM, found:52
Our October 2025 NEJM submission reanalyzing STAR*D’s Celexa-augmentation article found:54
Both of our 2025 RIAT reanalysis submissions were accompanied by letters to Dr. Eric Rubin, NEJM’s current Editor-in-Chief.55,56 Key points in the letters were:
The STAR*D trial is the most influential depression treatment study ever conducted and continues to shape clinical practice guidelines worldwide. Yet its core findings are unreliable due to systematic protocol violations. The Committee on Publication Ethics (COPE) guidelines state that retraction is warranted in cases of pervasive error and/or research misconduct that render findings invalid and misleading.58 Both apply here:
Given our correspondence, Dr. Rubin has now twice made grave, ethically-flawed editorial decisions: first, by refusing to send our RIAT re-analyses for peer review, and second, by refusing to retract the articles NEJM published despite compelling evidence that their conclusions are invalid. The question now is not whether STAR*D’s central claims are invalid, but whether the journals that gave them global authority will do what integrity demands: Atone and Retract.
NEJM’s incompetent peer review 20 years ago led to a two-decade long delay in the public learning that (1) outcomes from both drug-switch and Celexa-augmentation treatments were far worse than claimed, and (2) patients were far more likely to become suicidal than remit and stay well.
I (and others),59-62 figured out that STAR*D was riddled with scientific mischief when first published. NEJM’s editors and peer reviewers would have seen the same problems I did, had they simply read STAR*D’s 2004 Research Design article15 and compared it to the STAR*D investigators’ two submissions. NEJM had the power to demand answers regarding the STAR*D protocol violations two decades ago—but failed to do so.
Dr. Rubin faced a choice when confronted with our 2025 submissions: expose NEJM’s editorial failures starting in 2006—including its rejection without peer review our 2008, 2012, and now both 2025 submissions—or bet that this fraudulent scheme NEJM became party to would remain largely hidden, as it has for 20 years.
Regrettably, Dr. Rubin chose the latter, effectively deciding that the NEJM-enabled fraud would stay buried—patient welfare be damned. Dr. Rubin knew that as NEJM’s Editor-in-Chief, his decision was final with no avenues for appeal.
The Hippocratic Oath emphasizes that medicine’s first duty is to do no harm, and fulfilling this duty requires honesty in medical research. Until the scientific publishing and editorial process permits independent correction of the scientific record—no matter how uncomfortable the truth may be—medical science will continue to protect its academic and corporate institutions at the expense of its integrity and the public’s safety.
For these reasons, we recommend the establishment of an independent, financially-unconflicted ombudsman to evaluate allegations of research misconduct. Patient safety and the advancement of evidence-based medicine requires a neutral arbiter, free of institutional and financial conflicts of interest, with the authority to investigate misconduct and mandate corrections or retractions when warranted.
In the meantime, NEJM should be shamed into cleaning up the mess they made. NEJM’s prestige—and its worldwide platform—were used to legitimize and disseminate a profoundly flawed and misleading narrative about antidepressant medications.
This same platform should now be used to inform clinicians and the public about these drugs’ true risks and their low likelihood of sustained benefit, as documented by our RIAT re-analyses. Failure to do so—despite nearly two decades of repeated notification that NEJM’s STAR*D publications contain demonstrably false and clinically-consequential claims—should expose the journal to legal consequences for knowingly allowing these claims to remain uncorrected.
For better or worse, high-impact medical journals are the gatekeepers of evidence-based medicine, where science is presumed to reign supreme. In practice, these journals supply the certified, high-strength concrete foundation of America’s 5.3 trillion dollar healthcare industry—accounting for approximately 18% of U.S. gross domestic product in 2024. What they publish determines which treatments are deemed legitimate, and therefore reimbursed by private insurers, Medicare, and Medicaid.
Yet from these institutions to whom we have entrusted such extraordinary influence, as a society, we demand remarkably little…
Like the president in the Supreme Court’s Trump v. United States decision, medical journals have broad legal immunity for official acts. They are essentially untouchable because the law treats them as protected speakers exercising editorial discretion, not as institutions whose publications are relied upon in clinical decision-making. Even after repeated, well-documented proof of research fraud, journals have no duty to investigate or correct the scientific record.
Even the First Amendment though does not protect yelling fire in a crowded theater and then, after learning the claim is false, choosing silence as people experience ongoing harm.
Medical journals do exactly that when they refuse to investigate and correct fraudulent research despite clear notice—converting editorial discretion into a mechanism for perpetuating harm and failing in their societal mission to provide a level platform for scientific inquiry and debate.
We would be far further ahead today in treating depression if STAR*D’s findings had been reported truthfully when first published two decades ago. While our preprints have not yet been certified by peer review, if verified they suggest very different evidence-based guidance for treating depression than that promoted by the STAR*D investigators.
Given the increased risk of emergent suicidality—and the low likelihood of sustained benefit—following a failed SSRI trial, the more responsible course is to stop and reconsider versus follow STAR*D’s try-and-try-again ad infinitum model of care with ever more aggressive antidepressant medications and their combinations.
When these findings are finally confirmed, the greatest failure will not be STAR*D’s flawed science, but the decisions by AJP and NEJM’s Editor-in-Chiefs to let its claims continue shaping clinical practice long after their validity was called into question.
It’s crazy. As a society, we’ve allowed an industry which is vital to America’s health to be legally immune from the adverse consequences of their actions. Legal reforms—such as a post-notification duty to respond, a recklessness standard for willful inaction, and injunctive relief to compel corrections—are needed to restore accountability without compromising scientific inquiry and debate.
For nearly two decades, AJP and NEJM have refused to correct the fraudulent STAR*D record and patients continue to be collateral damage.
When medical journals choose silence over truth, patients pay the price—and the law calls it editorial freedom. It is long past-time to end this practice. Only then will we as a society get what we expect and deserve from our investments in medical science.
*****
My co-PI and dear friend, Jay Amsterdam, died suddenly on July 28th as we waited to learn the fate of our appeal to BMJ Open. On May 14th, BMJ Open rejected our group’s reanalysis of STAR*D’s medication-switch study published in NEJM. Even worse, the editor did so without allowing us the opportunity to respond to the 5 peer reviewers’ criticisms.
This was devastating news. We first submitted the manuscript on March 25, 2025 and 14 months later, we learned it was rejected without the opportunity to respond. So, Jay and I got to read the criticisms—and learn the names of each reviewer—but not respond? What’s OPEN with that?
By December 31, 2025, the editor had received all 5 reviews—one of which was written by Dr. Maurizio Fava, a STAR*D PI. So the editor had Maurizio’s review, and we had to wait almost 6 months to read it…and then not respond?
On August 10th, the editor emailed me stating: “I appreciate your patience and am sorry for the wait you’ve experienced…we expect to provide you with a formal update within the next few days.” And 12 days later, I am still waiting…
Jay and I were very different. He was a true Renaissance man who quoted Latin while in essence, I dropped out of high school in 10th grade to join the work release program. What we shared was a commitment to integrity . . . and justice. Fortis Fortuna Adiuvat—fortune favors the brave—Jay would often say, when bolstering my nerve to press send.
Jay’s sudden passing has given me a measure of the courage he showed unto the very end. Carpe diem. Seize the day, we press on.
***
–
This post was originally published on Mad In America.
Trust this device for 30 days