The FBI requested the review that ended comparative bullet lead analysis, accepted the finding, and retired the method. Then it said nothing to the people already convicted by it. Two years later a newspaper and a television program had to go find them.
Comparative bullet lead analysis linked crime scene bullets to a suspect’s ammunition by measuring trace elements. A 2004 National Academy of Sciences review found the chemistry sound and the courtroom conclusions overstated to the point of being misleading. The FBI discontinued the method in 2005 and notified no one. In November 2007 journalists reported that affected defendants and courts had still not been told, with appeal deadlines running.
If the chemistry was accurate, what exactly went wrong?
The inference. Bullets from the same box did not always share a composition, and bullets from different boxes sometimes did. Measuring a composition accurately says nothing about how many other bullets in the world share it, which is the question a jury was being asked to answer.
What is chaining?
The practice of treating a match between bullet A and bullet B, and between bullet B and bullet C, as establishing a relationship between A and C. Each step introduces error and the practice compounds it without acknowledging that it has.
Did anyone inside the FBI raise concerns?
Yes. A former FBI metallurgist became one of the method’s most effective critics, and the challenges he developed contributed to early reversals. Internal expertise identified the problem before external review did.
Why does the notification failure matter more than the method failure?
Because the method failure was corrected in about a year once the question was asked. The notification failure ran for two more years, against appeal deadlines, and was resolved by people with no obligation to do it.
What I Notice First About This One
I want to start by giving the FBI credit, because this installment is going to be hard on the agency and the credit is real.
The FBI asked for the review. Not a court, not a defense organization, not a journalist. Facing sustained criticism of comparative bullet lead analysis, the bureau went to the National Academy of Sciences and requested an examination of its own method. That is the opposite of what happened in Part III, where an oversight body had to be created and then commissioned an outside expert, and the opposite of Part V, where the question reached an institution only because a complaint was filed on behalf of one man.
Then the Academy came back with an unfavorable answer, and the FBI accepted it and stopped using the method. On the science, this is the cleanest institutional performance in the entire series.
And it produced almost nothing for the people already in prison.
That is the whole installment. Everything the science-reform conversation asks institutions to do, the FBI did, in about the right order, on about the right timeline. The people convicted by the method were no better off for any of it until a newspaper went looking for them two years later.
The Method
Comparative bullet lead analysis, sometimes called compositional bullet lead analysis, was used by the FBI Laboratory beginning in the 1960s. Its first prominent application came in the investigation of the assassination of President Kennedy.
The technique measured the concentrations of trace elements in bullet lead, principally arsenic, antimony, tin, copper, bismuth, silver, and cadmium. The premise was that lead is produced in batches, that each batch carries a distinctive elemental profile, and that bullets sharing a profile therefore likely came from the same source, whether a manufacturer, a production run, or a box.
Applied to a crime scene, the argument ran like this. A bullet fragment recovered from a victim has composition X. A box of ammunition in the defendant’s home has composition X. Therefore the fatal bullet came from that box, or from one very much like it. Juries heard the second half of that formulation with the emphasis on the first clause.
This is a different failure mode from the disciplines examined so far. Hair comparison and bite mark analysis produced subjective judgments dressed as measurements. Bullet lead produced real measurements attached to an unsupported inference. The number on the instrument readout was correct. The sentence the examiner built around it was not.
What the Academy Found
The FBI referred the method to the National Academy of Sciences, and the resulting report was published in February 2004 as an assessment of bullet lead evidence.
The review largely validated the chemistry. The analytical instrumentation was appropriate, the measurements were being performed correctly, and the seven trace elements selected were reasonable choices for comparison. On the question of whether the laboratory could accurately determine the composition of a piece of lead, the answer was yes.
The problem lay entirely downstream of the measurement. Variation in manufacturing meant that bullets from a single box did not reliably share a composition, and bullets from unrelated boxes sometimes did. The statistical tests as applied by the FBI could produce confusion when conveyed to prosecutors or explained to a jury. The Academy concluded that the bureau’s decades of courtroom statements linking a particular bullet to a defendant’s ammunition were so overstated as to be potentially misleading under the rules of evidence.
Read the finding carefully, because its shape matters. The Academy did not say the FBI had been sloppy. It said the FBI had been accurate about something other than what it told juries it was accurate about. That distinction is the reason a conviction based on this testimony is difficult to attack case by case: nothing in the laboratory file is wrong.
The Lab holds free Clutch Justice tools for public records work and record comparison, including resources for locating and reading trial transcripts against the underlying laboratory documentation.
Explore The LabThe Retirement
In 2005 the FBI discontinued comparative bullet lead analysis. The New York Times reported the decision that September. The agency had requested a review, received a finding it did not like, and acted on it within roughly a year.
Measured as a scientific correction, that is a good outcome delivered quickly. Measured as a remedy, it delivered nothing at all.
Discontinuation is prospective by its nature. Ending a practice affects cases that have not happened yet. It does not touch a single conviction already entered, and it does not tell anybody that their conviction is now resting on a method the agency that performed it no longer offers.
A method is discontinued by a decision. A conviction is undone by a proceeding. The first requires one institution to act once. The second requires a person, a lawyer, a filing, a record, and a judge, and it requires all of them to learn that anything has changed.
Two Years of Silence
In November 2007 The Washington Post and 60 Minutes published a joint investigation reporting that the FBI laboratory had not taken steps to alert affected defendants or the courts, more than two years after the method was abandoned, while the windows for appealing those convictions were closing.
On November 19, 2007, the Innocence Network and the National Association of Criminal Defense Lawyers announced a joint task force to review convictions affected by the discredited analysis. The FBI then stated its intention to re-examine cases in which the testimony had been offered.
Reconstruct that sequence and the causation is unambiguous. The scientific finding arrived in 2004. The practice ended in 2005. The obligation to notify was recognized in 2007, in the same month as a national broadcast and a front-page investigation, and the work of identifying affected people was taken up by volunteers.
No rule required the notification. No statute set a deadline. No oversight body had authority to compel it. The gap between the 2005 discontinuation and the 2007 acknowledgment was not a period during which an institution failed to meet an obligation. It was a period during which no obligation existed, and it ended because journalism substituted for one.
Bullet lead comparison is first applied prominently in the investigation of the Kennedy assassination, then adopted into routine FBI Laboratory casework.
The FBI Laboratory performs bullet lead analysis on cases submitted by law enforcement agencies. The technique plays a significant role in many of them, including capital cases.
At the FBI’s own request, the National Academy of Sciences finds the chemical measurements sound but the statistical interpretation and courtroom characterizations overstated to the point of being potentially misleading under the rules of evidence.
A New Jersey appellate court orders a new trial in a case where the bullet lead testimony was challenged, among the first reversals to rest on the developing critique of the method.
The FBI abandons comparative bullet lead analysis. The decision is reported nationally. No case identification, notification program, or review of prior testimony accompanies it.
A joint Washington Post and 60 Minutes investigation reports that affected defendants and courts have not been alerted while appeal windows close. Days later the Innocence Network and NACDL announce a joint task force, and the FBI states it will re-examine affected cases.
Against Part V
Set this installment beside the previous one, because the comparison produces the finding that neither case yields alone.
An outside commission acted on a complaint about the method, recommended a moratorium, and in the same action ordered identification of every affected conviction in the state.
Obligations one and two performed. Practitioners were still defending the discipline at the time.
The agency requested the review, accepted the finding, and ended the practice. It identified no affected convictions and notified no one for two further years.
Obligation one performed cleanly. Obligation two performed only after press exposure, by volunteers.
In Texas the entity that stopped the method was not the entity that had used it, and it had authority to order a review across cases.
The variable is not good faith. It is whether an institution other than the practitioner had the power to require case identification.
The uncomfortable conclusion is that the agency acting in good faith produced a worse outcome for affected people than the agency acting under external pressure. Not because the FBI was less sincere than the Texas commission. Because a body that stops its own practice has completed the task from its own perspective, and only an outside body experiences case identification as part of the same job.
What This Means for Michigan
Michigan has no laboratory that performed comparative bullet lead analysis, because the method was performed solely by the FBI Laboratory. Michigan cases are nonetheless in the affected population, because state and local agencies submitted evidence to the FBI for analysis and then tried the resulting cases in Michigan courts.
How many, nobody can say. The same absence documented in Parts II and V applies. There is no Michigan record of which convictions rested on which forensic discipline, no body with authority to order that question be answered, and no mechanism by which a federal discontinuation reaches a state court file.
A receiving duty for external forensic invalidations. When a federal agency or another state discontinues a method, some Michigan body should be obligated to determine whether Michigan convictions relied on it. At present a discontinuation announced in Washington produces no action in Lansing, because nobody in Lansing is assigned to hear it.
Tolling of post-conviction deadlines on invalidation. The specific harm in 2005 to 2007 was that appeal windows continued to run while the affected population did not know anything had changed. A deadline that expires during a period of institutional silence is not a deadline. It is a transfer of the consequences of that silence onto the person least able to bear them.
Why This Matters Beyond One Method
I want to close on the thing that unsettles me about this installment, which is that it removes the easiest explanation for everything else in this series.
Reading Parts II through IV, it is possible to conclude that the failures were failures of will. The FBI scoped its hair review around its own examiners. Texas removed commissioners before a hearing. Someone, somewhere, did not want the answer.
Bullet lead does not permit that reading. The FBI asked the question nobody made it ask, got the answer it did not want, and published the retirement. Every element of institutional good faith is present in the record, and the outcome for people in prison was silence for two years while their deadlines expired.
Which means the problem is not, or not only, that institutions resist correcting themselves. It is that correcting a method and correcting a conviction are different tasks, and every institution in the American system is organized to perform the first one. The second belongs to nobody. When it gets done, it gets done by a newspaper, a clinic, a task force of volunteers, or a broadcast segment.
That is not a system with a gap in it. That is a gap with some institutions arranged around the edges.
Every discipline examined so far failed in the open, where anyone with access to the file could eventually examine it. Part VII takes up probabilistic genotyping, the software now interpreting complex DNA mixtures in American courtrooms, where the validation lag is no longer maintained by neglect but by intellectual property law, and where the one time a court ordered a source code release, the errors were found within a year.
Continue Your Investigation
If this reporting raised more questions, use the Clutch Justice ecosystem to keep going.