The Method/Entry 1.04/One technique, and what it can and cannot carry
STR multiplexes
Short repeated sequences at many sites, amplified together in one reaction, are what every modern profile and every national database is built on.

Every profile in every national database, every courtroom comparison, every cold-case hit — they all resolve to the same underlying thing: short tandem repeats, measured at multiple sites in a single amplification.

What an STR is, and why it varies
The human genome contains millions of short sequences repeated end to end — motifs of two, three, four or five bases that occur in tandem runs whose length differs between individuals. These are short tandem repeats, STRs. At any given location, or locus, a person carries two copies of the sequence — one on each chromosome — and the number of repeat units at each copy is their allele pair for that site. Someone might carry seven repeats on one chromosome and eleven on the other; their sibling might carry nine and twelve. The variation is generated by replication slippage during cell division, is largely heritable, and accumulates across generations until the population carries a broad distribution of allele lengths at each locus.
None of this is informative on its own. A single locus distinguishes poorly: too many people share the same pair of numbers. The power of the technique comes from combining sites. If each locus has a large number of alleles at credible frequencies, and the loci are on different chromosomes or sufficiently far apart to behave independently, then the probability of two unrelated people sharing the same profile across all of them drops to figures that, for the core sets used today, are routinely cited in the range of one in many billions or beyond. The NIST STR database ↗ maintains the population frequency data and variant allele catalogues on which those calculations depend.
From the register
How a multiplex works — the steps
- Biological sample collected and DNA extracted
- PCR amplificationprimers for all target loci, each fluorescently labelled, in a single reaction
- Capillary electrophoresisfragments separated by size through a polymer-filled capillary
- Laser detector reads fluorescent signal as fragments pass; dye colour distinguishes overlapping size ranges
- Software converts peak positions to allele calls; output is the electropherogram and a numerical profile
From single locus to multiplex
Early PCR-based forensic work — following the shift away from the Southern blot and its requirement for microgram quantities of DNA — amplified one locus at a time. That was slow, consumed sample, and produced results that needed to be combined from separate reactions. The step that made modern forensic DNA practical was the multiplex: a single PCR reaction containing primers for many loci simultaneously, each primer pair tagged with a fluorescent dye. The thermal cycler amplifies all targets in parallel, and the products are then separated by capillary electrophoresis, where each fragment migrates through a polymer-filled capillary and passes a laser detector. Size determines migration time; dye colour distinguishes loci whose size ranges overlap. The detector records a trace — a series of peaks — and software converts peak positions into allele calls. The result is a profile: a list of numbers, one pair per locus.
Commercial multiplex kits standardised this process. The Applied Biosystems SGM Plus kit, introduced in the late 1990s, set ten loci plus a sex-determining amelogenin marker as the basis for the UK's National DNA Database. Its American counterpart, the CODIS system run by the FBI Laboratory, originally specified thirteen core STR loci; in 2017 the US expanded that to twenty. The expansion was driven in part by a desire to reduce adventitious matches — coincidental profile agreements between unrelated people — as databases grew larger. SWGDAM, the Scientific Working Group for DNA Analysis Methods, coordinates the interpretation guidelines that US laboratories work to, and the National Institute of Standards and Technology ↗ supplies the reference materials and population data that underpin the statistics.

Reading the output
A clean single-source profile from a reference buccal swab looks straightforward on the electropherogram: two peaks per locus, or one if the person is homozygous at that site. But the output carries features that require interpretation even in the simple case. Stutter peaks — smaller peaks one repeat unit below the true allele — arise from replication slippage during PCR and are expected artifacts, not contamination. The ratio of stutter to true peak height is predictable enough that laboratories set thresholds below which a peak is attributed to stutter rather than treated as a genuine allele. Pull-up, or bleed-through between dye channels, is a second artifact. Off-ladder alleles — repeat counts that fall between the calibration standards — occur and must be handled. None of these invalidate the technique; they are all documented, and their management is specified in validation studies and guidelines.
What changes the interpretive difficulty is template quantity and sample composition. At low template — trace amounts from a touched surface, a shed hair root, a small blood stain — amplification becomes stochastic. Some alleles fail to amplify reliably, a phenomenon called drop-out, while background signal can occasionally produce a spurious peak, called drop-in. Laboratories set a stochastic threshold: a peak height below which heterozygous drop-out cannot be excluded and the profile must be treated with corresponding caution. These are not failures of the STR method as such; they are consequences of pushing any amplification-based technique toward the limits of its input. The National Academy of Sciences' 2009 report Strengthening Forensic Science in the United States identified low-template interpretation as one of the areas where practice had run ahead of systematic validation.
From the register
Key thresholds and artifacts
- Stutterexpected minor peak one repeat unit below true allele; managed by height-ratio threshold
- Drop-outfailure of an allele to amplify at low template; addressed by stochastic threshold
- Drop-inspurious low-level peak from background signal; distinguished from true allele by height and context
- Homozygous locusone peak instead of two; could also indicate drop-out of one allele
Mixture interpretation — when more than one person's DNA is in the same sample — compounds all of these considerations. Peaks overlap, minor contributors can be masked, and the number of contributors is often itself uncertain. Probabilistic genotyping software now handles the more complex mixtures by modelling the space of possible contributor genotypes and reporting a likelihood ratio: how much more probable is the evidence if this person is a contributor than if they are not. The approach is more defensible than the earlier practice of manual mixture interpretation, but it introduces its own complexity, because the software's models and their assumptions are not always transparent to courts.

The multiplex as infrastructure
The choice of which loci a multiplex includes is not merely technical — it is infrastructure. Profiles in a database are comparable only if they were generated at the same sites. When the US expanded its core loci in 2017, laboratories had to re-type samples in order to make older profiles searchable against newer ones. Cross-border comparison, managed through the Interpol gateway and bilateral agreements between national databases, requires agreed minimum overlapping sets. The European Network of Forensic Science Institutes runs collaborative exercises to verify that participating laboratories produce consistent results at shared loci, which is the only meaningful way to establish that a database comparison is trustworthy.
The STR multiplex also shapes what is retained and what is searched. National databases hold the numerical profile, not the underlying biological sample in most cases, and the profile at the agreed loci is what gets compared. Because STRs are non-coding — they sit outside the portions of the genome that determine traits or disease risk — the profile carries no medically meaningful information. That boundary is part of the legal and ethical framework for database retention, and it is why the method, rather than whole-genome sequencing, remains the standard for forensic databasing. A profile number tells you nothing about eye colour or health; it tells you, with a quantifiable probability, whether two samples came from the same person.
That is precisely what every database comparison, every courtroom statistic and every cold-case link depends on: not an image, not a sequence, but a list of numbers at named sites — powerful because the sites were chosen, validated and fixed as common infrastructure across laboratories and across borders.
From the register
Locus counts across major systems
- SGM Plus (UK, late 1990s)
- 10 STR loci + amelogenin
- Original CODIS (US, 1998)
- 13 core STR loci
- Expanded CODIS (US, 2017)
- 20 core STR loci
- Expansion reason
- reduce adventitious matches as database size increased