The image below presents a comparison that, at first glance, seems absurd.
A Chihuahua and a Labrador Retriever differ dramatically in size, anatomy, behavior, and function. One is a tiny companion dog; the other is a large working retriever bred for swimming and carrying game.
Yet their genetic differentiation (FST = 0.116) is almost identical to several ordinary human intercontinental comparisons: northern Italians vs Japanese (0.116), southern Han Chinese vs northern Italians (0.116), Beijing Han Chinese vs Mozabites (0.116), and French vs Japanese (0.113).
This apparent paradox illustrates a central theme of this article: genetic distance and phenotypic difference are not the same thing. As we will see, FST is often treated as a universal ruler of biological difference, but that interpretation can be highly misleading.
There are countless dog and cat breeds, and they differ a lot in hair color, size, and even intelligence and personality. Consider the stark behavioral differences we take for granted: a Border Collie lives to herd livestock, a Siberian Husky is driven to pull sleds for miles, and a Cavalier King Charles Spaniel is content simply being a lapdog. Meanwhile, in the feline world, a Siamese is famously vocal and demanding of attention, while a British Shorthair remains placid, reserved, and fiercely independent.
Looking at these dramatic differences in appearance and behavior, an intriguing question naturally arises: how much do they differ genetically? And more provocatively, are the genetic differences between a Bulldog and a Chihuahua similar to those between different human populations?
When researchers want to measure this kind of genetic distance, they often reach for a mathematical metric called FST. But as it turns out, relying purely on this number to map out biological reality can lead us straight into a trap.
The Illusion of the Absolute Ruler
FST is often treated as if it were a simple, objective genetic distance ruler. The logic seems straightforward: put two populations into the formula, get a number, and then rank the distances. While that is undeniably useful, it is also dangerous.
The truth is that an FST value is not just pure, unadulterated biology. It is highly sensitive to the mechanics of the study itself, including:
Sample size (N)
Total marker count
Missing data (”missingness”)
Ascertainment bias
The specific way the population labels were built in the first place
To demonstrate how easily these numbers can be misinterpreted, I compared pairwise Hudson FST values across three distinct within-species datasets:
Humans: Modern human groups from the Allen Ancient DNA Resource (AADR) v66 2M panel (179 population labels, yielding 15,931 pairwise values).
Dogs: Purebred dog SNP data from the Parker 2017 dataset after quality control, minority allele frequency (MAF) filtering, and linkage disequilibrium (LD) pruning (177 breed labels, yielding 15,576 pairwise values).
Cats: Cat breed and random-bred groups from the Matsumoto SNP dataset after QC and LD pruning (14 group labels, yielding 91 pairwise values).
The concrete result of this comparison is surprising: human, cat, and dog FST values live on a remarkably similar numerical scale.
What Exactly Is FST Measuring?
At its core, FST measures allele-frequency differentiation. If two groups share very similar allele frequencies across many single-nucleotide polymorphisms (SNPs), the FST is low. If those frequencies differ significantly, the FSTis high.
However, because we must estimate allele frequencies from finite samples, small sample sizes introduce a massive amount of statistical noise. Because FST behaves like a distance metric, random noise inherently pushes groups apart, creating the illusion of large genetic separations where none exist.
The Data: Uncomfortably Close Distributions
When we look at the raw distributions, the overlap between species is striking. Cats are especially close to the human median, and while dogs are shifted higher, as you would expect from strictly closed, inbred breed structures, the vast majority of dog breed-pair values still sit firmly within the central human range.
Table 1: Pairwise Hudson FST Summary Statistics
Using the human 5th-to-95th percentile range (0.017 to 0.468) as our baseline reference interval, we get a clear picture of just how much these distributions overlap:
Table 2: Overlap with the Human F_ST Distribution
This means that 98.9% of cat breed comparisons and 97.3% of dog breed comparisons fall squarely inside the normal genetic distance range found between human populations.
The visual density plot below illustrates this dramatic overlap, highlighting how the mathematical distributions occupy the exact same territory:
A Chihuahua and a Labrador can be separated by roughly the same FST as an Italian and a Japanese person.
So why don’t they look anything alike?
To answer that question, I match dog breeds to human population pairs with nearly identical genetic distances, reveal the complete cat FST matrix and an expanded matrix of major dog breeds, test the sample-size trap that inflates genetic distances, and explain the quantitative genetic logic behind Lewontin’s Fallacy.
Upgrade to a paid subscription to continue reading and access the full archive.
Your support directly funds the independent data pipelines, quality control, and computation required to turn raw genomic data into analyses like this one.
The heatmaps make this less abstract. The cat panel shows all 14 cat groups in the dataset. The dog panel uses a readable subset of 12 familiar breed labels, because a full 177-by-177 dog heatmap would be useless in a post. The point is not that every cat or dog pair matches every human pair. It is that ordinary within-species comparisons can easily produce FST values in the same broad range people often treat as intuitively large.
and the dogs:
Real-World Collisions: Putting Faces to the Numbers
To understand just how severely the FST metric flattens biological reality, it helps to put actual faces to these numbers. When we match specific dog breed pairs against human population pairs with nearly identical FST values, the paradox becomes undeniable.
Comparison 1: The Moderate Distance FST approx 0.11)
The Human Pair: Han Beijing vs. Tuscan Italy (FST = 0.110) or Japanese vs. Tuscan Italy (FST = 0.112).
The Dog Pair: Chihuahua vs. Labrador Retriever (FST = 0.116).
The Phenotypic Reality: Genetically, the distance between a Chihuahua and a Labrador is virtually identical to the distance between a European and an East Asian human population. Physically and behaviorally, however, they occupy different universes. The Chihuahua is a tiny, compact companion animal built for alert toy-dog ecologies. The Labrador is a large, powerfully built retriever meticulously selected for swimming, carrying heavy game, high biddability, and intense cooperative work with humans.
Comparison 2: The Intermediate Distance (FST approx 0.15)
The Human Pair: Tuscan Italy vs. Yoruba (FST = 0.151) or Gambian vs. Tuscan Italy (FST = 0.146).
The Dog Pair: Beagle vs. Dachshund (FST = 0.156).
The Phenotypic Reality: This value matches standard European-to-African human population distances. Yet, the phenotypic contrast between these two hounds is immense. The Beagle is a medium-sized scent hound optimized for long-distance trailing and pack hunting. The Dachshund features a radically altered body plan—uniquely long-bodied and short-legged, engineered specifically to face down badgers inside the tight, subterranean confines of a burrow.
Comparison 3: The Intercontinental Scale (FST approx 0.18)
The Human Pair: Japanese vs. Yoruba (FST = 0.185) or Gambian vs. Japanese (FST = 0.182).
The Dog Pair: Beagle vs. Labrador (FST = 0.189).
The Phenotypic Reality: To find a dog comparison that matches a deep, African-to-East Asian human baseline, we look not at radically different-looking toy dogs, but at a Beagle and a Labrador.
This is precisely why running these raw numbers is such an important exercise. It provides a vivid demonstration of the quantitative reality: vast, unmistakable differences in anatomy, size, and behavioral programming can easily exist within a single species at FST values that perfectly overlap with ordinary human continental comparisons.
Falling into the Sample-Size Trap
Why are these numbers so similar? Part of the answer lies in a major methodological pitfall: the sample-size trap.
In the human dataset, the average sample size of a pair correlated negatively with its calculated FST (Pearson r = -0.192). This isn’t because small, isolated populations are magically vastly different from everyone else in every scenario; it’s because small samples are terrible at accurately estimating allele frequencies. The resulting mathematical noise destabilizes the statistical “high tail,” inflating the final distance metric.
The number of observed markers matters just as much. Across all three species, a lower observed SNP count was strongly associated with a higher FST value. The Pearson correlation between observed SNPs and FST was -0.575 for humans, -0.585 for dogs, and -0.430 for cats.
Table 3: Correlations Between Contextual Variables and Calculated FST
The Downsampling Check: Isolating Noise from Bias
To make the sample-size point concrete, I used two large human pairs with FST above 0.1 and downsampled them. The pairs were Japanese vs Yoruba and Gambian vs Tuscan Italy. When I reduced N alone but kept the huge AADR marker panel, the mean FST barely moved. The important change was volatility. When I combined N=2 with tiny SNP panels, the high outliers became large. For Japanese vs Yoruba, the full FST was 0.185; with N=2 and 100 SNPs, one replicate reached 0.328. For Gambian vs Tuscan Italy, the full FST was 0.146; with N=2 and 100 SNPs, one replicate reached 0.277.
Conclusion: FST, Lewontin’s Fallacy, and why Pixels Aren’t the Picture
Does the striking mathematical overlap between humans, cats, and dogs mean that FST is a reliable guide to how physically or behaviorally different two groups are? Absolutely not. In fact, this comparison perfectly illustrates what the geneticist A.W.F. Edwards termed Lewontin’s Fallacy.
In 1972, Richard Lewontin famously argued that because most human genetic variation occurs within populations rather than between them (reflected in a low average FST), human population differences were biologically insignificant. But this is a mathematical illusion. FST is an average metric across many individual, unlinked loci. It looks at single pixels, not the emerging picture.
When it comes to complex, polygenic traits—like a dog’s height, a cat’s coat color, or behavioral tendencies—phenotypic differences are not driven by single genes acting in isolation. They are driven by the correlations among allele frequencies across thousands of loci.
This leads to two critical reasons why low FST distances can coexist with massive phenotypic differences:
The Cancellation Effect: Because FST is an average distance metric, it completely flattens the direction of natural or artificial selection. In a standard FST calculation, an allele that makes an animal larger and an allele that makes it smaller both contribute positively to the “genetic distance” number. The signs of the different alleles on the actual trait cancel out in the math, making the groups look close genetically even if selection has aggressively pushed their physical traits apart.
QST > FST (Phenotypic vs. Genetic Differentiation): In quantitative genetics, we compare FST (differentiation at neutral genetic markers) to QST (differentiation in quantitative, polygenic traits). When artificial selection aggressively breeds a Chihuahua for tiny size and a Great Dane for massive size, it shifts the allele frequencies in the same phenotypic direction across thousands of relevant loci. This allows the group polygenic scores to diverge radically. Consequently, QST becomes vastly larger than FST.
This is exactly what the dog and cat data illustrate so beautifully. A French Bulldog and a Border Collie share a massive amount of their evolutionary history, meaning their baseline genetic background is highly similar and their pairwise FST remains firmly within the typical human range. Yet, because human breeders have spent centuries sorting and stacking directional alleles for behavior and morphology, their physical reality is worlds apart.
The ultimate lesson of the FST trap is that a low genetic distance coefficient is not proof of phenotypic uniformity. FST measures the average background drift of unlinked genetic pixels; it is completely blind to the powerful, coordinated genetic architecture that shapes the actual organisms we see, live with, and love.
References
Mallick, S., et al. (2024). The Allen Ancient DNA Resource (AADR) a curated compendium of ancient human genomes. Scientific Data, 11, 182.
Parker, H.G., et al. (2017). Genomic analyses reveal the influence of geographic origin, migration, and hybridization on modern dog breed development. Cell Reports, 19, 697-708.
Matsumoto, Y., et al. (2021). Genetic relationships and inbreeding levels among geographically distant populations of Felis catus from Japan and the United States. Genomics, 113, 104-110.
Matsumoto, Y. (2020). Feline 60K SNP chip data originated from the domestic cat in Japan. Dryad/Zenodo dataset.
Hudson, R.R., Slatkin, M. & Maddison, W.P. (1992). Estimation of levels of gene flow from DNA sequence data. Genetics, 132, 583-589.










