Monday, 27 July 2026

Distance Without Discrimination

 

A review of Resolving the Geochemical Provenance of Stonehenge Bluestones: A Sample-Size-Independent Multivariate Framework (G. W. Taylor, June 2026)


Taylor's paper, DOI:10.13140/RG.2.2.17039.96164 , sets out to demonstrate that the Stonehenge bluestones are geochemically identical to outcrops in the Mynydd Preseli, using rare earth element (REE) ratio data and three complementary techniques: PERMANOVA, SIMPER, and Euclidean nearest-neighbour distance mapping. It is generous with its data. Table 1 gives the full REE ratio matrix for all forty-two analyses, Table 2 the complete pairwise PERMANOVA output, and Table 3 the entire 20 × 22 distance matrix with nearest-neighbour assignments. That openness makes a substantive review possible, and everything below is derived from those three tables and from the published source of the underlying measurements.

The conclusion the paper reaches is, in outline, correct. The problem is that the method used to reach it cannot distinguish that conclusion from its opposite — and the paper's own data, read against their source, contain the demonstration.

1. What the paper argues

The argument runs in three stages.

A global PERMANOVA on the pooled data fails to reject the null hypothesis of no difference between Stonehenge and Welsh samples (p = 0.85, pseudo-F = 0.074). The author treats this as positive support for common origin. Because the pairwise PERMANOVA matrix contains a handful of rejections, and because archaeological sample sizes are small, those pairwise results are judged unreliable and set aside.

To bypass the sample-size problem, all Stonehenge analyses are pooled into one group and all Welsh analyses into another, and SIMPER is used to compare group means. The means agree closely across all twelve ratios — La/Lu at 15.2 against 15.3, La/Yb at 2.12 against 2.12 — which is presented as demonstrating homogeneity.

Finally, Euclidean distances are computed between every Stonehenge analysis and every Welsh analysis. Several minima are very small, and these are offered as sample-level confirmation, with the two smallest presented as headline results.

Stated at its strongest, the argument is: three methods at three different scales all fail to find a difference, so there is no difference to find.

2. The data are from Bevins, Pearce and Ixer (2021), uncited

The paper does not say where its measurements come from. They can be identified exactly.

Taylor's Table 1 consists of twelve ratios, La divided by each of the other REE. Taking sample CGD1 from Table 2 of Bevins et al. (2021) — La 3.36, Ce 9.05, Pr 1.42, Nd 7.50, Sm 2.42, Eu 0.97, Gd 2.75, Tb 0.50, Dy 3.30, Ho 0.65, Er 1.78, Yb 1.71, Lu 0.24 ppm — and performing those divisions reproduces Taylor's row to every reported digit:

RatioComputed from Bevins et al. (2021)Taylor Table 1
La/Ce0.3712710.371271
La/Pr2.3661972.366197
La/Nd0.4480000.448
La/Sm1.3884301.38843
La/Eu3.4639183.463918
La/Gd1.2218181.221818
La/Tb6.7200006.72
La/Dy1.0181821.018182
La/Ho5.1692315.169231
La/Er1.8876401.88764
La/Yb1.9649121.964912
La/Lu14.00000014

The same reconstruction succeeds for OU10, OU10 rpt, PCM7 and PCM7 rpt. Taylor's entire dataset is the REE table of Bevins et al. (2021), converted to La/X ratios.

That paper appears nowhere in Taylor's bibliography, which cites three methodological references (Clarke, 1993; Anderson, 2001; Gower, 1966) and no archaeological or geological source at all. The dataset represents years of laboratory work, sample access negotiated with the Natural History Museum and the Salisbury and South Wiltshire Museum, and analytical development at Aberystwyth. It should be cited.

Two errors follow from the missing attribution. Taylor's Table 1 caption describes the measurements as LA-ICP-MS data. Bevins et al. (2021) used solution nebulisation ICP-MS on acid-digested bulk powders — a different technique producing a different kind of measurement. And the ratios are described as "mineral element proportionality ratios"; they are ratios of whole-rock elemental concentrations, and carry no information about mineral proportions.

3. The groups are sampling categories, not geochemical ones

Taylor's group codes — PO1, SLF1, SODC1, SOF1, SLF2v, SLF2vi, PO2ii, PO2iii, PO2iv, PO3, SLF3 — have no stated derivation. They are Bevins et al.'s group numbers prefixed with their Sample source field:

Taylor codeBevins groupBevins sample sourcen
PO1Preseli Group 1Preseli outcrop3
SLF1Stonehenge Group 1Stonehenge Landscape fragment7
SODC1Stonehenge Group 1Stonehenge orthostat drill core4
SOF1Stonehenge Group 1Stonehenge orthostat fragment1

Bevins et al.'s Stonehenge Group 1 is a single population of twelve analyses, argued on REE grounds to derive from one intrusive body and attributed to Carn Goedog. Taylor has divided it into three separate "geochemical groups" according to whether each sample was a drill core, a chip off an orthostat, or a surface find from the Stonehenge Landscape — and then run PERMANOVA between them.

This explains the two singleton groups. SOF1 contains one analysis because exactly one Group 1 sample happened to be an orthostat fragment (SH67); SLF2vi contains one for the same reason. These are not rare geological units. They are the intersection of a geochemical group with a collection method.

The scheme is also applied inconsistently: SLF2v contains a drill core (SH62), and SLF3 contains three orthostat fragments, which suggests labels carried down from the first row of each sorted block rather than a deliberate classification.

4. Twelve Welsh analyses and one Stonehenge analysis have been removed

Bevins et al. (2021) report 32 Preseli analyses and 23 Stonehenge analyses. Taylor uses 20 and 22 respectively. The omissions are unstated, and they are not random.

Dropped from the Welsh set: all four Group 2i samples (PCP12, PCS13, PCTF14, PCA15), together with PCM6, PMB9, PMB10, PCB16, PCB19, PCGF27, PCGF27 rpt and PCAW47.

Group 2i matters more than any other. Its distinctive concave-down, MREE-enriched patterns are what allow Bevins et al. to exclude Craig Talfynydd, Carn Sian, Carn Arthur and Carn Bica as sources for any Stonehenge dolerite. It is the clearest discriminating result in the source paper, and it is the one group removed in its entirety.

The twenty retained Welsh analyses span La from 3.04 to 4.26 ppm. All twelve dropped analyses fall outside that band — eight above it, four below. The retained set is, exactly, the central twenty of thirty-two by REE abundance.

Dropped from the Stonehenge set: SH42, at 5.50 ppm La the most REE-rich Stonehenge analysis, and the single sample that Bevins et al. identify as falling outside the envelope of the Preseli patterns in their Fig. 4.

No reason is given for any of this. Whatever the cause, the effect is that both tails were removed from the Welsh distribution and the one non-conforming Stonehenge analysis was removed from the other, before a test for homogeneity was run.

5. The conclusion is almost certainly right

It is worth being clear about this before going further. That the great majority of the Stonehenge bluestones derive from the Mynydd Preseli has been the settled position since Thomas (1923), and successive work by Bevins, Ixer, Pearce and colleagues has narrowed several lithologies to individual outcrops, with excavation at Carn Goedog recovering a Neolithic quarry (Parker Pearson et al., 2019). Nobody reading this review needs persuading of the Preseli connection.

That is precisely why the paper repays close reading. We already know the answer. A method applied to a case with a known answer, which cannot recover that answer reliably, has been shown not to work — and the demonstration is far cleaner than it would be on an open question. What follows is not a defence of some rival provenance. It is an argument that this framework would have produced the same confident result had the stones come from anywhere.

6. What a provenance method has to demonstrate

A geochemical fingerprint is only useful if it discriminates. Showing that a Stonehenge sample resembles a Welsh outcrop establishes nothing on its own; the question is always whether it resembles that outcrop more than it resembles the alternatives, and by a margin larger than the measurement error.

Two requirements follow:

  1. A comparison set. At least one candidate source outside the favoured region, so that a match can be shown to be selective rather than universal.
  2. A resolution limit. An estimate of how much two measurements of the same rock differ, so that "close" can be distinguished from "indistinguishable given the noise."

The paper meets neither. No non-Welsh source appears anywhere, so the specificity of the match is never tested. And no error estimate is offered — the framework is presented as a way of avoiding variance estimates rather than as a way of quantifying them.

The second gap can be closed from the source data, and doing so is the substance of the next section.

7. The replicates give the resolution limit directly

Bevins et al. (2021) include three analytical replicate pairs: OU10 and OU10 rpt, PCM7 and PCM7 rpt, PCGF27 and PCGF27 rpt. Each pair is one rock analysed twice. Two of the pairs survive into Taylor's dataset; the third was among the samples dropped.

Two measurements of one rock should, if the method has any resolving power, be closer to each other than either is to a genuinely different rock. The distance between replicates is therefore an estimate of the method's noise floor. Computing Taylor's metric on the Bevins concentrations gives it exactly:

Replicate pairEuclidean distanceIn Taylor's dataset?
PCGF27 / PCGF27 rpt0.1475dropped
OU10 / OU10 rpt0.1909retained
PCM7 / PCM7 rpt0.2952retained

Now set those against the paper's results.

The headline match is smaller than the noise floor. The abstract, methodology and conclusion all cite OU11 to CGD2 at d = 0.143 as the exemplary near-zero pairing. Every one of the three replicate distances exceeds it. A sample in this dataset sits further from itself than the flagship match sits from its claimed source.

Most of the reported matches fall inside the noise floor. Of the twenty-two nearest-neighbour distances in Taylor's "Closest Euclidean" row, thirteen are below 0.295 and four are below 0.191. For these the assignment carries no information.

The replicates disagree with each other about provenance. OU10 and OU10 rpt are one stone:

AnalysisNearest Welsh outcropGroupdRunner-upGroupd
OU10CGD2PO10.203463PCM30PO30.317994
OU10 rptPCM30PO30.222013CGD2PO10.345259

The same physical stone is assigned to Carn Goedog on one analytical run and to a Group 3 locality on the other. On Taylor's Fig. 1 these are different red boxes. The ranking is reversed by the difference between two analyses of one rock.

The Welsh-side replicate behaves the same way. Against SH33, PCM7 rpt ranks first at 0.196369 while PCM7 — the same rock — ranks fifth at 0.422311, behind PCM31, PCGF29 and PCM32.

It is worth being precise about the cause. OU10 and OU10 rpt differ in La by 3.63 against 3.67 ppm, about one percent, and both report Lu as 0.24. That one percent alone shifts La/Lu by 0.167, which supplies 76% of the squared distance between them. The noise floor here is substantially an artefact of two-significant-figure reporting of the heavy REE, amplified by the choice to divide by them.

The margins between competing sources are far smaller than the noise. SH62 is 0.550687 from PCDL25 and 0.561924 from PCAW49 — a separation of 0.011, a twentieth of the OU10 replicate distance. SH61 sits 0.023 from a decision between PCC11 and PCAW49. SH37 has four candidates within 0.08 of one another, spanning two different Welsh groups. Every assignment in Table 3 is a coin toss.

8. The assignments against the published attributions

Reading the bottom rows of Taylor's Table 3 produces the paper's actual provenance result: each Stonehenge analysis matched to a Welsh source. Compared with Bevins et al. (2021):

Carn Goedog. Bevins et al. assign twelve Stonehenge analyses to Group 1, sourced to Carn Goedog and corroborated by excavation of a Neolithic quarry there. Taylor's nearest-neighbour method sends four of the twelve to PO1. The other eight go to Carn Ddafad-las, the ground between Cerrigmarchogion and Mynydd Bach, and Group 3 outcrops.

SH45. The strongest positive attribution in Bevins et al. (2021) is SH45 to Preseli Group 2iii, the Cerrigmarchogion samples, described as a near-identical REE composition. Taylor's matrix assigns SH45 to PCC11 at d = 0.5404, with PCM2 — the nearest Group 2iii sample retained — next at 0.5735. The margin is 0.033, a ninth of the OU10 replicate distance. The method gets the source paper's cleanest result wrong, by a margin well inside its own noise.

Group 2iv. Bevins et al. explicitly exclude the outcrops between Cerrigmarchogion and Mynydd Bach as a source for any Stonehenge Group 2 dolerite. Two of Taylor's tightest matches — SH33 at 0.1964 and OU19A at 0.2328 — are both to PCM7 rpt, which belongs to that group.

Group 2v. SH62 and OU6 form one Stonehenge group in Bevins et al., linked by a shared marked positive Eu anomaly and tentatively associated with Carn Ddafad-las and Garn Ddu Fach. Taylor sends SH62 to PCDL25 at Carn Ddafad-las — agreeing with Bevins, but by a margin of 0.011 over the runner-up — and sends OU6 to PCM3 at Cerrigmarchogion, some eight kilometres west.

The spotted / non-spotted test. Bevins et al.'s Groups 1 and 3 are spotted dolerite; Group 2 is non-spotted. The distinction is visible without instruments. Seven of Taylor's twenty-two assignments cross it: six spotted Stonehenge analyses (OU8, OU12, OU14, OU19A, SH33, SH67) are sent to non-spotted Preseli sources, and one non-spotted stone (SH45) to a spotted one. No statistics are needed to see that these attributions cannot be right.

Reading Taylor's Table 3 alongside the published attributions is the most direct test available of whether the framework works, and it fails it in every case where the two make comparable claims.

9. What the distances are actually measuring

The methodology states that the distances are computed on normalised REE ratios. They are not. The calculation reproduces exactly from the untransformed values: taking OU11 and CGD2 and summing squared differences across all twelve ratios gives 0.020460, whose square root is 0.14304 — the reported 0.143. Full working is in the appendix.

This matters because the twelve variables are on wildly different scales. Across the dataset La/Lu ranges from about 13.8 to 16.8, a spread of 3.1, while La/Ce ranges from 0.355 to 0.407, a spread of 0.05. In an unstandardised Euclidean distance each variable contributes as the square of its difference, so La/Ce can contribute at most about 0.0025 to a squared distance while La/Lu can contribute over 9.

Decomposing the largest distance in the matrix, SH49 to PCDL26 at 3.7325: La/Lu supplies 75.5% of the squared distance, La/Tb a further 13.0%, La/Ho 6.1%. Three of the twelve variables account for 94.6% of the result. The remaining nine are, for practical purposes, absent. The twelve-variable multivariate distance is largely a single-variable comparison of La/Lu, dressed in twelve dimensions.

Three further problems attach to the variable set.

No chondrite normalisation. Bevins et al. work throughout with chondrite-normalised patterns, which removes the Oddo–Harkins alternation whereby even-atomic-number REE are roughly ten times more abundant than their odd-numbered neighbours. What remains after normalisation is the shape of the pattern, which is where the petrogenetic information lies. Taylor's raw La/X ratios retain the alternation, so a substantial part of what the distances measure is the periodic-table artefact that normalisation exists to remove.

No sensitivity to the Eu anomaly. Eu/Eu* is a local deviation from the value interpolated between Sm and Gd, and it is the single most discriminating parameter in Bevins et al.'s Group 2 analysis — it is what links SH62 and OU6 to Carn Ddafad-las. La/Eu conflates that local deviation with the overall LREE-to-HREE slope. Taylor's variable set cannot see a Eu anomaly at all.

Induced correlation. All twelve variables share La as numerator. Dividing one quantity by twelve others induces strong correlation among the results by arithmetic alone (Chayes, 1949). The PCA reporting 87.47% of variance on the first component is not evidence that the projection is dependable; it is largely a measurement of that induced correlation. There is also an unresolved inconsistency: an unstandardised PCA on these ratios would place considerably more than 87% on PC1, so the PCA appears to have been run on the correlation matrix while the distances were not standardised at all. The two analyses are not on the same footing and cannot corroborate one another.

Ratios of compositional parts are in any case not amenable to ordinary Euclidean geometry; centred log-ratio transformation exists for exactly this case (Aitchison, 1986). Taylor's appendix presents the absence of transformation as a methodological virtue.

10. The PERMANOVA worked. Its results were discarded

This is the section a reviewer is most likely to get wrong, and the arithmetic repays care.

The pairwise matrix contains five p-values below 0.05. All five are discarded as artefacts of small sample sizes. Checking each against the group assignments of Bevins et al. (2021):

RejectionpGroups comparedGenuinely different?
SODC1 vs PO2ii0.029Stonehenge Gp 1 vs Preseli Gp 2iiYes
SODC1 vs SLF30.035Stonehenge Gp 1 vs Stonehenge Gp 3Yes
PO2ii vs PO30.020Preseli Gp 2ii vs Preseli Gp 3Yes
PO2ii vs SLF30.009Preseli Gp 2ii vs Stonehenge Gp 3Yes
PO2iv vs SLF30.009Preseli Gp 2iv vs Stonehenge Gp 3Yes

Five out of five. Not one false positive.

The non-rejections that matter also come out right. Preseli Group 1 against Stonehenge Group 1 — the Carn Goedog attribution — gives p = 0.972. Preseli Group 3 against Stonehenge Group 3 gives p = 0.178. Stonehenge Group 1 drill cores against Stonehenge Group 1 landscape fragments, which are the same population divided by collection method, give p = 0.184.

The test is underpowered and misses real differences: Preseli Group 1 against Preseli Group 3 fails to reject at p = 0.595 when it should not. But every difference it detected is real, and it recovered the published group structure wherever it had the samples to do so. Taylor deleted precisely the entries carrying the signal.

The showcase example the paper offers for discarding them is worth examining. It cites SOF1 against SODC1, where a pseudo-F of 5.317 accompanies a non-significant p of 0.197, as evidence that the pairwise results are unreliable. But SOF1 and SODC1 are both Stonehenge Group 1 — the same geochemical population, separated only by whether the sample was an orthostat fragment or a drill core. Failing to reject is the correct answer, and it says nothing about provenance. It is also an unavoidable answer: with group sizes of 1 and 4 there are only five distinct partitions of the data, so the smallest attainable p-value is 0.200, and 0.197 is the floor of the test.

That floor effect runs through every comparison involving the singleton groups. In a permutation test with group sizes n₁ and n₂ there are C(n₁+n₂, n₁) distinct partitions, and no p-value below 1/C(n₁+n₂, n₁) can be returned however large the true difference:

ComparisonGroup sizesPartitionsMinimum possible pReported pReported F
SOF1 vs SODC11, 450.2000.1975.317
SOF1 vs PO2iii1, 340.2500.24537.24
SOF1 vs PO2iv1, 340.2500.25192.41
SODC1 vs PO2ii4, 3350.0290.02916.3
SLF3 vs PO2ii7, 31200.0080.0099.532

The rows involving SOF1 and SLF2vi should be deleted rather than interpreted: PERMANOVA on a group of one has no within-group variance to estimate, and those groups exist only because of the sampling-category split described in section 3. The rows at the lower floor are the opposite case — the most extreme outcome the design permits, and correct.

The paper's diagnosis of the problem also needs correcting. It argues that a single critical F value of 2.1532 cannot apply to pairs of differing sizes because the degrees of freedom differ. The deeper point is that the pseudo-F in PERMANOVA does not follow an F distribution at all — that is the reason for permuting (Anderson, 2001). There is no critical F for any pair, at any sample size. The recommendation to work from permutation p-values is right; the reasoning offered for it is not.

11. Pooling for SIMPER

The SIMPER analysis pools all Stonehenge analyses into one group and all Welsh analyses into another, on the grounds that this removes small-sample noise. The paper's own Table 2 records Stonehenge Group 1 differing from Stonehenge Group 3, and Preseli Group 2ii differing from Preseli Group 3. Neither pool is internally homogeneous, by the paper's own test. Averaging each into a single mean profile produces two composite figures corresponding to no actual rock, whose agreement is guaranteed by the mixing rather than by shared origin.

The appendix result confirms it: a pooled pseudo-F of 0.074 means that variation between the two groups is about 7% of the variation within them. That is not a signature of common origin. It is a statement that the grouping explains essentially nothing, which is what happens when heterogeneous populations are pooled.

SIMPER (Clarke, 1993) is also built on Bray–Curtis dissimilarity, defined for abundance data; applied to element ratios it has no clear interpretation. The paper's own observation that the high-magnitude ratios dominate the dissimilarity profile is a symptom of this rather than a finding.

Underlying all of it is the load-bearing assumption that failing to reject H₀ is evidence for H₀. It is not, and the stated justification — that sample sizes are small and variances tight — describes the conditions under which a non-significant result is least informative.

12. Errors of fact

The text states that SH65 matches CGD1 at d = 0.1585. In Table 3 the distance from SH65 to CGD1 is 1.3004. The figure 0.1585 belongs to PCM30 — a different outcrop in a different group, and more than eight times nearer than CGD1. One of the paper's two headline pairings names the wrong source.

The note beneath Table 3 states that only two or three pairings reject the null hypothesis. Table 2 contains five p-values below 0.05, two of them within-side comparisons.

Figure 1 is an annotated reproduction of a published geological map of southwest Wales — its own legend refers to sources proposed by Thomas — presented without attribution. It is not the map from Bevins et al. (2021); its origin should be established and credited.

Separately, the submitted document contains material that reads as unedited drafting: passages addressed in the second person within a first-person paper, unrendered LaTeX in the body and appendix, two alternative titles both retained, and one sentence in the appendix describing an intention to draw the reader's attention away from a discrepancy in the author's own figure. That sentence cannot be what the author meant to publish, and should be removed.

13. What a working version would look like

The instinct behind the sample-level analysis is sound. Nearest-neighbour matching in geochemical space is a legitimate provenance technique, and preferring it to group-level tests when groups are small is a reasonable judgement. The execution is what fails. A version that would carry weight would need:

  • Full attribution of the dataset to Bevins et al. (2021), with the analytical method correctly described.
  • The complete dataset, including Group 2i and SH42. The samples that don't fit are the ones that establish resolution.
  • Chondrite-normalised concentrations, centred-log-ratio transformed, rather than twelve raw ratios sharing a numerator.
  • Standardisation before any distance is computed, so that all variables contribute and the result is not a proxy for La/Lu.
  • Shape-sensitive parameters — Lan/Smn, Gdn/Ybn, Eu/Eu*, and the λ coefficients of O'Neill (2016) — rather than a set blind to the Eu anomaly.
  • The published geochemical groups, not sampling categories. The pairwise structure is the provenance signal, and the rejections are the informative entries, because exclusion is what geochemistry can establish.
  • Candidate sources outside Preseli, so that specificity can be demonstrated rather than assumed.
  • A permutation null for the nearest-neighbour distances, answering whether the observed minima are smaller than would arise by chance from twenty candidate outcrops.
  • The replicate distance reported as the resolution limit, with any assignment whose margin over the runner-up falls below it declared undetermined.

Applied honestly, that last step alone would leave most of the assignments in Table 3 unresolved. That is not a failure of the study; it is the correct result for this dataset, and stating it would be a genuine contribution.

14. Why this matters beyond one paper

Bevins et al.'s Fig. 4 shows the Stonehenge Group 1 and 3 REE patterns falling inside the envelope of the Preseli patterns, with one exception — SH42. They present this as consistency with a Preseli origin and as a limit on what REE alone can resolve, which is why the group definitions rest on compatible-element geochemistry (Bevins et al., 2014) and why the rhyolite work required zircon chemistry. Their section on sample sizes and analytical homogeneity states in advance that differences of ten percent to a factor of two are to be expected between Preseli field samples and Stonehenge drill cores, and that exact matching should not be looked for.

Taylor's headline result is that same overlap, with SH42 and Group 2i removed, restated as proof of identity and presented as a novel framework. The source paper's stated conclusion — that these data are consistent with Preseli origin but cannot on their own resolve outcrops — has been converted into its opposite by removing the uncertainty rather than by adding information.

The pattern is not unusual. "No statistically significant difference" is quietly doing the work of a positive finding across a good deal of provenance literature, and the smaller the sample the more confident the claim tends to become. The framework in this paper is an unusually explicit version of a common move: replacing tests that can fail with descriptive statistics that cannot, and describing the absence of an error estimate as independence from sample size.

The replicate check offers a cheap and general guard against it. Most analytical programmes run duplicates already — Bevins et al. ran three. Computing the distance between two analyses of the same sample, and refusing to report any assignment whose margin is smaller than that distance, costs nothing and would prevent a great deal of overclaiming. It is the one thing this paper's data do establish, and the source data supplied the means to establish it.


Appendix: worked arithmetic

A1. Reproducing d(OU11, CGD2) from untransformed ratios.

RatioOU11CGD2DifferenceSquared
La/Ce0.3821730.3731810.0089920.0000809
La/Pr2.4645672.4093960.0551710.0030438
La/Nd0.4735250.4602560.0132690.0001761
La/Sm1.4626171.4360000.0266170.0007085
La/Eu3.5568183.626263−0.0694450.0048226
La/Gd1.2987551.2913670.0073880.0000546
La/Tb7.1136367.0392160.0744200.0055383
La/Dy1.0468231.052786−0.0059630.0000356
La/Ho5.3965525.439394−0.0428420.0018354
La/Er1.9562501.983425−0.0271750.0007385
La/Yb2.1006712.124260−0.0235890.0005564
La/Lu14.90476014.958330−0.0535700.0028698
ÎŁ0.0204605

√0.0204605 = 0.14304, matching the reported 0.143. The distances are computed on untransformed ratios, not normalised values.

A2. Decomposition of the largest distance, d(SH49, PCDL26) = 3.7325.

RatioDifferenceSquared% of total
La/Lu3.2427610.515575.5
La/Tb1.347841.816713.0
La/Ho0.924240.85426.1
La/Eu0.517400.26771.9
La/Yb0.385920.14891.1
La/Er0.375960.14131.0
remaining six0.18701.3
ÎŁ13.9313

√13.9313 = 3.7325. Three variables account for 94.6% of the result.

A3. Replicate distances, computed from Bevins et al. (2021) Table 2.

Ratios were formed as in Taylor's Table 1 and the same untransformed Euclidean distance applied.

PairLa (ppm)Lu (ppm)Δ(La/Lu)Σ of squaresd
OU10 / OU10 rpt3.63 / 3.670.24 / 0.240.166670.0364400.1909
PCM7 / PCM7 rpt3.04 / 3.100.21 / 0.210.285720.0871120.2952
PCGF27 / PCGF27 rpt5.16 / 5.200.32 / 0.320.125000.0217500.1475

In each case the La/Lu term supplies the large majority of the total: 76%, 94% and 72% respectively. All three exceed the paper's headline match of d = 0.143.


References

Aitchison, J., 1986. The Statistical Analysis of Compositional Data. Chapman & Hall, London.

Anderson, M.J., 2001. A new method for non-parametric multivariate analysis of variance. Austral Ecology 26, 32–46.

Bevins, R.E., Ixer, R.A., Pearce, N.J.G., 2014. Carn Goedog is the likely major source of Stonehenge doleritic bluestones: evidence based on compatible element geochemistry and Principal Component Analysis. Journal of Archaeological Science 42, 179–193.

Bevins, R.E., Pearce, N.J.G., Ixer, R.A., 2021. Revisiting the provenance of the Stonehenge bluestones: Refining the provenance of the Group 2 non-spotted dolerites using rare earth element geochemistry. Journal of Archaeological Science: Reports 38, 103083. https://doi.org/10.1016/j.jasrep.2021.103083

Chayes, F., 1949. On ratio correlation in petrography. Journal of Geology 57, 239–254.

Clarke, K.R., 1993. Non-parametric multivariate analyses of changes in community structure. Australian Journal of Ecology 18, 117–143.

O'Neill, H.St.C., 2016. The smoothness and shapes of chondrite-normalized rare earth element patterns in basalts. Journal of Petrology 57, 1463–1508.

Parker Pearson, M., Pollard, J., Richards, C., Welham, K., Caswell, C., French, C.A.I., Shaw, D., Simmons, E., Stanford, A., Bevins, R.E., Ixer, R.A., 2019. Megalithic quarries for Stonehenge's bluestones. Antiquity 93, 45–62.

Thomas, H.H., 1923. The source of the stones of Stonehenge. Antiquaries Journal 3, 239–260.

Sunday, 26 July 2026

Who Built Stonehenge and When? Meet the Real Builders | Stonehenge Access All Areas, Ep 7

Counting R1b: an audit of "The Great Ancient DNA Illusion"

On 25 July 2026, Robert John Langdon published a post arguing that the Bell Beaker migration model rests on a statistical illusion. Its central evidence is a table of 43 pre-Bell Beaker R1b individuals compiled, he says, from the Allen Ancient DNA Resource, together with a Mesolithic R1b individual from Aveline's Hole in Somerset said to move the starting point of British population history back by several millennia.

I have re-run the query against the source file. The table can be checked, and it does not hold. The Aveline's Hole claim has a specific and traceable explanation, and it is not the one offered.

What follows is the audit. The dataset is the AADR 1240k annotation file, release v66.p1, which anyone can download and reproduce.

Verify before believing

What the post gets right

It is worth being clear about this first, because the post's opening three sections are sound and the rebuttal does not depend on disputing them.

Ancient DNA is fragmentary. Genomes are partially reconstructed rather than read. Coverage varies by orders of magnitude between individuals who appear side by side in a published table. Radiocarbon results are probability distributions, not dates. Bayesian chronological models are conditional on their priors, and a published range can be more precise than the underlying measurements alone would justify. Databases are samples, not censuses, shaped by preservation, excavation history and research priorities.

All of that is true. None of it is news — these caveats appear routinely in the supplementary material of the papers the post is criticising — but stating them for a general readership is a service rather than a fault.

The post is also correct that R1b existed in Europe before the Bell Beaker period. That has been published, mainstream, uncontroversial data since 2017.

The difficulty is what is built on top.

The category error

"R1b" is used throughout the post as though it named a population. It does not. It names a macro-haplogroup roughly eighteen to twenty thousand years deep, with two primary branches, and the distinction between its subclades is the entire substance of the question at issue.

The individual from Villabruna in northern Italy, dated to about 14,000 years ago, is R1b1a-L754. The Iron Gates Mesolithic individuals from Serbia and Romania are L754-derived. The Balkan Chalcolithic individuals from Pietrele, which supply eight rows of the post's own table, are basal or V88-type lineages. Blätterhöhle in Germany is V88.

None of these are ancestral to R1b-M269 > L51 > P312, the clade that expands across north-western Europe after 2500 BCE and that accounts for the overwhelming majority of modern British paternal ancestry. They are collateral branches. Their presence in Mesolithic and Neolithic Europe is a well-known fact that no researcher has ever disputed, and it bears on the Beaker question in roughly the way that the presence of wolves in Pleistocene Europe bears on the origin of the Labrador retriever.

The post's argument requires the reader not to notice this. Once noticed, the argument does not survive it.

There is a nomenclature trap here that catches honest readers too. AADR haplogroup strings have shifted across ISOGG releases, and "R1b1b" has denoted different nodes at different times. Anyone working from the strings rather than the primary publications will find them genuinely confusing. That is a reason for care, not a defence of the conclusion.

The chronological conflation

The post treats "pre-Bell Beaker" as though it meant "pre-steppe." It does not, and the gap between the two is where a fifth of its evidence sits.

Nine of the 43 are Bohemian: Plotiště nad Labem, Obříství, Vliněves, Stadice, Konobrže, dated between roughly 2919 and 2636 calBCE and carrying P312-derived subclades. These are Corded Ware burials.

Papac and colleagues, publishing in Science Advances in 2021, reported that Corded Ware appeared in Bohemia by about 2900 BCE and that R1b-L151 was already the most common Y-lineage among the earliest Corded Ware males there — six of eleven — with P312 most likely diversifying from within that pool. The same study documents the subsequent volatility with some precision: P312 rises to complete fixation in late Bell Beaker Bohemia and then falls to twenty per cent in preclassical ĂšnÄ›tice, implying a minimum eighty per cent influx of new Y-lineages at the onset of the Early Bronze Age.

The Bohemian individuals in the post's table are therefore not counter-evidence to the steppe migration. They are its evidence, dated and labelled, filed under the wrong heading.

Aveline's Hole

Chapter 4 of the post rests on the identification of an R1b lineage from Aveline's Hole dating to the Mesolithic. No such individual appears in the post's own appendix. The only two British entries are Neolithic.

The sample exists. It is I3004, from Aveline's Hole, Burrington Combe, directly dated to 8606–8300 calBCE (OxA-34339), published in Brace et al. 2019. Its Y-haplogroup in AADR is R1b1a1b1a1a2a5a~, terminal SNP R-Y11281.

Its assessment in AADR is CRITICAL.

The hapConX X-chromosome contamination estimate is [0.212, 0.364] — between twenty-one and thirty-six per cent. The documented threshold at which AADR marks a sample CRITICAL or FAIL is a lower bound above 0.03. This sample exceeds that by a factor of seven. The second warning on the record is high.popgen.heterozygosity, the other standard signature of contamination. The individual carries 112,612 SNPs on 1240k targets.

The Y-call itself is the tell. R-Y11281 sits deep within P312 — which is what the modern British population overwhelmingly carries. A low-coverage British Mesolithic sample with a fifth to a third of its X-chromosome reads coming from somewhere else, returning a deeply derived P312 subclade, is not a discovery about the Mesolithic. It is a description of laboratory contamination, and the curators recorded it as such.

For completeness, the other Aveline's Hole individuals in AADR are I3006 (Mesolithic, female), I3005 and I3010 (both Early Neolithic, both female). This matches Brace et al. 2019 and the subsequent site-specific reassessment in the Proceedings of the University of Bristol Spelaeological Society: four individuals yielding genomic data, two Mesolithic with the expected Western Hunter-Gatherer signature, two Early Neolithic with Aegean farmer ancestry, separated by nearly five millennia and identified only when the crania were directly dated.

That last detail deserves emphasis, because the post has the moral of the story backwards. Aveline's Hole is the textbook case of undated cave material proving to be thousands of years younger than assumed. It is an argument for stringent chronological control, not against it.

The count

The post states that it examined every published prehistoric male dated before 2500 BC in the AADR, and reports 1,351 individuals of whom 43 are R1b, an observed frequency of 3.18 per cent.

https://prehistoric-britain.co.uk/ancient-dna-illusion-archaeological-facts accessed 14:41 26/7/2026


I cannot reproduce those figures. Deduplicating on Individual ID as the AADR README instructs — the same individual can appear under several Genetic IDs, and counting rows inflates everything — and taking ancient males with a mean date at or before 2500 BCE:

FilterMalesR1b% of all males% of males with a Y call
Global2,60935713.6814.30
Europe2,09834216.3016.72
Europe, passing assessments only1,90631916.7417.08
Europe excluding Russia, Ukraine, Moldova, Belarus1,4521117.647.87
Europe, rows not deduplicated2,61149919.1119.64

The fourth row is the closest approximation I can construct to the post's denominator. Restricting to non-steppe Europe gives 1,452 males, near enough 1,351 that something of this kind was probably done.

That same filter contains 111 R1b individuals, not 43.

Sixty-nine are missing from a table captioned as listing all confirmed pre-Bell Beaker R1b individuals: six Iron Gates Mesolithic individuals from Serbia, eight from Latvia across the Kunda and Narva groups, twelve from Bulgaria including four from Varna, nine from Denmark, three further Czech Corded Ware, and others from Spain, Italy, Germany, Hungary, Slovakia, Sweden, Switzerland, Poland and Estonia.

The observed frequency is therefore not 3.18 per cent under any reading I can construct. It is 7.6 per cent on the post's own apparent geography and 16.3 per cent for Europe as a whole.

This does not support the post's argument. It damages it. The bulk of pre-2500 BCE European R1b sits in Yamnaya burials from Samara, Kalmykia, Orenburg, Rostov and Moldova, in Khvalynsk and Ekaterinovka, in Afanasievo, and in Corded Ware. Making the number larger makes the steppe expansion more visible, not less. The post's central statistic understates its own database by a factor of two and a half, and correcting the error strengthens the model it was assembled to refute.

The arithmetic that follows

Even taking 3.18 per cent at face value, the extrapolation to "approximately 8,000 to 16,000 R1b individuals" does not work.

A Y-haplogroup frequency is a frequency among males. Applying it to a total population of 250,000 to 500,000 counts everybody twice; the internally consistent figure would be roughly 4,000 to 8,000. The denominator spans about 7000 to 2500 BCE, mostly Neolithic and Chalcolithic, while the population estimate is Mesolithic — and European population grew by an order of magnitude across that interval, so combining them is not conservative but incoherent. Those population estimates themselves span more than a factor of two in the published literature and rest on sparse ethnographic analogy and climate-envelope modelling.

More fundamentally, multiplying any non-zero frequency by a large population yields a large number. That establishes nothing about descent. The Beaker model is not a claim that no R1b existed before 2500 BCE. It is a claim about which subclade expanded and left descendants. Since most of the 43 belong to lineages with negligible modern Western European paternal descent, the headcount is irrelevant to the proposition it is deployed against.

The post also invokes sampling bias in one direction only. If the database undercounts R1b males, it equally undercounts I2a, G2a and everyone else. A frequency is a ratio, and bias in the denominator does not preferentially inflate the numerator.

The 43

All 43 identifiers exist in v66.p1. The table is not fabricated, and it is worth saying so.

On AADR's own annotations, however:

  • Eight carry ASSESSMENT = Questionable: I6912, PNL001, I18101, JK2804, ATP3, I3035, I2611, OC.
  • Thirteen have contextual dates only, with no direct radiocarbon determination. The AADR file distinguishes these explicitly; the post's table preserves the distinction in its "cal BCE" versus "BCE" notation without remarking on it.
  • Two carry the outlier group label England_N-o — both British entries.
  • VLI011 and VLI015 are father and son, recorded as a first-degree pair. Two of the 43 are one observation. KON003 has two second-degree relatives at the same site; I14176 has a 2.5-degree relative.
  • Twenty-five both pass AADR quality control and are directly dated. Twenty-four carry the literal Pass value; the twenty-fifth, I1590 from Blätterhöhle Cave, is MERGE_PASS.

"Forty-three confirmed" is not a description the source file supports. The point matters because the post's rhetorical structure depends on accumulation — that independent discoveries across many countries and laboratories cannot all be anomalies. Kinship, QC flags and undated contexts all reduce the number of independent observations.

The two British entries deserve individual attention, since they are the ones bearing on Britain.

I2611, from Summerhill, Blaydon, directly dated 3092–2905 calBCE and called R-L21, is the single most striking item in the table. Its AADR warning field reads: carbon date is unexpected for the archaeological context and genetic profile, with hapConX at [0.016, 0.038]. It is grouped as England_N-o, an outlier, and published in Patterson et al. 2021. The curators identified the anomaly before the post did, and recorded their view of it.

I3035, from Fox Holes Cave, has a contextual date only of 4000–3500 BCE and a U106-derived call. Its warnings record ANGSD contamination at [0.015, 0.028], hapConX at [0.014, 0.022], a library with technical problems, and that the observed and expected predicted dates differ by more than a thousand years. Another cave assemblage, another chronological mismatch.



What the evidence actually shows

The Beaker model does not rest on Y-haplogroups. It rests on genome-wide ancestry proportions, and the post does not mention them once.

Olalde and colleagues, in Nature in 2018, estimated that around ninety per cent of Britain's gene pool was replaced within a few centuries of the Beaker horizon, on the basis of autosomal ancestry. The alternatives the post proposes in its seventh section — gradual admixture, regional survival, cultural diffusion, multiple episodes — are not alternatives to the mainstream picture; they are components of it. Papac's Bohemian study documents Corded Ware males assimilating females of diverse local backgrounds. Work on Bronze Age Orkney has shown Neolithic I2a lineages persisting locally long after Beaker-derived ancestry arrives elsewhere in the genome.

But none of the four alternatives explains the observation the model was built to explain, and one of them cannot in principle. Cultural diffusion does not transmit autosomes. Pottery styles travel without people; steppe ancestry does not.

There is a final irony in the post's framing. It presents the Beaker migration model as an entrenched orthodoxy of more than two decades, self-reinforcing and resistant to revision. The model dates from 2015 and 2018. Before roughly 2010 the dominant view held R1b to be a Palaeolithic or Mesolithic Western European lineage — close to the position the post presents as suppressed heterodoxy. The consensus moved because the evidence moved, which is precisely the process the post claims does not occur.

Method, and where I might be wrong

The audit uses the AADR 1240k annotation file, release v66.p1, dated 8 June 2026. The post cites "AADR v66.1," which is not a version string the project has issued; the sequence runs v66.0 (12 April 2026), subsequently decommissioned, then v66.p1. The difference between those two releases concerns 161 modern Papuan individuals and cannot affect any count of prehistoric European males, so nothing substantive turns on it.

Two filter judgements are mine and are contestable. I deduplicated on Individual ID rather than counting rows, and I used the mean-BP date field rather than range endpoints. Both follow the README's guidance, but neither is the only defensible choice, which is why the sensitivity table above shows several variants rather than a single number. The gap between 43 and 111 survives all of them.

I could not reproduce 1,351 exactly, and it is possible I have misread the filter that produced it. The remedy is straightforward: publish the query. A count presented as a direct read of a public database, in a post whose thesis is that archaeologists hide their assumptions inside models, should be reproducible by anyone who downloads the file.

Sources

  • Brace, S. et al. (2019) Ancient genomes indicate population replacement in Early Neolithic Britain. Nature Ecology & Evolution 3, 765–771.
  • Schulting, R., Booth, T., Brace, S. et al. (2019) Aveline's Hole: an unexpected twist in the tale. Proceedings of the University of Bristol Spelaeological Society.
  • Olalde, I. et al. (2018) The Beaker phenomenon and the genomic transformation of northwest Europe. Nature 555, 190–196.
  • Papac, L. et al. (2021) Dynamic changes in genomic and social structures in third millennium BCE central Europe. Science Advances 7, eabi6941.
  • Patterson, N. et al. (2022) Large-scale migration into Britain during the Middle to Late Bronze Age. Nature 601, 588–594.
  • Mathieson, I. et al. (2018) The genomic history of southeastern Europe. Nature 555, 197–203.
  • González-Fortes, G. et al. (2017) Paleogenomic evidence for multi-generational mixing between Neolithic farmers and Mesolithic hunter-gatherers in the Lower Danube basin. Current Biology 27, 1801–1810.
  • Fu, Q. et al. (2016) The genetic history of Ice Age Europe. Nature 534, 200–205.
  • Haak, W. et al. (2015) Massive migration from the steppe was a source for Indo-European languages in Europe. Nature 522, 207–211.
  • Dulias, K. et al. (2022) Ancient DNA at the edge of the world: continental immigration and the persistence of Neolithic male lineages in Bronze Age Orkney. PNAS 119, e2108001119.
  • Allen Ancient DNA Resource, v66.p1 (8 June 2026), 1240k annotation file. Harvard Dataverse.

Playing with the Stones: The New Locative Games at Avebury

 Avebury has always been a place that rewards slow looking. The scale of the bank and ditch, the awkward tilt of individual sarsens, the way the landscape folds around the circle — these things reveal themselves gradually. Yet for many visitors the experience can still feel static: walk, look, read a panel, move on. In 2025 a group of researchers and designers tried a different approach. They turned the monument into a series of short, location-aware smartphone games.

The result is an anthology of roughly twenty free apps collectively known as the Avebury Adventures (or the Avebury Anthology). They form part of the larger European research project LoGaCulture — Locative Games for Cultural Heritage — run by Bournemouth University, the University of Southampton and the National Trust, with funding from the EU Horizon programme and UK Research and Innovation.


What the games actually do

Most of the apps use the phone’s GPS so that content appears only when you are standing in a particular place. Some add augmented reality, so virtual objects or characters sit on top of the real stones through the camera. Others are simpler narrative or puzzle experiences that unfold as you walk.

Examples include:

  • Ages of Avebury (University of Southampton) — you take the role of a surveyor searching for missing stones and uncover the story of a disappearing archaeologist, with an AR reconstruction of a complete circle.
  • Henge Hunts — follow trails and “hunt” the giant cattle, bears and wolves that once lived in the Neolithic landscape.
  • Replicas, Uncovering Avebury, Avebury Research Challenge — more archaeological in flavour, involving collecting fragments, helping a fictional Professor Stone, or digging up knowledge around the stones.
  • Lighter or more playful titles such as Don’t Harm the Sheep!, Journey Home (aimed at children), and even a cyberpunk-themed cyber@avebury.

There are also creative experiences where you collect stickers or visual elements to design your own Avebury postcard, and a couple of titles that can be partly experienced away from the site.

The games launched with QR-code boards in the visitor barn in late July 2025. Many remain available on both major app stores.

The thinking behind it

Locative games sit at the intersection of several research strands. First is the long-standing question of how people develop a sense of place and presence at heritage sites. Traditional interpretation often treats the visitor as a passive recipient of information. Locative and mixed-reality experiences try to make the visitor an active participant whose physical movement through the landscape drives the narrative.

Researchers in this field talk about “entanglement” — the idea that the digital layer, the physical stones, the visitor’s body and the stories become interwoven rather than sitting side by side. One of the 2026 papers from the project, The Relational Entanglement Diary, even proposes a method for evaluating these multi-layered experiences.

There is also a practical research question: can playful interaction increase engagement, especially among younger visitors or those who find conventional interpretation dry, without trivialising the archaeology? Early results from the Avebury trials suggest many people stay longer, notice details they would otherwise have missed, and report a stronger emotional connection to the site. Whether that translates into deeper historical understanding is harder to measure and remains an open question the academic papers are still exploring.

The technical side draws on earlier work in hypertext narrative, location-based storytelling systems, and the design of authoring tools that let writers and artists create these experiences without needing to be programmers. The Bournemouth team in particular developed frameworks (including the LUTE toolkit) to make locative game creation more accessible.

Where to find them

The simplest way to browse the Android versions is the Google Play search the project itself recommends:

https://play.google.com/store/search?q=avebury&c=apps

On iOS, search the App Store for “Avebury” together with titles such as Replicas, Henge Hunts, Ages of Avebury or cyber@avebury. Most of the apps are credited to LoGaCulture / Bournemouth University or the University of Southampton.

The central project site is https://logaculture.eu/. It contains background on the wider research, design patterns, and links to the games and related tools.

Key people involved include Dr Charlie Hargood (Bournemouth University, games technology), Professor David Millard (University of Southampton), and Dr Ros Cleal (National Trust curator at Avebury).

The academic papers

Two papers presented at the 2026 ACM Creativity & Cognition conference discuss the Avebury work in more formal terms:

  • In Search of Lost Times: Reimagining Mixed Reality for an Ancient Site (Rimington, Baker, Blount, Jones, Jordan, Malinov & Millard) — DOI: 10.1145/3803784.3809283
  • The Relational Entanglement Diary: A Novel Method for Evaluating Entangled Transmedia Engagement (Ferreira et al.) — DOI: 10.1145/3803784.3809275

They are written for a specialist audience and use the dense theoretical language common in interaction design research. The practical outcome — the apps themselves — is more approachable.