The initial question or criticism exists independently of Deardorff’s TJ analysis.
Deardorff’s probability argument, without the jargon
What does PHoax mean?
PHoax is Deardorff’s estimate of which direction a Matthew–TJ comparison points: toward the TJ being written from Matthew, or toward Matthew being written from an earlier TJ-like source.
Why these comparisons are unusually valuable
The criticism of Matthew often came first
Deardorff did not simply invent every problem after reading the TJ. A substantial and clearly labeled part of the comparison set begins with criticisms published by Gospel scholar Francis W. Beare in The Gospel according to Matthew. Deardorff then asked whether the same criticism also applies to the corresponding TJ passage.
He compares the corresponding wording and asks whether the same problem remains.
If the TJ avoids or explains the problem, Deardorff judges how strongly that weighs on literary direction.
The provenance is visible in Deardorff’s key: blue passages use Beare; green passages use other Gospel scholars; red passages are Deardorff’s own criticisms developed with TJ hindsight. Beare therefore supplies an important independent control, but not every observation or numerical PHoax estimate.
Read Deardorff’s original explanation and color key →Read one value
The scale runs from 0 to 1
Deardorff normally estimated to the nearest 0.05. A value such as 0.10 is his judgment that the comparison strongly weighs against the hoax explanation; 0.60 leans modestly toward it.
The method in three steps
How separate comparisons become one result
- 01
Score one comparison
Read the Matthew and TJ passages, Deardorff’s stated problem, and his proposed explanation. Assign one p value to that evidence unit.
- 02
Keep evidence units separate
Only comparisons treated as conditionally independent are supposed to receive separate scores. Related verses are omitted or grouped together.
- 03
Multiply the odds
Each value is converted into hoax-versus-no-hoax odds. Those odds are multiplied, then converted back into a probability.
For values p₁, p₂ … pₙ
P = (p₁ × p₂ × … × pₙ) ÷ [(p₁ × p₂ × … × pₙ) + ((1−p₁) × (1−p₂) × … × (1−pₙ))]A 0.50 contributes equal weight to both sides, so it does not change the combined result.
Try the calculation
Deardorff’s simple example
His technical page uses three estimates: 0.10, 0.20, and 0.40. Each leans against a hoax, but with different strength. Combined by his formula, they produce approximately 0.018—about a 1.8% hoax probability under the model.
This demonstrates the arithmetic. Whether the answer is persuasive still depends on whether the original scores and independence assumptions are well justified.
Use values greater than 0 and less than 1. Try adding 0.50 to see why a neutral item has no effect.
The important distinction
What the final number does—and does not—say
It does say
- how strongly Deardorff’s chosen scores combine under his two-hypothesis model;
- how repeated evidence pointing in the same direction can accumulate;
- which passages he regarded as especially difficult for the hoax hypothesis.
It does not automatically say
- that the individual scores were measured objectively;
- that every comparison is truly independent of the others;
- that hoax and direct Matthean dependence are the only conceivable explanations;
- that the reported result is an experimentally observed error rate.
Deardorff’s reported all-chapter result
Approximately
10−111
After first combining passage estimates into chapter results and then combining all 28 Matthew chapters, Deardorff reported a hoax probability near 10−111. Read this as the output of his scoring model. The extraordinary size of the claim makes independent review of the inputs, model choices, and dependence between comparisons essential.
September 21, 2026 · neutral AI audit
Substantial research. Suggestive direction. Not independent proof.
The audit found that Deardorff’s result is not created solely by criticisms he originated: the strict external-scholar subset remains extremely directional when his scores and model are retained. A separate neutral rescore, however, did not reproduce that certainty.
193 Beare/other-scholar units, but still interpreted and numerically scored by Deardorff.
Thirty-six disclosed test comparisons; sensitivity range 0.20–0.74 crosses the undecided point.
The corpus merits serious study, but it does not yet authenticate an ancient TJ or establish an objective hoax probability.
Why the difference? The observations can be worthwhile even when their literary direction is reversible. Deardorff’s scoring choices and the assumption that hundreds of judgments contribute independent evidence drive much of the astronomical result.
Read the full audit, item ledger, adverse findings, and alternate scores →Go deeper