# Protocol for independent scoring

Version 1, September 21, 2026. This is a prepared protocol, not a completed panel study.

Recruit several reviewers who state whether they initially accept, reject, or remain undecided about TJ authenticity. Those attitudes are descriptors, not scoring instructions. Include competence in Gospel textual criticism and German; reserve Aramaic claims for someone able to assess the asserted source-language construction. The organizer should obtain scores independently before holding discussion.

Use a fixed edition. Start with photographed or otherwise verifiable 1978 German TJ wording and the specified RSV text, with relevant Greek and German Bible parallels available. Keep 1992, 2001, 2007, and 2011 variants in a separate apparatus. Do not let a later correction count as a pre-criticism prediction. Record first demonstrable attestation of each distinctive TJ reading.

For an actually masked first pass, prepare a fresh packet showing the compared texts, context, and references, without Deardorff's scores, aggregate, conclusion, or argumentative prose. The existing `review-packet-masked.json` hides numeric scores only and is **not** a sufficiently blinded packet. Randomize order, preserve an organizer-only ID key, and reveal the scholar provenance after the initial observation judgment. Keep matched comparisons that do not favor TJ and a prespecified control set from ordinary paraphrases or adaptations.

Each reviewer should record three distinct responses:

1. Observation accuracy: supported / partly supported / unsupported / unresolved. State the exact wording and whether the purported problem depends on translation, genre, historical assumptions, or a supposed author's beliefs.
2. Explanation: rate how the observation bears on direct modern adaptation, earlier TJ-like source, common source, composite text, and mixed direction. State a best counterexplanation. “Cannot distinguish” is permitted and must not be treated as a failure to assess the evidence.
3. Weight: if willing, provide an equal-prior two-model score and a judgment interval, with a likelihood-ratio interpretation. If not, leave the numeric score unassigned. Do not force every genuine textual observation to favor one direction.

Annotate overlapping evidence before seeing aggregates. Use both local passage groups and shared mechanisms, allowing an item to have multiple tags. Shared mechanisms include translation vocabulary, a repeated doctrinal replacement, an expanded motivation, a scene transition, and dependence on the same assumed editor profile. Adopt a prespecified rule for overlapping groups or report alternative plausible structures.

Publish each person's raw judgments, source corrections, missing data, and aggregates. Report three-way directional agreement, absolute score differences, and an agreement statistic with uncertainty only when the data and sample size support it. Assess agreement separately for observation accuracy and direction. A shared belief can produce agreement without correctness, so agreement is a reliability measure rather than authentication.

For an optional pooled sensitivity model, average reviewer log likelihood ratios per evidence unit, rather than multiplying reviewers as independent evidence about the text. Reviewers saw the same evidence. Then apply the declared dependence adjustment to the units. Publish each reviewer's result as well as the pooled model, and include leave-one-reviewer-out analyses. Missing scores remain missing; they are not silently replaced with 0.5.

Predeclare the sampling frame or score all counted units. The current 36-unit coverage sample is useful for a pilot and disagreement analysis, but it is not a random sample and cannot estimate a corpus-wide proportion with ordinary sampling confidence intervals.

Success is a reproducible, falsifiable account of which comparisons survive source checking and how they discriminate among histories. Neither religious agreement nor a desired small aggregate is a success criterion.
