Deliverable 03 Specification · v1.0 MODEL: 3PL · SCALE: LOGITS

Psychometric equivalence specification

Giving every candidate a different paper is only defensible if the papers are the same test. This document sets out exactly what "the same test" means, how it is enforced before delivery, and what evidence would prove the claim false.

ModelThree-parameter logistic, hierarchical over item families
ToleranceΔb ≤ 0.05 · Δa ≤ 0.01 · Δt ≤ 4 s
Form drift< 0.25 marks across the whole ability range
EnforcementPre-delivery rejection, not post-hoc apology

Section 01The claim, stated so it can be tested

Claim. For any two candidates i and j of identical ability θ, sitting two different generated forms of the same examination, the expected score difference attributable to form is smaller than the measurement error of the instrument by at least an order of magnitude.

This is deliberately a statement about expected scores on the ability scale, not about surface similarity. Two papers that look different are not necessarily unequal, and two that look identical are not necessarily equal. Only the item parameters settle it.

What would falsify it. A reliably detectable relationship between which variant a candidate received and the score they obtained, after conditioning on ability. We test for exactly this after every administration and publish the result — §10.

Section 02The measurement model

Every item is described by a three-parameter logistic model. The probability that a candidate of ability θ answers item k correctly is:

Pk(θ) = ck + (1 − ck) / (1 + e−1.7·ak·(θ − bk)) a — discrimination: how sharply the item separates candidates around its difficulty.
b — difficulty, on the same logit scale as ability, so "as hard as a candidate at θ = 0.38" is a precise statement.
c — pseudo-guessing asymptote: the probability a candidate of very low ability still answers correctly, driven by option count and distractor plausibility.
The constant 1.7 places the logistic on the normal-ogive metric, which is the convention in Indian and international large-scale testing.

Two items are psychometrically equivalent when their (a, b, c) triples coincide within tolerance and their expected response times coincide. Everything else about an item — its wording, its names, its numbers — is free to move.

Base item, calibrated (dotted) Generated variants, n = 8 (solid) ±0.05 logit parity envelope
FIG 1  Item characteristic curves. The eight variants are not approximately equal — they are inside a band the calibrator refuses to let them leave.

Equivalence of difficulty is not enough on its own. An item that is equally hard but less discriminating measures less, and a paper built from such variants would be quietly less reliable. So the same tolerance is applied to information:

Ik(θ) = 2.89 · ak2 · [(1 − Pk) / Pk] · [(Pk − ck) / (1 − ck)]2 Item information: how much a single item reduces uncertainty about θ. A variant must peak at the same place, and to the same height, as its parent.
Base item (dotted) Generated variants (solid)
FIG 2  Item information. Equal difficulty with unequal information would still be an unfair swap; holding Δa ≤ 0.01 keeps the peaks coincident.

Section 03Radicals and incidentals

The generation rule comes from item-generation theory, not from prompt engineering. Every feature of an item is classified as a radical — a feature that drives difficulty — or an incidental — a feature that does not. Generation may only move incidentals.

FeatureClassMay the generator move it?
Numeric values inside declared rangesIncidentalYes — provided the arithmetic stays in the same magnitude band
Person and place namesIncidentalYes — drawn from a curated, regionally balanced list
Surface phrasing and sentence orderIncidentalYes — within a readability band of ±0.4 grade levels
Option orderIncidentalYes — with key position balanced across the cohort
Cognitive operation requiredRadicalNo
Number of solution stepsRadicalNo
Number of optionsRadical (sets c)No
Distractor structure — which misconception each distractor encodesRadicalNo; distractors are regenerated from the same misconception map
Quantity of irrelevant informationRadicalNo
Unit systems and conversions requiredRadicalNo

An item template is therefore not a prompt. It is a typed object: declared parameters with ranges and units, a closed-form solution, a misconception map for distractors, and an explicit list of transforms that provably preserve the answer. The model's job is confined to producing fluent natural language over that structure.

The known failure mode. A change intended as incidental can turn out to be radical — a "cyclist" swapped for a "goods train" changes the plausible magnitude of the answer and therefore the guessability of the distractors. This is why every variant is machine-solved and calibrated rather than trusted, and why the readability and magnitude bands exist.

Section 04Cold-start: parameters for an item nobody has ever sat

The obvious objection: IRT parameters are normally estimated from response data, and a variant generated four seconds ago has none. Four mechanisms answer this, in the order they are applied.

  1. Family-level priors (hierarchical IRT). Variants of one template are not independent items; they are siblings drawn from an item family. The family has its own parameter distribution, estimated from every sibling ever administered. A new variant inherits the family posterior as its prior — so it starts with an informative estimate, not a guess. Formally, bk ~ N(μfamily, σ2family), and the family's σ is itself an object of study: a family with a wide σ is a badly specified template and is sent back to its author.
  2. Feature-based prediction (LLTM-style). Difficulty is regressed on the item's declared radicals — number of solution steps, operation type, unit conversions, distractor structure, readability. Because generation cannot move radicals, the predicted difficulty of a variant is by construction the predicted difficulty of its parent. This model is fitted on the whole calibrated bank, not on one family, and its residual standard error is reported with every release.
  3. Seeding on a non-scoring cohort. New families enter live papers as unscored seed items first — the standard practice for pre-testing — accumulating real response data before any candidate's result depends on them. A family is not released for scoring until its parameters are estimated from at least 1,200 responses per variant class.
  4. Online re-estimation and retrospective correction. During delivery, responses update the family posterior continuously. If a variant's realised difficulty drifts outside tolerance mid-administration, it is withdrawn from further delivery and every candidate who already received it is re-scored on the corrected parameters — upward or downward — before results are released.

The honest version. Mechanism 3 is the one that carries the scientific weight; 1 and 2 make cold-start estimates good enough to be safe, and 4 catches what the first three miss. A vendor claiming that a language model can assign IRT parameters directly, with no response data anywhere in the loop, is claiming something psychometrics does not support.

Section 05Where ±0.05 logits comes from

The tolerance is not a round number chosen for comfort. It is derived from the effect a difficulty shift can have on a candidate's expected score, and bounded so that effect stays far below the test's own measurement error.

The steepest point of the ICC bounds the damage a difficulty error can do:

maxθ |∂P/∂b| = 1.7·a·(1 − c) / 4 For the worked item (a = 1.240, c = 0.200): 1.7 × 1.240 × 0.800 ÷ 4 = 0.4216 per logit.
A difficulty error of 0.05 logits therefore moves the probability of a correct response by at most 0.021 — about two percentage points, and only for candidates sitting exactly at the item's difficulty.

Aggregated over a 60-item paper, three quantities matter, and the specification bounds all three:

QuantityValueAs a fraction of one SEM
Typical form effect — variants uniform on ±0.05, independent across itemsσ ≈ 0.094 marks3.0%
Systematic effect — bounded by the mean-deviation rule |b̄form − b̄reference| ≤ 0.008≤ 0.202 marks6.4%
Adversarial worst case — every item at the tolerance edge, same direction (prevented by the rule above)1.27 marks40%
Standard error of measurement, 60 items, reliability 0.90, score SD 103.16 marks100%

SEM = SD × √(1 − reliability). The comparison is the point: a candidate's score already carries about ±3.2 marks of irreducible measurement noise. The form they were given contributes under a tenth of one mark of that.

Two rules follow, and both are enforced by the calibrator rather than by policy:

  • Per-item: reject any variant with |Δb| > 0.05, |Δa| > 0.01, |Δc| > 0 or |Δt̄| > 4 s.
  • Per-form: reject any assembled form whose mean signed difficulty deviation exceeds 0.008 logits. This is what stops many small permitted errors from stacking in one direction.

Section 06Form-level equivalence

Item-level parity is necessary but not sufficient; candidates sit forms, not items. The test characteristic curve — expected raw score as a function of ability — is the form-level statement of fairness, and it is checked for every assembled form before release.

Form A (solid) Form B, independently generated variants (dashed)
FIG 3  Test characteristic curves for two complete 60-item forms assembled from different variants of the same templates. Largest expected-score difference anywhere on the ability range: marks.

Assembly constraints applied to every form, in addition to the parity rules:

  • Target information function. Every form must match the reference form's information curve within 3% at each of nine ability anchor points, so precision is equal where cut-scores fall.
  • Content blueprint. Topic, sub-topic and cognitive-level counts are fixed by the blueprint and cannot be traded against difficulty.
  • Answer-key balance. Key positions are balanced across the cohort so option order carries no signal.
  • Time model. The sum of expected solve times must fall within ±90 seconds of the reference form.

Section 07Linking and scale maintenance

Scores are reported on a stable scale across shifts, sessions and years. Three mechanisms hold that scale in place.

MechanismHow it worksFrequency
Common-item anchoringA fixed anchor set of unvaried, fully calibrated items appears in every form, in identical wording, and carries the scale between forms.Every form, ≥ 12 anchor items
Concurrent calibrationAll forms in an administration are calibrated together in one run, so parameters are on one metric by construction rather than by transformation.Per administration
Scale drift monitoringAnchor-item parameters are compared against their historical values; drift beyond a control limit halts reporting and triggers re-linking.Per shift

Where shifts must be compared — as in any multi-session national examination — equating is performed on the ability metric, not by percentile normalisation of raw scores. This matters: percentile normalisation across shifts of unequal ability composition is the source of much of the public distrust in existing multi-shift examinations, and IRT linking removes the need for it.

Section 08Exposure control

Generation does not by itself solve exposure. If one template's variants are drawn far more often than another's, that template becomes worth memorising. Exposure is therefore managed explicitly.

  • Per-template exposure ceiling. No template may exceed a target proportion of the cohort, enforced by the sampler at draw time.
  • No repeats within a cohort. Two candidates in the same administration never receive the identical parameter tuple from the same template.
  • Family retirement. Cumulative exposure across administrations is tracked; families are retired on a schedule and replaced from the authoring pipeline.
  • Late-arrival parity. A candidate starting twenty minutes late is drawn from the same live pool under the same ceilings — there is no leftover paper and no advantage to arriving late.

Section 09Differential item functioning

Parity in aggregate can still hide unfairness for a subgroup — a generated variant set in an unfamiliar context may be harder for rural candidates at the same ability, or a Hindi rendering may not be of equal difficulty to its English source. DIF analysis is run after every administration and its results are published, not filed.

Observed Δb between subgroups 95% confidence interval ±0.05 tolerance
FIG 4  Subgroup contrasts for the worked template. Illustrative values shown to demonstrate the analysis; real figures are published per administration.
  • Method. Mantel–Haenszel with ETS A/B/C classification for flagging, plus IRT-based Δb comparison for magnitude, matched on ability rather than raw score.
  • Contrasts. Language of delivery, urban/rural, gender, device tier, and state. Device tier is included because it is the contrast most specific to remote delivery and the one a centre-based exam never had to consider.
  • Consequence. A family showing C-level DIF is withdrawn and its authors notified. Candidates who received the affected variant are re-scored without it.
  • Translation. Hindi and English renderings are calibrated as separate items in the same family, never assumed equal because they are translations of one another.

Section 10Validation protocol and limitations

Before an examination

  1. Family seeding on a non-scoring cohort, minimum 1,200 responses per variant class.
  2. Solver verification of every template transform, including adversarial parameter values at range boundaries.
  3. Cognitive-lab review of a random variant sample by subject experts, blind to which is the parent item.
  4. Simulated cohort run: 100,000 synthetic candidates across the ability range, checking recovery of true θ under randomised form assignment.

After an examination

  1. Concurrent calibration of all delivered variants; comparison of realised against predicted parameters.
  2. Regression of score on variant identity, conditioning on θ. The coefficient must be indistinguishable from zero. This is the falsification test in §1.
  3. Full DIF sweep across the five contrasts, published with the results.
  4. Publication of every delivered variant and its parameters, as answer keys are published today.

Limitations, stated plainly

LimitationWhy it existsWhat we do about it
Cold-start estimates are estimatesA never-sat variant has no response data; family priors and feature models are informed guesses with quantified error.Seeding before scoring; online re-estimation; retrospective re-scoring when a variant drifts.
The 3PL assumes unidimensionalityMost competitive-exam subtests approximate it, but not perfectly.Dimensionality checks per subtest; multidimensional models where the assumption fails, rather than ignoring it.
Response-time models are weaker than difficulty modelsTime-to-solve varies more with candidate strategy than difficulty does.A wider tolerance (±4 s) honestly reflecting lower precision, and form-level time bounds.
Constructed-response items are out of scopeParity for free-text answers depends on marking equivalence, a different and harder problem.v1.0 covers selected-response and numeric-entry only. We do not claim what we cannot show.
Templates can be badly writtenA template whose "incidentals" are secretly radical will produce a wide family σ.Family variance is monitored and reported back to authors; high-σ families are rejected, not tuned around.

Position. Dynamic generation raises the fairness bar, it does not lower it. A single printed paper is only "fair" because everyone got the same one — a much weaker guarantee than the one made here, and one that says nothing at all about whether the paper measured what it claimed to.