Deliverable 01 Version 1.0 · Draft for evaluation LAST REVISED 30 JULY 2026

Product Requirements Document

What TrustXD ExamOS must do, for whom, and what it must refuse to do. Written to be read by a procurement committee, an examinations controller and an engineer in the same sitting.

ScopeLevels 1–3 remote delivery, national and university scale
PersonasCandidate · Reviewer · Controller · Agency · Auditor
Requirements62 functional · 21 non-functional
Edge cases26 defined protocols, all candidate-favourable

Section 01Purpose and scope

TrustXD ExamOS is assessment infrastructure for examinations whose outcome materially changes a person's life. Its purpose is to make three specific attacks structurally impossible rather than statistically unlikely: advance paper leakage, proxy test-taking, and unaccountable automated judgement.

In scope: candidate onboarding and identity assurance, dual-device capture, screen binding, just-in-time paper generation with psychometric parity, delivery, evidence capture, human review, and candidate-facing transparency — for Levels 1 through 3 as defined in §6.

Out of scope for v1.0: subjective long-form marking, practical or laboratory examinations, viva voce, and any exam requiring physical apparatus. Also explicitly out of scope, permanently: emotion recognition, gaze-based "attention scoring", keystroke biometrics, and any probabilistic integrity score.

The negative requirement is a requirement. If a future release can produce a number that purports to represent a candidate's honesty, the product has failed its own specification. This is enforced by schema: no such field exists in the evidence model, and the review API rejects any payload containing one.

Section 02Product principles

  1. Structure beats surveillance. Prefer a design where cheating is impossible over a design that watches harder for it. The paper that does not exist cannot leak.
  2. Evidence, never inference. The system records what happened with timestamps. It does not conclude what a person intended.
  3. Ambiguity favours the candidate. Every undecidable state resolves to Not assessable with a free re-sit, never to a penalty.
  4. The candidate sees what we see. Self-view, evidence trail, reviewer access log and retention clock are all candidate-visible by default.
  5. Assurance is declared, not assumed. An institution picks a level; the level appears on the consent screen and on the result.
  6. Access is a design constraint. A requirement that only works for candidates with two new phones and fibre broadband is a defect, not a feature.

Section 03Personas

PersonaContextPrimary needFailure they fear
Aditi — candidate
18, NEET aspirant, Tier-3 town
Shared 4G, one laptop borrowed from a cousin, a mid-range Android phone. To sit the exam without an eleven-hour journey, and to know exactly what is recorded. Being failed by software for a dropped connection she did not cause.
Rakesh — reviewer
Contracted, independent pool
Reviews flagged sessions across many institutions; never sees candidate names first. Complete, ordered evidence and a decision that is defensible in writing. Being asked to rubber-stamp a machine's opinion.
Dr Menon — controller of examinations
University board
Answerable to a vice-chancellor, a senate and the press on exam day. Predictable delivery, a live operational picture, and an audit file. A viral video of a candidate wrongly ejected.
Agency programme office
National testing body
Millions of candidates, multiple shifts, statutory obligations. Scale, evidence, cost per candidate, and legal defensibility. A leak, and the cancellation that follows it.
Statutory auditor
CAG / data-protection board
Post-hoc inspection, sometimes years later. Reproducible records, purge evidence, and consent receipts. A system that cannot show its working.

Section 04User stories

Twelve representative stories; the full backlog carries 84. Acceptance criteria are written to be testable without interpretation.

Candidate

US-C-01MUST · L1–L3

As a candidate, I want to check my device and network days before the exam, so that I discover problems while I can still solve them.

Acceptance criteria
  • Readiness check runs from the same client build as exam day and reports upload bandwidth, camera resolution, battery health and rig geometry.
  • A failed check returns a specific remedy, never a generic error.
  • Results are stored against the registration and re-run automatically at T−45.
US-C-02MUST · L1–L3

As a candidate, I want to grant each permission separately and see a receipt, so that consent is a record and not a checkbox.

Acceptance criteria
  • Camera, microphone, ID scan and face matching are four independent grants.
  • Consent UI renders in English and Hindi at parity; language choice persists.
  • A signed receipt (Ed25519) is issued to the candidate before first capture and is downloadable at any later date.
US-C-03MUST · L2–L3

As a candidate, I want to see my own phone camera stream at full size throughout, so that I always know what is in frame.

Acceptance criteria
  • Self-view is on by default, cannot be reduced below 240 px on the shortest edge, and can be repositioned but not disabled.
  • Any capture state change (paused, resumed, degraded) is announced visually and to assistive technology.
US-C-04MUST · L1–L3

As a candidate whose session was flagged, I want to read the same evidence the reviewer read, so that I can respond to facts rather than to a score.

Acceptance criteria
  • The candidate report contains the full ordered event trail with timestamps.
  • Any clip shown to a reviewer is playable by the candidate.
  • The report lists every account that accessed the file, with role and timestamp.
  • A dispute can be opened from within the report; opening one freezes the retention clock.

Reviewer

US-R-01MUST · L2–L3

As a reviewer, I want flagged sessions presented as an ordered evidence trail, so that I decide from the record rather than from an impression.

Acceptance criteria
  • Events are shown in strict server-clock order with the triggering rule named.
  • Candidate identity is masked until a verdict is committed.
  • The UI offers exactly three verdicts and requires a written rationale for any verdict other than Clear.
  • No aggregate score, percentage or ranking is displayed anywhere in the reviewer UI.
US-R-02SHOULD · L3

As a reviewer, I want to escalate a session I cannot decide, so that difficult cases reach a second independent reviewer instead of a coin flip.

Acceptance criteria
  • Escalation is a first-class action, not an absence of decision.
  • The second reviewer cannot see the first reviewer's rationale until they commit their own.
  • Disagreement routes to a three-person panel; the panel's reasoning joins the candidate's record.

Controller and agency

US-A-01MUST · L1–L3

As a controller, I want a live operational picture of delivery health, so that I can answer questions during the exam rather than after it.

Acceptance criteria
  • Concurrency, gate latency, queue depth and disposition update at ≤ 10 s freshness.
  • Regional degradation is surfaced as an aggregate, never as a list of named candidates.
  • Every panel is exportable as a timestamped CSV for the audit file.
US-A-02MUST · L2–L3

As an agency, I want to declare an incident that affects a region and reschedule affected candidates in bulk, so that a power grid failure is not a candidate's problem.

Acceptance criteria
  • Incident declaration takes a geography, a time window and a reason code.
  • Affected sessions are marked Not assessable — administrative, with a re-sit slot issued automatically.
  • The incident and its blast radius are written to the audit ledger.

Item author and psychometrician

US-P-01MUST · L2–L3

As an item author, I want to declare which parts of an item may vary and within what ranges, so that generation cannot change what the item measures.

Acceptance criteria
  • Templates declare typed parameters with ranges, units and answer-preserving transforms.
  • Any generated variant is machine-solved and compared against the template's closed-form answer before release.
  • A variant that changes the number of options, the cognitive operation or the required formula is rejected at build time.
US-P-02MUST · L2–L3

As a psychometrician, I want every delivered variant's IRT parameters recorded, so that I can prove cohort fairness after the fact.

Acceptance criteria
  • a, b, c and expected solve time are persisted per delivered variant per candidate.
  • Any variant outside tolerance is rejected pre-delivery and the rejection is logged with its parameters.
  • Post-administration DIF analysis across language, region and gender is produced within 72 hours.

Auditor

US-X-01MUST · L1–L3

As an auditor, I want to replay a session's evidence years later and get an identical trail, so that the record is evidence rather than a rendering.

Acceptance criteria
  • Event records are append-only, content-addressed and signed at write time.
  • Replaying the same inputs through the same engine version produces byte-identical output.
  • Engine version, rule-set version and model version are recorded on every session.
US-X-02MUST · L1–L3

As an auditor, I want proof that data was deleted, so that a retention policy is a fact rather than a promise.

Acceptance criteria
  • Each purge emits a signed record naming the object, its hash and the policy that triggered deletion.
  • Purge records survive the data they describe and are themselves retained for seven years.

Section 05Functional requirements

Abbreviated to the load-bearing set. M = must, S = should, C = could.

Module A — identity and spatial verification

IDRequirementLevelPri
FR-A-01Parse Aadhaar, passport or driving licence entirely on device; transmit field hashes and a match decision only.L1–L3M
FR-A-02Run active liveness with randomised prompts drawn from a set of at least eight movement classes, including an ear-inspection rotation.L1–L3M
FR-A-03Construct a 468-point facial mesh and verify depth symmetry; reject flat-surface presentations.L2–L3M
FR-A-04Require the secondary device to sit between 60° and 120° from the primary optical axis, verified geometrically before session start.L2–L3M
FR-A-05Correlate head-motion vectors from both cameras over a rolling 8-second window; hold correlation ≥ 0.94 to remain in Level 3.L3M
FR-A-06Require hardware attestation (App Attest / Play Integrity) on every media payload from the secondary device; verify the receipt chain server-side.L3M
FR-A-07Refuse sessions from emulators, rooted or jailbroken devices, and virtual camera drivers, before any exam material is addressed.L3M
FR-A-08Support religious head coverings, prosthetics and facial differences without additional friction; verification adapts, the candidate does not.L1–L3M

Module B — screen binding

IDRequirementLevelPri
FR-B-01Derive a per-session secret server-side and render a visual token with a 3-second lifetime on the primary display.L2–L3M
FR-B-02Read the token on the secondary device and return it over a WebSocket independent of the primary client's transport.L2–L3M
FR-B-03Validate arrival against the server's NTP-disciplined clock within a 2000 ms window; never trust a client-supplied timestamp.L2–L3M
FR-B-04Require three consecutive successful cycles before the delivery gate will evaluate.L3M
FR-B-05Record every lapse as a discrete event with observed sequence number and measured delta; never aggregate lapses into a score.L2–L3M
FR-B-06Degrade the token to a high-contrast, large-format glyph set when the secondary camera reports low resolution or poor light.L2–L3S

Module C — JIT generation and delivery

IDRequirementLevelPri
FR-C-01Hold the item bank encrypted at rest; derive the session decryption key only after Modules A and B both hold, and never persist it client-side.L2–L3M
FR-C-02Generate a per-candidate variant set at T−0, including for late arrivals, from declared parameter ranges only.L2–L3M
FR-C-03Vary numeric parameters, named entities, phrasing and option order; never vary option count or the cognitive operation under test.L2–L3M
FR-C-04Machine-solve every variant against the template's closed form and reject any mismatch before delivery.L2–L3M
FR-C-05Submit every variant to the IRT calibrator; reject any variant outside Δb ±0.05 logits, Δa ±0.01 or Δt̄ ±4 s.L2–L3M
FR-C-06Stream items individually; never place the complete paper in client memory or storage at one time.L3M
FR-C-07Fall back to a pre-generated, pre-calibrated variant pool if the generation service degrades, without pausing delivery.L2–L3M
FR-C-08Apply the candidate's declared accessibility profile to the timing model at generation time, not as an afterthought.L1–L3M

Modules D and E — evidence, review and candidate rights

IDRequirementLevelPri
FR-D-01Emit deterministic, append-only, signed event records; identical inputs must produce byte-identical output.L1–L3M
FR-D-02Expose exactly three verdicts. Reject at the API layer any payload containing a probability, score or ranking.L1–L3M
FR-D-03Route every flagged session to a trained human reviewer; no verdict may be committed by an automated process.L1–L3M
FR-D-04Mask candidate identity in the reviewer UI until a verdict is committed.L2–L3M
FR-D-05Require a written rationale for any verdict other than Clear, stored with the record and visible to the candidate.L1–L3M
FR-E-01Capture itemised consent per permission, bilingually, with a signed receipt issued before first capture.L1–L3M
FR-E-02Display the candidate's secondary stream to the candidate at full size for the entire session.L2–L3M
FR-E-03Destroy unflagged media at session seal; retain flagged clips for 30 days encrypted, then purge with a signed record.L1–L3M
FR-E-04Show the candidate every access to their evidence, with role and timestamp.L1–L3M
FR-E-05Allow a dispute to be opened from the candidate report; freeze the retention clock while it is open.L1–L3M

Section 06Non-functional requirements

IDAttributeTarget
NFR-01Peak concurrency2,500,000 simultaneous Level 3 sessions per shift, across at least three availability zones in-country.
NFR-02Delivery gate latencyp50 ≤ 45 s, p95 ≤ 90 s from Level 3 clear to first item rendered.
NFR-03Token round tripp95 ≤ 900 ms on a 1.5 Mbps uplink; hard validity window 2000 ms.
NFR-04Minimum viable network1.5 Mbps sustained upstream; graceful degradation to audio + 5 fps below that, with no verdict consequence.
NFR-05Client footprintExam client runs on a 4 GB RAM laptop of 2018 vintage; phone app supports Android 10+ and iOS 15+.
NFR-06Availability99.99% during declared exam windows; no scheduled maintenance within 72 h of a national administration.
NFR-07RecoveryRPO 0 for responses and event records; RTO ≤ 60 s for session resumption.
NFR-08AccessibilityWCAG 2.2 AA across candidate, reviewer and controller surfaces; screen-reader parity for all exam navigation.
NFR-09LocalisationEnglish and Hindi at launch; architecture supports all scheduled languages without client changes.
NFR-10Data residencyAll candidate data processed and stored in-country on infrastructure the agency may inspect.
NFR-11CryptographyTLS 1.3 in transit; AES-256-GCM at rest with per-session keys; Ed25519 for receipts and event signing.
NFR-12DeterminismEvidence engine output byte-identical across replays for seven years; engine and rule-set versions pinned per session.
NFR-13Review capacityMedian time to first reviewer ≤ 15 min during a live administration; reviewer pool sized to a 4% flag rate.
NFR-14CostTarget delivered cost per candidate below the incumbent physical-centre cost at equal or better assurance.

NFR-04 is the equity requirement. Any change that raises the minimum viable network, device generation or bandwidth is treated as a scope change requiring programme-office approval, because it silently removes candidates from the exam.

Section 07Edge-case protocols

Twenty-six defined states. The governing rule: an environmental failure is never a candidate's fault unless a human reviewer, looking at the evidence, concludes that it was.

IDTriggerDetectionProtocolCandidate impact
E01Primary network dropsHeartbeat miss > 400 msTimer halts, answers sealed locally, client retries 15 min; on resume, elapsed time credited in fullNone — event on record only
E02Incoming phone callOS audio-session interruptCamera session held, timer paused; answered call logged as stream interruption and routed to reviewReview, not penalty
E03Secondary battery exhaustedBattery telemetry < 5% or stream end10-minute restoration window; beyond it session becomes Not assessableFree re-sit
E04Phone mount slipsGeometry outside 60°–120° for 2 framesReposition prompt; window excluded from exam clockNone
E05Mains power failureClient on battery + region correlationDegrade to audio + 5 fps; region-wide pattern escalates to agency incidentReschedule if regional
E06Second person enters frameAdditional face detectedEnvironmental event with clip attached; a human distinguishes passing from seatedReview only
E07Comfort or medical breakDeclared in advanceBreak granted, clock excluded; re-entry requires fresh livenessNone
E08Assistive technology in useDeclared at registrationStatus agent allowlists AT; timing model adjusted at generationNone
E09Only one device availableReadiness checkCandidate assigned a supervised civic node with loan hardwareNone
E10Browser or OS crashSession heartbeat lostResume token valid 15 min; answers restored from sealed local storeNone
E11Secondary app backgroundedApp lifecycle eventTwo warnings, then pause; repeated backgrounding is an event for reviewReview only
E12Bandwidth below floorSustained < 1.5 Mbps upAutomatic degradation ladder; token glyphs enlarge; no verdict consequenceNone
E13Loud ambient environmentAudio SNRNoted as environmental context; never a flag on its ownNone
E14Two candidates in one householdShared NAT + concurrent sessionsPermitted and expected; each session evaluated independentlyNone
E15Spectacle glareMesh confidence dropReposition-light prompt; profile camera carries verification meanwhileNone
E16Religious head coveringDeclared or detectedVerification uses the visible facial region; no request to remove anythingNone
E17Facial difference or prostheticDeclared at registrationEnrolment-based matching against the candidate's own baseline, not a population modelNone
E18Severe backlightFace SNR below floorGuided lighting fix; if unresolved after two attempts, Not assessableFree re-sit
E19Client clock skewServer-side comparisonIgnored — server clock is authoritative for every validity decisionNone
E20Attestation provider outageVendor API failureCached receipts honoured for 24 h; new sessions drop to Level 2 with agency notificationLevel declared on result
E21Generation service degradedLatency or error budgetFall back to pre-calibrated variant pool; delivery never blocks on the modelNone
E22Variant fails calibrationIRT tolerance checkVariant discarded, next candidate-specific variant drawn; rejection logged with parametersNone
E23Candidate disputes a verdictDispute opened in reportRetention clock frozen; independent panel review within 10 working daysRight of reply
E24Reviewers disagreeVerdict mismatchThree-person panel; reasoning joins the candidate recordTransparent outcome
E25Regional internet shutdownCorrelated loss by geographyAgency incident declared; affected sessions Not assessable — administrative, re-sit issuedAutomatic re-sit
E26Candidate is a minorDate of birth at registrationGuardian co-consent required under DPDP; retention reduced to 15 daysAdditional protection

Section 08Verdicts and appeals

The verdict set is closed. Adding a fourth outcome is a breaking change requiring board approval and public notice.

VerdictMeaningWho may assign itConsequence
Clear No flags raised, or flags raised and resolved as benign by a reviewer. System (no flags) or reviewer Result released normally. Media purged.
Review required One or more events need human interpretation. This is a routing state, not a finding. System Result held ≤ 10 working days. Candidate notified with the trail.
Not assessable Evidence quality is insufficient to say anything at all. Says nothing about the candidate. Reviewer or agency (administrative) Free re-sit at the next available slot. No mark, no penalty, no record of suspicion.

What a reviewer may conclude

A reviewer resolves Review required into Clear or Not assessable. A finding of misconduct is not a verdict this system issues — it is a decision for the institution's own disciplinary process, to which TrustXD supplies the evidence file and nothing else. The product deliberately has no opinion on guilt.

Section 09Data and retention

ClassExamplesRetentionCandidate visibility
Identity artefactsID field hashes, enrolment face templateDuration of the exam cycle + 90 daysVisible, exportable
Unflagged mediaPrimary and secondary video, ambient audioDestroyed at session sealLive self-view only
Flagged clips±30 s around a flagged event30 days (15 for minors), frozen during appealPlayable in candidate report
Event recordsTimestamped protocol events, verdicts7 years, append-onlyFull trail visible
Consent receiptsSigned grants and withdrawals7 yearsDownloadable any time
Purge recordsObject hash, policy, timestamp7 yearsNotified on purge
Delivered variantsItem text, IRT parameters per candidate7 yearsReleased with the answer key
// Evidence record — the only shape the review API accepts. // Note what is absent: no score, no probability, no ranking, no confidence. { "session": "S-77401", "engine": { "evidence": "3.2.1", "rules": "nta-l3-2026.4" }, "event": { "t": "2026-07-30T11:42:06.114+05:30", // server clock, authoritative "type": "secondary_stream.interrupt", "rule": "B-04.heartbeat_timeout", "observed": { "rtt_ms": 4180, "gap_s": 11.4 }, "clip": "sha256:9c1f…a07f", "routes_to": "human_review" }, "verdict": "review_required" // one of exactly three }

Section 10Rollout and success measures

PhaseScopeExit criteria
P0 · ShadowOne university entrance test, ~5,000 candidates, run alongside the existing centre-based exam.Zero delivery incidents; verdict distribution reviewed by the institution's own psychometrician; candidate satisfaction ≥ 4.2/5.
P1 · ParallelTwo university tests + one banking recruitment shift, ~120,000 candidates, results issued from TrustXD.Gate latency p95 ≤ 90 s; flag rate ≤ 4%; appeal upholding rate published.
P2 · RegionalOne state common entrance test, ~600,000 candidates, with civic nodes in aspirational districts.Civic-node coverage ≥ 98% of candidates lacking a second device; no correlation between district and Not-assessable rate.
P3 · NationalA national competitive examination at full scale.Statutory audit clean; independent DIF review published; incumbent cost per candidate matched or beaten.

Success measures

  • Zero papers in existence before T−0. Binary, auditable, non-negotiable.
  • Flag rate under 4% with an appeal-upholding rate published every cycle — a rising upholding rate means the rules are wrong, not the candidates.
  • No demographic skew in Not-assessable rates across district, language, gender or device tier. This is measured and published, not asserted.
  • Median time to reviewer under 15 minutes during live administration.
  • Candidate-reported clarity ≥ 4.2/5 on "I understood what was being recorded and why".

Section 11Risks and open questions

RiskAssessmentMitigation
Household coercion — someone in the room pressuring the candidateThe hardest unsolved problem in at-home assessment. Two cameras see the room; they cannot see intent.Ambient audio + environmental events; a discreet in-exam duress signal; civic nodes as an opt-in alternative for any candidate who wants one, no reason required.
Generative model produces a subtly harder variantMedium likelihood, high impact on fairness.Machine-solve gate, IRT tolerance gate, response-time model, and post-hoc DIF. Rejection is cheap; delivery of a bad variant is not.
Attestation monoculture — dependence on two vendorsStructural. Apple and Google can change policy unilaterally.Cached receipts, documented Level 2 fallback, and an active track on open attestation standards.
Device inequality reframed as meritHigh. The most likely way this product does harm.NFR-04 as a hard floor, civic nodes funded as programme cost, and published Not-assessable rates by district.
Reviewer capacity at national peakMedium. A 4% flag rate at 2.5 M candidates is 100,000 reviews.Flag-rate budgeting per rule, tiered triage, and a contracted pool sized against the worst observed rate, not the average.
Scope creep toward behavioural analyticsMedium — it will be requested by customers.Schema-level refusal, published product principles, and a contractual commitment not to ship integrity scoring.

Open questions for the evaluation committee

  1. Should Not assessable re-sits be scheduled within the same cycle, or does that create a second-mover advantage that the variant engine must then neutralise?
  2. Who contracts the reviewer pool — the agency, an independent trust, or a statutory body? Independence is a design requirement, and procurement structure decides whether it is real.
  3. What is the statutory basis for civic nodes, and which budget line funds them?
  4. Should delivered variants and their parameters be published after every administration, as answer keys are today? We propose yes.