HomeAboutContents

Technical Appendix: JSON Legibility and the Routing Boundary

Technical appendixLLM Reliability2026-10-02Eric Fruhinsholz

Preregistration, complete-block estimands, transition counts, controls, and reproducibility details for the 720-call prospective study.

This appendix documents the prospective experiment behind JSON Legibility Is Not Control: When Schema Names Move a Routing Boundary.

Prospective Status

The preregistration, frozen stimuli, runner, analysis, randomization seed, and semantic orders were committed before the first model call. The source commit was created at 2026-09-29T18:45:02-07:00; run.json was created 17 seconds later at 2026-09-30T01:45:19.321Z. Before sending the first call, the runner wrote the 720 planned job positions and source hashes to run.json.

The preserved run records runner SHA-256 a40763322832385ecd5d43b84296ab4753648f234ebfd95d98a9287ff0a11be2. The published runner is the exact executed source and has that hash.

All estimates, intervals, and success decisions reported here come from one prospective 720-call run.

Fixed Interface

Jev 1.13.0 received a state, one question, and three choices. Each choice had a model-facing JSON key and a complete description. Every returned key was mapped onto one stable semantic outcome: full work, partial work, or no work.

The review score and routing rule were fixed as:

review_score = score(partial) + score(none)

human_review if review_score >= 0.50
automatic otherwise

No temperature parameter was sent.

Design

  • 2 key families: duty_prevention, work_impediment
  • 2 primary states: r05, r06
  • 2 range controls: r04, r08
  • 2 clear-state guardrails: clear_full, clear_partial
  • 60 complete blocks
  • 12 calls per block
  • 720 calls total
  • all 6 semantic choice orders used exactly 10 times
  • both families assigned the same semantic order inside each state and block
  • randomized call order inside each block
  • randomization seed: 29092001
  • up to 3 retries, limited to HTTP 429, 503, and 529

The run completed with 720 valid calls, zero errors, and zero retries. Every response reported jev-1.13.0.

Earlier Exploratory Control With Opaque Keys

An earlier 600-call exploratory follow-up included opaque A/B/C identifiers alongside four descriptive key families. Across the four boundary states, the mean review score was 48.85% with opaque keys, 45.86% with the prevention family, and 57.24% with the impediment family. All families preserved the expected routing on the two clear-state guardrails.

This follow-up was separate from the preregistered 720-call prospective test. No equivalence margin was preregistered, so these results do not establish that opaque keys are equivalent to either descriptive family. They provide secondary evidence that removing semantic labels does not produce a simple, privileged baseline. The public exploratory artifacts contain the plan, raw responses, analysis, and audit trail for that earlier run.

Primary Estimands

Within each block, the analysis calculated work_impediment - duty_prevention for r05 and r06, then averaged the two differences equally. The primary estimate is the mean of those 60 block effects.

The routing estimand used the same construction after coding human review as 1 and automatic routing as 0.

Both 95% intervals came from 10,000 percentile-bootstrap resamples of the 60 complete blocks with seed 29092002 for scores and 29092003 for routing.

Primary estimand Estimate Complete-block bootstrap 95% interval
Review-score difference +6.88 pp [+5.63, +8.13] pp
Human-review-rate difference +25.00 pp [+15.00, +35.83] pp

State Results

State Prevention score Impediment score Difference Review-rate difference
r05_maybe_one_slight_delay 46.63% 52.15% +5.52 pp +15.00 pp
r06_unsure_one_slight_delay 47.68% 55.93% +8.25 pp +35.00 pp
r04_close_to_normal_pace 13.63% 25.97% +12.33 pp 0.00 pp
r08_one_duty_slightly_slower 88.48% 94.65% +6.17 pp 0.00 pp
clear_full 0.00% 0.00% 0.00 pp 0.00 pp
clear_partial 100.00% 100.00% 0.00 pp 0.00 pp

Primary Transition Table

The 120 matched units below pair separate calls by state and randomized block. They do not represent shared latent model events or literal counterfactuals.

Prevention route Impediment route Pairs
Automatic Human review 33
Human review Automatic 3
Automatic Automatic 32
Human review Human review 52

The precise statement is:

Thirty-three of 120 block-matched primary call pairs crossed the threshold in the prevention-to-impediment direction, while three crossed in the reverse direction.

Guardrails

State Family Correct routing
clear_full duty_prevention 60/60
clear_full work_impediment 60/60
clear_partial duty_prevention 60/60
clear_partial work_impediment 60/60

Frozen Decision

All six preregistered conditions passed:

  1. the primary score shift exceeded +5 points;
  2. its complete-block interval was above zero;
  3. both primary states shifted positively;
  4. the review-rate difference and its interval were above zero;
  5. every guardrail cell exceeded 95% correct routing;
  6. every response reported one Jev version.

The frozen decision is therefore supported for this test.

Audit Surface

The experiment directory contains:

  • PREREGISTRATION.md: frozen claim, design, estimands, criteria, and limits;
  • stimuli.v1.json: exact question, states, keys, descriptions, semantic orders, seeds, and policy;
  • run.json: the complete planned schedule, run-level timestamps, frozen stimuli, arguments, and source hashes;
  • calls.jsonl: per-call payloads, timestamps, attempts, and raw responses;
  • analysis.json and analysis.md: machine-readable and readable results;
  • the runner and analysis scripts needed to reproduce the schedule and calculations.

The complete audit surface is in the public experiment repository at commit 6a79286.

The result remains limited to this task, state set, threshold, interface, and returned Jev version. The scores are not interpreted as calibrated probabilities, and the experiment makes no claim about Jev’s internal mechanism.