JSON Legibility Is Not Control: When Schema Names Move a Routing Boundary
A preregistered experiment with Jev tests an engineering assumption: if the choice descriptions stay identical, should renaming JSON keys change what the application does?
I thought I was looking at a harmless refactor: change the JSON choice keys sent to a model, keep their complete descriptions identical, and leave the application’s routing logic alone.
I tested that expectation with Jev, a model designed for fast, structured decisions inside software. Give it application state and a set of typed choices, and it returns a decision with probabilities that code can consume directly. Both sets of keys were reasonable labels for the same outcomes, with identical explicit definitions. I expected those definitions to carry enough meaning that the rename would make little practical difference.
Instead, the human review rate went from 45.8% to 70.8% on cases deliberately selected near the routing threshold. Every response was valid. The descriptions and routing rule had not changed.
As an engineer, I was used to treating those keys as identifiers. JSON is such a familiar software contract that I had carried the same mental model into a model-facing interface. The experiment made me reconsider that assumption: when a model reads a JSON key, it also reads the words in it.
Why the rename looked safe
TypeSafe presents Jev as a “frontier-intelligence function call”: application state in, typed probabilistic decisions out. It is designed to slot into ordinary software as a “smart if-statement” that can classify, score, or route, with probability thresholds determining when code acts and when a human reviews. TypeSafe also emphasizes similar answers for similar inputs. That is precisely the engineering expectation this experiment probes.
To a developer, this resembles a typed service. The request has a schema. The response has a schema. Invalid structures can be rejected. If every choice retains the same complete description, changing an identifier can look like an ordinary refactor.
Each choice in my experiment included an explicit definition. The partial capacity outcome, for example, applied when “at least one essential duty cannot be completed at normal quality or pace.” I expected definitions like that to carry most of the meaning. The alternative keys were neither opaque nor misleading, so I expected little or no operational difference between them.
But a model-facing identifier is not only an enum value. It is also language, and that matters when the returned score sits near a routing threshold.
The two schemas
I compared these key families:
{
- "no_essential_duties_prevented": "The person can complete every essential job duty at normal quality and pace without assistance.",
- "some_essential_duties_prevented": "The person can complete some essential job duties, but at least one essential duty cannot be completed at normal quality or pace.",
- "all_essential_duties_prevented": "The person cannot complete any essential job duty."
+ "essential_work_unimpeded": "The person can complete every essential job duty at normal quality and pace without assistance.",
+ "essential_work_impeded": "The person can complete some essential job duties, but at least one essential duty cannot be completed at normal quality or pace.",
+ "essential_work_halted": "The person cannot complete any essential job duty."
}
Both followed Jev’s documented Choice structure. Its OpenAPI schema describes the criteria as choice names paired with descriptions of when each applies. It also notes that a choice without a description is interpreted by its name alone. Here, every choice had a complete description.
The two families are not strict synonyms, and that is the point. No / some / all prevented forms an explicit scale. Unimpeded / impeded / halted forms an asymmetric scale, and impeded can reasonably suggest a lower bar than prevented. To a model, that difference is meaningful. To a developer, it can look like a naming choice: exactly the kind of rename that might be suggested in code review for readability and accepted because each outcome retains the same explicit description.
That is the assumption I tested: not whether words carry meaning, but whether keeping each choice’s complete description identical makes two reasonable schemas interchangeable, as renaming an enum member would be in ordinary code.
A prospective test
Before the first call in the preregistered run, I froze and committed the preregistration, stimuli, analysis, and randomization plan.
The experiment used six state descriptions. Two were primary probes chosen during test-bed development because their scores sat near the fixed 0.50 routing boundary. Two range controls covered less ambiguous cases on either side of that boundary. Two endpoint guardrails represented unequivocal full and partial capacity.
| State | Experimental role | Expected behavior |
|---|---|---|
| Possible occasional delay | Primary threshold probe | Boundary-sensitive |
| Explicit uncertainty | Primary threshold probe | Boundary-sensitive |
| Close-to-normal pace | Below-boundary range control | Automatic |
| Definite slight delay | Above-boundary range control | Human review |
| Full capacity | Endpoint guardrail | Automatic |
| Partial capacity | Endpoint guardrail | Human review |
These labels summarize the frozen state texts. The preregistration and public evidence preserve the exact identifiers and wording.
I ran 60 randomized blocks, each containing all six states under both key families. This produced 12 calls per block and 720 valid calls in total, with zero errors and zero retries. The two primary states produced two matched comparisons per block, for 120 matched pairs in total. All six semantic choice orders were balanced across blocks, and every response reported jev-1.13.0. Jev’s documented interface did not expose a temperature parameter, so I did not send one or assume an undocumented default.
The pseudocode below routes a response to human review when the combined review score reaches 0.50:
const reviewScore =
answer.probabilities[partialChoice] +
answer.probabilities[noWorkChoice];
const route = reviewScore >= 0.50
? "human_review"
: "automatic_full_capacity";
The technical appendix and public repository preserve the exact policy, randomization and bootstrap seeds, randomized schedule, and source hashes.
What changed
Across those 120 matched pairs, the mean review score moved from 47.16% under prevention keys to 54.04% under impediment keys.
The estimated shift was 6.88 percentage points, with a preregistered 95% complete-block bootstrap interval from 5.63 to 8.13 points, entirely above the preregistered 5-point effect-size threshold.
At the fixed 0.50 routing threshold, the score movement produced a larger change in action:
| Model-facing key family | Human review |
|---|---|
| Prevention | 55/120, or 45.8% |
| Impediment | 85/120, or 70.8% |
Within those same 120 matched pairs, 33 moved from automatic routing under prevention keys to human review under impediment keys. Three moved in the other direction. The net increase was 30 reviews out of 120, or 25.0 percentage points, with a 95% complete-block interval from 15.0 to 35.8 points.
These were separate model calls, not known counterfactual outcomes for 30 real cases. The review rates do not estimate production impact because these two states were deliberately selected near the experimental boundary. The test covers one task, one returned Jev version, and six selected state texts.
Within that boundary, the result is clear: a reasonable model-facing schema rename remained behaviorally material even though the complete descriptions, semantic outcomes, and deterministic routing policy were fixed.
The threshold exposed the shift; it did not create it
The routing difference is visually dramatic, so I checked whether 0.50 had manufactured the result by turning small score noise into binary decisions.
All four non-saturated states moved in the same direction. The two primary probes shifted by 5.52 and 8.25 points. The two range controls shifted by 12.33 and 6.17 points. Neither range control changed route because both remained far enough from 0.50.
The endpoint guardrails behaved as expected. Full capacity stayed at 0%, partial capacity stayed at 100%, and all 240 guardrail calls routed correctly across both key families.
Across those four non-saturated states, the score distribution moved in the same direction. The routing policy exposed the portion that crossed its boundary. This does not imply that every rename changes behavior, or that either key family is intrinsically better.
What typed output guarantees
Typed output can constrain a response to declared keys, reject malformed structures, and simplify parsing. It cannot guarantee semantic invariance among valid responses.
The experiment involved two different boundaries. The syntax boundary determined whether the response matched the declared shape. The routing boundary converted a model score into an application action. The syntax boundary held, but the value entering the routing boundary moved when the model-visible identifiers changed.
In conventional software, JSON is a familiar structural contract. At the model boundary, that mental model does not transfer unchanged.
JSON still makes inputs legible, auditable, transformable, and easy to validate. But its keys are also language, and in this experiment, changing that language changed scores and routing. A model-facing schema is therefore part of the prompt as well as the application contract.
The experiment is behavioral. It does not identify an internal mechanism, show that the keys dominated the descriptions, or establish which wording was more correct. It shows that the descriptions did not neutralize the framing carried by the keys.
The technical appendix contains the complete estimand, transition table, guardrails, randomization, retry rules, and audit trail. The public experiment repository contains the preregistration, frozen stimuli, executed runner, planned job order, raw responses, and reproducible analysis.
I would now keep stable internal enums for application identity, treat model-facing labels as versioned prompt language, and regression-test the routing boundary whenever that language changes.