surprise 0.1

world model · surprise driven · every belief written down

Predict. Be Surprised.
Rewrite The Model.

An agent that learns any world by predicting what will happen, noticing when it is wrong, and patching its own explicit world model. Every belief is a readable, versioned diff.

Open the live playroom See the measured results

surprise.log· · ·

        
one real surprise from a recorded run

the loop

Perceive, Predict, Act, Compare

1 · perceive
The environment is serialized into relations such as on(red,blue) or urgent(caller). Only relations whose concept is active are visible.
2 · predict
For each yes/no question about the next action, the predictor gives a calibrated probability and a confidence.
3 · act
The action runs in the world: a block is dropped, a caller is transferred, a ticket is reassigned.
4 · compare
Truth is evaluated after the world settles. Each prediction ends as one of three outcomes.
matched → assimilate
The rule that predicted gets one more piece of evidence. Confidence grows with counts.
surprised → accommodate
A confident guess failed. Claude Fable 5.1 reads the full trace, including what the agent could not see, and returns a patch: an explanation, concepts to activate, rules to add or remove.
unsure → escalate
Low confidence means a human is asked. The answer becomes an example the model can predict from next time.

two models share the work

A fast, cheap model makes every moment-to-moment prediction. A slow, careful model steps in only when a confident guess fails.

jev · TypeSafe AIsystem one

Calibrated probabilities for every question, every tick. About $0.042 per million input tokens. Given the current rules as context it predicts from them; without them it predicts zero-shot from the state alone.

predict(state, action, questions, rules) → { p(yes), confidence }

claude fable 5.1system two

Called only on surprise. It explains why the prediction failed and writes the patch as structured output. A hand-written rule predictor and a heuristic accommodator run the whole loop with no keys, and remain the fallback when a call fails.

accommodate(surprise, worldModel) → { explanation, conceptsAdd, rulesAdd, rulesRemove }

environments · one loop, one world model, seven worlds

The Loop Is Not A Physics Trick

  • physics3D playroomNaive physics curriculum in the order children acquire it: objects fall, object permanence, containers, support, ramps, tools.
  • peoplePhone callA synthetic caller with an intent, a mood and hidden dynamics. Urgent results must be transferred; asking twice annoys people.
  • softwareWeb shop checkoutAn undocumented flow with session and validation state the agent must discover.
  • organisationTeam ticketsWill the ticket close by Friday? Capacity and dependencies are latent until learned.
  • peopleOne person's habitsWhich slot will they take? Weather, energy and history explain the pattern.
  • gameRock, paper, scissorsAn opponent with a habit. The agent has to acquire an opponent model to beat chance.
  • assistiveVoice navigation aidA spoken guide for blind users, built on the same predict-and-correct loop.

A Unity build of the playroom is at ./unity.html. It may show build instructions instead of a scene.

four numbers, all measured

prediction accuracy

Did the guess match what happened? Reported per curriculum stage and per third of the run.

calibration

Does 80% mean 80%? Predicted probability against observed hit rate, in five buckets.

accommodation value

Did each rewrite help? Accuracy on the next predictions before and after every patch.

escalation rate

How often the agent asked a human. It should fall as the model fills in.

results · measured on this workload

Seven Worlds, Three Predictors

Each scenario ran once per predictor: hand-written rules, Jev zero-shot from the state alone, and Jev given the learned rules. Brier score is the mean squared error of the probabilities, so lower is better. Escalation should fall as the model fills in.

Loading ./results/index.json …

summary · all scenarios
Per predictor: mean Brier, mean accuracy, Jev latency, cost and errors

Latency and cost were measured on this workload, not taken from a datasheet. Jev measured about 0.9 seconds per call from this machine in early runs, against the vendor's claimed roughly 100 ms; the table above reports what the recorded runs saw. The browser demo cannot call Jev directly because the API disallows browser origins (CORS), so every Jev number here comes from Node runs written to ./results.

honest tradeoffs

What This Is Not

Neurosymbolic, not latent
Learning lives in the explicit schema and in the accommodator's revisions, not in any model's weights. That is the point, and the limit.
The serializer is load-bearing
The starting vocabulary caps what can be learned. A relation the serializer cannot emit can never appear in a rule. Growing that vocabulary is the main lever.
Jev cannot do arithmetic
All numbers are bucketed into categories before they reach it: near or far, few or many, soon or late.
Schema bloat needs pruning
Every surprise can add rules. Uninformative rules are pruned on each revision, and that pruning is a heuristic, not a proof.
Vendor benchmarks re-measured
Latency and cost quoted on this page are what this machine saw on this workload. Where they differ from the vendor's numbers, the page says so.