TypeSafe AI released a model called Jev on Monday and it's been stuck in my head ever since, so I'm going to think out loud here rather than sit on it.
Quick background for anyone who hasn't seen it: Jev is what they're calling a "System One Model." Instead of generating text token-by-token like a normal LLM, it takes unstructured input and returns only values from a schema you define ahead of time — all in parallel, each with a calibrated confidence score. No free text at all. The claim is it structurally can't hallucinate because there's nowhere for it to improvise. They're also claiming ~70–500ms response times and pricing around $0.042/MTok. Big caveats: it's four days old, self-published benchmarks, no independent validation yet.
What interests me isn't Jev specifically — it's what its design implies for ICM, and specifically for the tool I've been building on top of it.
Where I'm coming from:
For Competition #11, it, I created a Security Cartographer. I am still playing with it, as I think there is value in the concept. It walks a body of work and maps everything that can influence an AI agent — files, websites, retrieval results, persistent memory, tool definitions, hidden content, agent handoffs, action paths — and outputs three things: an interactive HTML map, a Markdown field guide, and a machine-readable trust graph. ICM is the backbone: verified Instructions carry bounded workflow authority, Context is data only, Memory preserves state without instruction authority.
The attack that motivated it: an approved Markdown workflow points to a legitimate website, and a bad actor later changes that site. The file is unchanged, the hash is unchanged, but the agent gets different instructions. So the tool records exact approved content, seals policy and baseline, detects later changes, quarantines replacement content, and produces a run-specific exact-hash ICM bundle. There's also an independent authorization boundary between agent reasoning and consequential actions — uploads, code execution, credentials, memory modification, tool installation. It has 23 adversarial tests against synthetic data and is positioned as a reference monitor / defense-in-depth component, not a claim to eliminate prompt injection.
Mapping Jev onto the ICM buckets:
- Context is where it fits most naturally. Context is supposed to carry zero instruction authority, and a schema-locked model enforces that by construction — there's no text output channel for hostile content to hijack into a new instruction. It could still be tricked into misclassifying something (scoring malicious content as benign with high confidence), but that's a narrower failure mode worth tracking as its own node in the trust graph.
- Instructions get relocated rather than eliminated. The schema now does the work an instruction prompt used to do. A gap or bias in schema design is a security gap, not a UX one.
- Memory doesn't really apply — Jev is stateless per call. Which is actually what makes it interesting for the authorization boundary: a fast, cheap classifier gating whether an agent's proposed action matches an expected pattern before it crosses into consequential territory.
Do you need Jev to get this?
For the safety properties, I don't think so. Any LLM with structured output / function calling can produce the same shape of response — fixed categories, typed values, confidence scores, no free text. What you can't replicate without their architecture is the speed and cost. So you get the reliability shape without the performance shape.
The actual idea I want feedback on:
Right now Security Cartographer's trust graph captures what can influence the agent and whether it changed. What it doesn't formalize is the classification decision itself — how something gets assigned to Instructions vs. Context vs. Memory in the first place. In my current setup that's still rules plus some free-text reasoning, which is uncomfortable given it's arguably the most consequential decision in the whole chain. It's the part I've formalized the least.
The idea: add a fourth artifact alongside the map, field guide, and trust graph — a schema-locked classification layer. Every item the tool walks gets the same typed questions (Is this I/C/M? Confidence? Anomaly flags?) with no ability to answer outside those options. Confidence becomes a queryable field in the trust graph rather than a narrative aside, and low-confidence or flagged items go into the same quarantine path as replaced web content.
The reasoning: if Context carries zero instruction authority, the thing deciding "this is just Context" shouldn't itself be persuadable into the wrong answer the way free-text reasoning can be.
What I'm not sure about:
- Does this actually reduce attack surface, or just relocate it into schema design?
- Has anyone treated the I/C/M classification step as its own trust boundary with dedicated tooling? Did it hold up under adversarial testing?
- Should the classifier's confidence scores be treated as ground truth, or do they need the same adversarial scrutiny as any other component gating the authorization boundary? My instinct is the latter, which means those 23 tests need company.
Would genuinely appreciate people poking holes in this. Feels like it's either a reasonably clean extension of what's already there, or I'm missing something obvious.