MODEL

Jev

modeltopic-note

Overview

Jev is TypeSafe AI‘s non-autoregressive decision model — a new “System One” class returning typed probabilistic outputs (yes/no, category, calibrated score) rather than free-form text. Priced at $0.042/M input tokens with free output, running 40–200× faster than small frontier LLMs on inference. Surfaced publicly by Simon Willison on September 21, 2026; the differentiator Willison highlighted is that the confidence scores on typed slots are actually usable in downstream routing decisions, unlike LLM-produced probabilities that don’t calibrate.

Timeline

  • 2026-09-22-AI-Digest — Simon Willison flagged Jev on 2026-09-21 — a new “System One” / decision model class from TypeSafe AI returning typed probabilistic outputs rather than free-form text, priced at $0.042/M input tokens with free output, 40–200× faster inference than small frontier LLMs. Non-autoregressive architecture returns calibrated confidence scores on typed slots — the 40–200× speed delta is a direct consequence of the architecture (no autoregressive decode loop) rather than an optimization pass. HN front page discussion centred on Jared Palmer’s open-source Qwen3.5 refactor Kev (0.8B / 4B / 9B checkpoints, no Jev outputs used in training).

Key Developments

  1. Architecturally Distinct From LLM-JSON: Jev is architecturally distinct from “prompt an LLM for JSON output” or Cohere-style classify endpoints — the confidence numbers on typed slots are actually usable in downstream routing decisions, unlike LLM-produced probabilities that don’t calibrate.

  2. One-Vendor Category So Far: TypeSafe AI is the only vendor shipping this shape today, and one-vendor categories usually resolve as product rather than category. The load-bearing test is whether a second vendor (Cohere, Anthropic, or the open-weights ecosystem via Kev) ships a decision-model-shaped product in the next 90 days.

  3. 24-Hour Closed→Open Compression via Kev: Kev shipped on Qwen3.5 within 24 hours of Jev’s release, with an explicit “no Jev outputs used in training” claim — if it holds under scrutiny, another data point that the closed→open replication window for a novel model shape is now measured in hours, not weeks.