Analysis · Open Models

Laya lands as Jev's open alternative

TypeSafe proved System 1 decisions were a real product. Convai proved they did not need to be closed.

R
RAR Editor
Published September 2026 · 5 min read

Drafted by an AI agent · reviewed and approved by a human editor before publication. How this works.

The Quick Version
  • TypeSafe AI released Jev on 15 September 2026 — a 'System 1' model returning typed probabilities in 70 to 500 ms, priced at $0.042 per million input tokens.
  • Convai Innovations shipped Laya under Apache 2.0 on 18 September — the open-weights alternative, free to host, free per token.
  • Flowtivity's independent benchmark gives Laya the edge on four of six measured dimensions — speed, cost, hard-label accuracy and calibration.
  • Jev still leads on Banking77's wide option set (0.870 vs 0.425) and on soft accuracy (0.580 vs 0.471), per the same benchmark.
  • The open-weights release lands three days after Jev's closed launch and is anchored to an arXiv paper the Laya team says predates TypeSafe's work.

A new model category, opened twice in three days

TypeSafe AI released Jev on 15 September 2026: a small “System 1” model that returns typed probabilities instead of text, with no string output and no risk of malformed JSON (TypeSafe). Three days later, on 18 September, Convai Innovations shipped Laya under Apache 2.0: the same architectural idea, with weights downloadable and a one-line pip install laya (Flowtivity benchmark). Flowtivity published a benchmarked head-to-head on 21 September that scores the showdown across six dimensions.

The headline numbers belong to Laya. Open weights answer in tens of milliseconds on commodity GPU hardware; Jev’s closed service runs in 70 to 500 ms. Jev charges per token; Laya costs nothing per token because you host it yourself. Jev sits behind a waitlist; Laya’s weights, training code and router are on Hugging Face and PyPI today.

4 of 6measured dimensions where Laya beats Jev — Flowtivity benchmark, 21 September

The trade Jev makes for type-safety

TypeSafe, founded by ex-OpenAI researcher Diogo Almeida, calls the new class “System 1 Models” — fast, structured, automation-first. The trade is explicit: Jev cannot generate strings. You send it state — text, JSON, a ticket, an email — plus typed questions whose answer spaces you define. It returns probabilities and confidence values across every question in parallel, in one forward pass.

That sounds like a downgrade until you wire it into a workflow. A hallucinated tool call is annoying in a chat assistant and catastrophic inside a latency-critical dependency chain. Jev’s pitch is that type-safety is table stakes for automation. The cost: you cannot ask Jev to write me a polite apology email — only given this complaint, is the customer churn risk above 50%? How urgent?

Where Laya pulls ahead, and where it does not

Flowtivity tested both and the scorecard is mixed. Laya’s routed stack beats Jev on four of six measured dimensions:

  • Speed: roughly 7.8× faster per question
  • Cost: $0 self-hosted vs $0.042 per million input tokens metered
  • Typed-decisions accuracy (picking the single best answer from your defined options): 0.766 vs 0.727 on a fine-tuned Laya checkpoint; 0.362 zero-shot (base model, no extra training)
  • Calibration (how often the model’s stated confidence matches its real accuracy): 0.081 vs 0.246 expected error after temperature refitting
  • Language coverage: 45 of 51 languages usable on Laya vs no published benchmark for Jev

Jev still leads on two:

  • Banking77 (a wide-option classifier): Jev 0.870, Laya 0.425 — option count above 20 starts to degrade Laya’s accuracy
  • Soft accuracy (partial credit): Jev 0.580, Laya 0.471

Flowtivity flags the asterisk on the headline accuracy figure: the 0.766 number is from a Laya checkpoint fine-tuned on the benchmark’s training split; zero-shot base scores 0.362, below the 0.461 majority-class baseline. The Hugging Face model card is candid: Laya is a fast base to specialise, not a zero-shot decision engine.

The origin story, and why it matters

Convai founder Nandakishor Mukkunnoth claims the underlying idea predates TypeSafe’s work. Flowtivity cites two arXiv papers from March and September 2025 — one on reinforcement-learned sales-conversion decisions and a follow-up formalising schema-based decisions guided by reinforcement learning. TypeSafe launched Jev without technical papers, open weights, or training datasets.

Two caveats from Flowtivity’s review: the three days between launches separate publication dates, not verified development timelines; the David-and-Goliath framing is marketing-adjacent. The part that is simply true is the one that matters operationally: the code and weights are downloadable, and you can inspect, refit and verify everything yourself.

Three shifts worth watching

  1. The category is real but unproven at scale. TypeSafe’s claimed 193.6× speed and 444.6× cost wins over frontier models come from the company’s own evaluation. Independent head-to-heads against GPT-6 Astra and Fable 5.1 on production automation workloads are sparse. Watch for benchmarks outside TypeSafe’s workflow suite — particularly long-context and multi-step decision tasks where Flowtivity noted Jev’s full numbers have not been independently measured.
  2. The option-count ceiling is the open question. Laya degrades past around 20 options; Jev scales to 77 on Banking77. If decision models become the routing layer beneath LLM agents, the routing layer needs to handle wider option spaces without accuracy collapse. Watch for either team publishing numbers past 100 options.
  3. Sovereign local decision-making is now a category. An Apache 2.0 model on hardware a UK small team already owns, returning typed probabilities at no per-token cost, is a different kind of automation primitive. It is, in TypeSafe’s framing, a frontier-intelligence function call — but with the keys in your hand.

Sources & quotes

Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →

  1. Introducing System One Models & Jev — TypeSafe AI Blog
  2. Laya: The Open-Source Jev Alternative, Benchmarked Honestly — Flowtivity
Filed under Analysis · Open Models

Continue Reading