Reviewed by Jonathan West · Updated Oct 5, 2026

AI Decision Models Like Jev and Laya You Can Use Today

Jev from TypeSafe AI plus five open projects you can self-host, and the job each one fits.

Reviewed by Jonathan West · Updated Oct 5, 2026

An AI decision model answers a typed question (pick an option, give a score, or say yes or no) with probabilities in one pass instead of writing text. At Layer3Labs, we build routing and document pipelines inside client software, and we send bounded decisions to fast models so the slower generative calls run less often.

These systems, often called System One models, solve bounded operational tasks like ticket triage, intent routing, and guardrail validation in one forward pass.

Technical decision models differ fundamentally from business management frameworks like decision trees or matrices. They are neural networks optimized for direct probability estimation across discrete options without running an autoregressive decoding loop. Six exist today. Jev is the managed cloud service from TypeSafe AI. The five open-source projects are Laya from Convai Innovations, Kev by Jared Palmer, SemIf by TheoLeeCJ, NanoJev by TianyuCodings, and the jevlike research starter by vinnylarouge.


The Decision Models You Can Use Today

Engineers evaluating this space can deploy one managed cloud service or choose among five distinct open-source projects depending on their hosting requirements and schema complexity.

The primary dividing line across these options is operational overhead versus architectural control. Managed endpoints handle infrastructure scaling at a per-token price, while open-weight models let engineering teams self-host inference workers on private Graphics Processing Units (GPUs) or local hardware.

  • Jev (TypeSafe AI): closed-source cloud API with an undisclosed parameter size. Use Jev when you want zero infrastructure maintenance, support for up to 255 choice options, and calibrated probability outputs out of the box.
  • Laya (Convai Innovations): open-weight ModernBERT-large and mmBERT-base checkpoints (322 million to 421 million parameters). Use Laya when you want self-hosted speed (32.8 to 39.5 milliseconds per single question on a Tesla T4 GPU) and your questions have 20 or fewer options.
  • Kev (Jared Palmer): open-weight Qwen3.5 and Qwen3.8 adaptations (0.8 billion to 27 billion parameters). Use Kev when you need a self-hosted, wire-compatible drop-in for Jev's POST /v1/systemone endpoint on private infrastructure.
  • SemIf (TheoLeeCJ): open-source, MIT-licensed interface that reads logits directly from open models like Qwen and MiniCPM. Use SemIf when you want to extract single-pass decisions from arbitrary foundation models without custom pointer heads.
  • NanoJev (TianyuCodings): 0.6 billion parameter model with parallel decision heads under the MIT license. Use NanoJev only to study parallel decision heads: it needs CUDA, and its only published results are ViZDoom game-control tests.
  • jevlike (vinnylarouge): byte-embedding research baseline under the MIT license. Use jevlike as an educational starter to study how single-pass option scoring functions without large decoder overhead.
Choose Jev if your team avoids running hardware, Laya for 32.8 to 39.5 ms answers on a Tesla T4 with 20 or fewer options, or Kev if you require an open-source drop-in for Jev's API.

Differences Between Decision Models and Generative LLMs

Generative Large Language Models (LLMs) operate autoregressively by predicting the next token in an open-ended sequence, requiring dozens or hundreds of sequential forward passes to generate a complete response.

In contrast, an AI decision model processes both the prompt and the discrete candidate choices simultaneously in a single forward pass. Instead of emitting conversational text that requires JavaScript Object Notation (JSON) parsers or regex extractors, the architecture calculates normalized probability distributions across the provided options directly.

A single pass removes JSON formatting errors. A decision model cannot hallucinate unstructured prose, drift off topic, or drop JSON closing brackets because it lacks a generative decoding loop entirely.

  • Inference mechanics: single-pass forward evaluation instead of step-by-step token generation.
  • Output format: probability scores and typed class selections rather than unstructured text.
  • Primary roles: intent routing, customer support ticket triage, prompt injection guardrails, and automated spam filtering.
  • Functional constraints: decision models cannot write email replies, summarize contracts, or explain why they selected a given probability score.
Decision models act as high-speed system reflexes for bounded classification; they do not replace generative models for document drafting or analytical reasoning.

Jev and the Managed System One Baseline

TypeSafe AI released Jev in early access on September 15, 2026, as its first System One model, a managed API with an undisclosed model size.

The service charges $0.042 per million input tokens, with output tokens listed as completely free. You call Jev through a single endpoint, POST /v1/systemone, through TypeSafe's official Python Software Development Kit (SDK).

One defining technical capability of Jev is its capacity to evaluate up to 255 candidate options within a single choice query. That lets one question pick from up to 255 billing or support tags in a ticketing system.

Within a week of Jev's release, developers built about 30 independent open reproductions. For benchmarks and migration paths, see our breakdown of Jev alternatives.

  • Provider and release: launched by TypeSafe AI on September 15, 2026.
  • Pricing model: $0.042 per million input tokens with free output delivery.
  • Interface design: accessed via POST /v1/systemone using the TypeSafe Python SDK.
  • Option capacity: evaluates up to 255 discrete candidate labels per query.
  • Underlying architecture: closed-source model with undisclosed parameter count and hosted infrastructure.
Jev offers up to 255 options per question and the simplest setup for teams that cannot manage private GPU clusters.

Laya and the Open-Weight Encoder Architecture

Convai Innovations published Laya as an open-weight decision model under the Apache 2.0 license, providing weights on Hugging Face and code on GitHub.

Rather than adopting a large generative decoder backbone, Laya uses bidirectional encoder backbones based on ModernBERT-large and mmBERT-base. The project distributes three specialized checkpoints: an English model (421 million parameters, 512-token context), a multilingual model (322 million parameters, 1,024-token context), and a typed-decisions checkpoint (421 million parameters, 1,024-token context).

Convai timed Laya on a single NVIDIA Tesla T4 GPU. A single question processes in 39.5 milliseconds on the base English checkpoint and 32.8 milliseconds on the multilingual checkpoint. Throughput scales from 103 to 332 questions per second under batching on the same T4 card.

Convai's documentation reports that accuracy degrades when a question has more than 20 options at the default 256-token head budget. On the 77-label Banking77 benchmark, Jev scored 0.870 while Laya achieved 0.425. For complete technical performance breakdowns, see our guide to Laya benchmarks and our direct Laya vs Jev comparison.

  • Checkpoints: 421M English base, 322M multilingual (100+ languages), and 421M typed-decisions.
  • Hardware speed: 32.8 to 39.5 milliseconds for single queries on an Nvidia Tesla T4 GPU.
  • Batch capacity: achieves 103 to 332 decisions per second on a single T4 card.
  • Option boundary: accuracy drops when evaluating more than 20 options per prompt.
  • Licensing: Apache 2.0 open-weight model with zero per-token software fees.
Laya answers a single question in 32.8 to 39.5 milliseconds on a Tesla T4 GPU, but teams routing across large taxonomies must respect its 20-option boundary.

Kev and the Drop-In Open Model Family

Kev, developed by Jared Palmer under the Apache 2.0 license, is an open-weight family of decision models built on Qwen foundation models to provide a self-hosted alternative to Jev.

The project supplies four checkpoint tiers: Kev-0.8B (Qwen3.5-0.8B-Base), Kev-4B (Qwen3.5-4B-Base), Kev-9B (Qwen3.5-9B-Base), and Kev-27B (Qwen3.8-27B, 51 gigabytes). The smaller variants use a rank-16 Low-Rank Adaptation (LoRA) adapter with a pointer head on a frozen backbone, while Kev-27B trains all model weights.

Kev is engineered specifically as a drop-in substitute for TypeSafe's contract. It serves the POST /v1/systemone endpoint, enabling applications built with the TypeSafe Python SDK to switch targets simply by updating the base URL to point at a local Kev server.

On the Kev author's new-source benchmark, which uses held-out datasets, Kev-0.8B scored 0.648, Kev-4B 0.817, Kev-9B 0.820 and Kev-27B 0.851. Jev's reference score is 0.857. In independent testing by opper.ai on September 25, 2026, Kev-4B scored 95.0% to Jev's 96.9% on arXiv categories and 98.3% to 97.5% on Stack Exchange sites. On GitHub bug/feature labels, Kev-4B scored 93.9% to Jev's 95.1%. Jev's probabilities were better calibrated. For head-to-heads, see our Jev vs Kev comparison and Laya vs Kev comparison.

  • Model tiers: Kev-0.8B, Kev-4B, Kev-9B, and full-weight Kev-27B.
  • Protocol compatibility: exposes the POST /v1/systemone endpoint for instant TypeSafe SDK compatibility.
  • Author benchmarks: Kev-27B achieved 0.851 accuracy on held-out sources compared to Jev's 0.857.
  • Third-party validation: opper.ai confirmed Kev-4B tracks Jev classification accuracy within normal sampling margins.
  • Hardware requirements: Kev-4B answers 5 questions in 721 ms on an Apple M5 (MLX) and runs at 18.1 ms model time on an H100, while Kev-27B requires an 80-gigabyte GPU or 96 to 128 gigabytes of shared memory on Apple Silicon.
Kev is the most direct self-hosted replacement for Jev, allowing teams to keep their existing client code while running inference on private infrastructure.

SemIf, NanoJev, and Experimental Open Implementations

Beyond Laya and Kev, the open-source ecosystem includes several specialized projects exploring alternative architectures for non-autoregressive decision scoring.

SemIf, formerly OpenJev, is an MIT-licensed project by TheoLeeCJ. It makes runtime-defined decisions by reading option probabilities straight from an open model's logits in one forward pass. Tested across Qwen3-0.6B, MiniCPM5-2B, and Qwen3.5-4B, SemIf bypasses generative text output and JSON parsing without training custom pointer heads. On an NVIDIA RTX 3090, its 4-bit Qwen3.5-4B setup scored 0.813 balanced accuracy on its own authored decisions. It was 5.21 times faster than compact JSON generation (1.023 s vs 5.332 s) and reached 20.03 decisions per second with parallel suffixes.

NanoJev, developed by TianyuCodings under the MIT license, is a compact 0.6 billion parameter model combining a Qwen3-0.6B backbone with parallel decision heads. Rather than text classification, its published benchmarks reflect game-control simulations in ViZDoom, where it achieved 128 of 128 basic task successes compared to Jev's 56 of 128.

jevlike, by vinnylarouge, is an MIT-licensed research starter. It uses byte embeddings trained from scratch, or an optional frozen Qwen2.5-0.5B encoder. One pass ran about 100 times faster than a small decoder, but its author calls it a research starter, so do not plan production work on it.

  • SemIf: extracts decision probabilities directly from raw model logits without requiring dedicated adapters.
  • NanoJev: 0.6B parameter parallel architecture benchmarked only on ViZDoom game-control tests.
  • jevlike: lightweight byte-embedding research prototype exploring multi-option scoring mechanics.
  • Maturity status: SemIf provides practical utility for custom model experiments, while NanoJev and jevlike remain exploratory research repositories.
For production software routing, stick to Jev, Laya, or Kev; SemIf, NanoJev, and jevlike are experimental frameworks for researchers exploring alternative scoring mechanics.

Traditional Classification Alternatives for Bounded Tasks

Engineering teams faced with classification and routing workloads do not always require dedicated System One decision models; established machine learning approaches remain competitive.

Bidirectional Encoder Representations from Transformers (BERT) classifiers and modern cross-encoders offer tiny memory footprints and rapid inference when trained on fixed label sets. If your routing taxonomy never changes at runtime, fine-tuning a small classification head on standard encoder embeddings is simpler to maintain than deploying newer decision architectures.

Zero-shot text classifiers, small LLMs prompted with strict structured JSON output schemas, and deterministic rule engines also satisfy bounded automation tasks. Deterministic regex or keyword rules remain the fastest option for high-volume routing when inputs contain explicit identifiers like account numbers or exact department codes.

Decision models become necessary when your application demands dynamic, runtime-defined choice sets with calibrated probabilities that static classifiers cannot support. To review broader classification designs, examine our guide to AI model routing.

  • Fine-tuned encoders: optimal for static taxonomies where labels never change after deployment.
  • Zero-shot text classifiers: viable for infrequent classification queries with minimal engineering investment.
  • Structured LLM prompting: suitable for low-volume applications that tolerate multi-second response latency.
  • Deterministic rule engines: fastest and most cost-effective method for exact keyword and regex matching.
If your category list is static and known at training time, a conventional fine-tuned encoder often matches decision model accuracy with lower computational requirements.

Architectural Comparison and Selection Criteria

Selecting the correct AI decision model requires balancing infrastructure control, query volume, candidate option counts, and software integration requirements.

Organizations operating high-volume routing layers must evaluate whether hosting dedicated GPU instances delivers a lower total cost of ownership compared to paying TypeSafe's $0.042 per million input tokens. Across the workflows we have automated for small and mid-sized business (SMB) teams, self-hosting makes economic sense only when query volume generates continuous server utilization.

  • Select Jev if: you require zero infrastructure maintenance, need to evaluate up to 255 options, and prefer a managed cloud API.
  • Select Laya if: you need self-hosted speed of 32.8 to 39.5 ms per question on a Tesla T4, your prompts feature 20 or fewer options, and you want an Apache 2.0 license.
  • Select Kev if: you need an open-source model running on private GPUs that drops directly into existing TypeSafe SDK implementations.
  • Select SemIf if: you want to experiment with logit extraction across arbitrary open-source foundation checkpoints without fine-tuning custom pointer heads.
The primary operational tradeoff is simple: choose Jev for zero-ops convenience across wide option lists, or deploy Laya and Kev on private GPUs to eliminate API fees.

Implementation Tradeoffs and Operational Boundaries

AI decision models are not suitable for organizations seeking conversational chatbots, customer-facing draft generation, or complex legal synthesis. Attempting to force an open-ended generative task into a single-pass scoring model will fail because decision architectures possess no conversational text decoders.

Our architectural recommendation flips if your application handles low or highly irregular query volume. At a few dozen decisions a day, a dedicated GPU for Laya or Kev sits idle most of the time. Jev's metered endpoint, or a standard generative API, costs less at that volume.

Start with our how to use Laya guide and run a sample of your own classification prompts through it. Then compare the other open options in Laya alternatives.

  • Who this is not for: engineering teams needing creative text generation, multi-turn conversational agents, or long-form document drafting.
  • Failure condition: forcing qualitative, multi-step logical synthesis into a single-pass probability question.
  • Verdict flip: low or sporadic query volumes favor metered cloud endpoints over self-hosted server instances.
  • Next operational step: benchmark your candidate choice counts against Laya and Kev in a local test container.
Never deploy a decision model to draft responses; deploy it upstream to filter, score, and route requests before calling generative services.

Frequently Asked Questions

  • An AI decision model is a non-autoregressive neural network that evaluates prompts against typed choices, numerical scales, or boolean conditions in a single forward pass. Rather than generating text token by token, it returns calibrated probabilities for bounded classification, scoring, and routing tasks.
  • Jev fits teams that want a zero-maintenance cloud API with up to 255 options. Laya fits self-hosted jobs with 20 or fewer options, at 32.8 to 39.5 milliseconds per question on a Tesla T4 GPU. Kev fits teams that want a self-hosted drop-in for Jev on their own GPUs.
  • System One is TypeSafe AI's name for a fast decision model that answers a typed question with probabilities in one pass instead of writing text. Jev was the first; Laya, Kev and SemIf do the same job as open projects.
  • Yes. Kev, created by Jared Palmer, is an open-source model family that serves Jev's exact POST /v1/systemone endpoint under the Apache 2.0 license. Additionally, Laya, SemIf, NanoJev, and jevlike provide open-source architectures for single-pass decision scoring.
  • SemIf, formerly known as OpenJev, is an open-source project by TheoLeeCJ released under the MIT license. It extracts runtime decision probabilities directly from the output logits of open foundation models like Qwen in a single forward pass without requiring custom adapter training.
  • Use a generative Large Language Model (LLM) when your application must synthesize complex information, draft open-ended prose, engage in conversational dialogue, or provide written explanations for its conclusions. Decision models only return probability scores across predefined options.

Need help implementing fast AI decision pipelines?

Layer3Labs designs and deploys automated routing layers that pair low-latency decision models with generative AI. Book a consultation to see where local models can cut your infrastructure costs.

Book a Consultation