Reviewed by Jonathan West · Updated Oct 5, 2026

Laya vs Kev: Comparing Open-Source Decision Models

A technical comparison of encoder and decoder architectures for self-hosted classification, routing, and scoring.

Reviewed by Jonathan West · Updated Oct 5, 2026

Pick Laya for routing on small hardware when questions have 20 or fewer options, and pick Kev for larger models or Jev compatibility. At Layer3Labs, we build and run automation systems where routing speed dictates overall throughput.

Both models bypass the token billing of generative language models. Laya uses an encoder architecture with ModernBERT and mmBERT backbones. Kev pairs a frozen Qwen decoder with a rank-16 Low-Rank Adaptation (LoRA) adapter and a pointer head.

Neither project has been tested against the other on a shared benchmark. Evaluating laya vs kev requires inspecting published numbers from Convai Innovations, Jared Palmer's Kev repository, and opper.ai. Review the full list of Laya alternatives for surrounding options.

Laya vs. Kev: Side-by-Side

DimensionLayaKev
Who Makes It and LicenseDeveloped by Convai Innovations under the Apache 2.0 license.Developed by Jared Palmer under the Apache 2.0 license.
Architecture and SizesNon-autoregressive encoder using ModernBERT-large (421M parameters) or mmBERT-base (322M parameters).Frozen Qwen3.5 decoder with a rank-16 LoRA adapter and pointer head in 0.8B, 4B, 9B, and fully trained 27B sizes.
API ContractDedicated schema served by the open-source laya package.POST /v1/systemone wire-compatible with the TypeSafe AI Python SDK.
SpeedConvai Innovations reports 32.8 ms for multilingual and 39.5 ms for English on a single Tesla T4 GPU.Jared Palmer reports Kev-4B at 18.1 ms model time on an NVIDIA H100, and 721 ms for 5 questions on an Apple M5.
Published AccuracyConvai Innovations reports 0.766 accuracy on its 2,000-decision test set and 0.950 on AG News.Jared Palmer reports test-split accuracy from 0.697 for 0.8B to 0.889 for 27B; opper.ai measured Kev-4B at 93.9% to 98.3%.
Calibration EvidenceConvai Innovations reports an Expected Calibration Error (ECE) of 0.081, but only after domain temperature fitting.opper.ai measured Kev-4B calibration error between 0.044 and 0.137, showing more drift than Jev.
Options per QuestionDegrades past 20 options at the default 256-token head budget based on Convai Innovations documentation.No option limit is published in the Kev repository.
Hardware to Self-HostRuns on a Tesla T4 GPU or CPU, with CPU execution times unpublished.Kev-4B runs on an Apple M5 or NVIDIA H100; Kev-27B requires an 80 GB GPU or 96 to 128 GB on a Mac.
Fine-Tuning PathKaggle notebook provided by Convai Innovations taking about 4 hours on free GPUs.CLI skill executing a full fine-tuning loop on Modal.
Best FitLow-latency classification workflows on smaller hardware with 20 or fewer options.Teams needing Jev SDK compatibility or larger model parameter sizes.

Are you one of these vendors? Update your listing


When to Pick Laya and When to Pick Kev

Target hardware and API compatibility determine the choice between Laya and Kev. Laya fits production systems needing fast classification on modest hardware. Kev fits pipelines requiring the Jev API standard or larger parameter footprints.

Neither model wins every routing task. Laya is built by Convai Innovations as an open-weight encoder. Kev is developed by Jared Palmer as an adapter family over Qwen.

Who these decision models are not for: neither model handles conversational text generation or report drafting. Teams needing generative prose should deploy standard Large Language Models (LLM) instead. Decision models focus strictly on categorical choices and numerical scores.

  • Pick Laya if you run on a Tesla T4 GPU and need classification latencies under 40 ms.
  • Pick Laya if your routing queries contain 20 or fewer options.
  • Pick Kev if you build with the TypeSafe Python Software Development Kit (SDK) and need a self-hosted backend.
  • Pick Kev if your workload requires larger parameter sizes up to 27B.

Two Different Designs

Laya and Kev solve decision tasks with different model structures. Laya is a non-autoregressive encoder architecture. It uses ModernBERT-large for English tasks and mmBERT-base for multilingual workloads.

Because Laya evaluates tokens simultaneously, it scores candidates without generating words sequentially. The English checkpoint contains 421M parameters with a 512-token context window. The multilingual checkpoint holds 322M parameters and defaults to a 1,024-token context window.

Kev relies on an autoregressive decoder backbone. Jared Palmer builds Kev on Qwen3.5 using a rank-16 LoRA adapter and a pointer head. The base model weights remain frozen while the adapter and pointer head train on decision data.

Kev offers four distinct checkpoints. These include Kev-0.8B, Kev-4B, Kev-9B, and Kev-27B. The three smaller variants use LoRA adapters, while Kev-27B trains all 51 GB of weights.

The difference shows up in model size. Laya's largest checkpoint has 421M parameters, while Kev's smallest has 0.8B and its largest 27B.

For foundational background on this model category, see our guide to System One decision models. Both projects release model weights under the Apache 2.0 license, and Convai publishes Laya's weights as safetensors files for direct local loading.


Speed and Hardware

Speed comparisons between Laya and Kev require careful interpretation. Nobody has benchmarked both models on identical hardware. Each project publishes separate timing metrics collected on different machines.

Convai Innovations tested Laya on a single Tesla T4 GPU, where the English checkpoint averages 39.5 ms for a single question. The multilingual checkpoint processes a single question in 32.8 ms.

Batching multiple questions improves Laya throughput on a Tesla T4 GPU. Evaluating 10 questions takes 158.6 ms on the English checkpoint, which equals 15.9 ms per question. On the multilingual checkpoint, 10 questions take 72.3 ms, or 7.2 ms per question.

At 50 questions, Laya takes 771 ms on English and 337 ms on multilingual. That scales to 6.8 ms per question for the multilingual checkpoint. Overall batched throughput reaches 103 to 332 questions per second on a single Tesla T4 GPU.

Laya also supports Central Processing Unit (CPU) execution. However, CPU execution speeds remain unpublished by Convai Innovations. Preloading model weights into memory is optional and eliminates initial call delays.

Jared Palmer measured Kev on different hardware configurations. On an Apple M5 using MLX, Kev-4B takes 721 ms for 5 questions on new text. When prompt tokens are cached, that latency drops to 136 ms.

On an NVIDIA H100, Kev-4B records 18.1 ms of model time. It handles about 101 requests per second with 64 concurrent clients. Serving Kev-27B requires an 80 GB GPU or 96 to 128 GB of memory on a Mac.


Accuracy Evidence on Each Side

Neither project has completed a shared head-to-head accuracy benchmark. The only shared reference point is Jev from TypeSafe AI. Evaluating accuracy between Laya and Kev requires reading through their published comparisons against Jev.

Convai Innovations published comparative scores against Jev 1.13.0 on multiple datasets. On Convai Innovations' own typed-decisions benchmark of 2,000 decisions, Laya achieved 0.766 accuracy while Jev reached 0.727. On AG News with 4 labels, Laya reached 0.950 while Jev scored 0.910.

On DAIR Emotion with 6 labels, Laya reached 0.595 against Jev's 0.480. On Banking77 with 77 labels, Jev achieved 0.870 while Laya dropped to 0.425. Convai Innovations measured every Laya number directly, whereas Jev figures came from third-party publications.

Base Laya checkpoints perform below the majority-class baseline without fine-tuning. On the 2,000-decision benchmark, base English Laya reached a soft accuracy of 0.332, and multilingual Laya scored 0.328. The majority-class soft-accuracy baseline sits at 0.461, showing that capability comes from fine-tuning.

Jared Palmer evaluated Kev against Jev across held-out splits. On new sources, Kev-0.8B scored 0.648 accuracy with a Brier score of 0.481. Kev-4B reached 0.817 accuracy with a 0.269 Brier score, while Kev-9B achieved 0.820 accuracy with a 0.289 Brier score.

Kev-27B reached 0.851 accuracy on new sources with a 0.225 Brier score. In Jared Palmer's test-split table, Kev rises from 0.838 at 4B to 0.889 at 27B. However, Kev-9B's Brier score of 0.289 on new sources is worse than Kev-4B's 0.269.

Independent testing exists only for Kev-4B. On September 25, 2026, opper.ai published an independent comparison between Kev-4B and Jev. The evaluation examined three distinct text classification tasks.

On arXiv category tagging, Jev scored 96.9% while Kev-4B reached 95.0%. On Stack Exchange site classification, Kev-4B achieved 98.3% while Jev scored 97.5%. On GitHub issue triage, Jev scored 95.1% while Kev-4B reached 93.9%.

opper.ai found accuracy within a couple of points either way and concluded that Kev is a close open alternative for straightforward classification. Readers can review further evaluation data in our Laya benchmarks guide.


Options per Question and Calibration

Laya and Jev differ most on how many options one question can hold. Convai Innovations reports that Laya accuracy degrades when questions exceed 20 options at the default 256-token head budget.

The Banking77 benchmark highlights this constraint. Jev supports up to 255 options per question and achieved 0.870 accuracy. Laya reached only 0.425 accuracy on that dataset because its questions exceed 20 options.

Kev's repository publishes no maximum option limit. Its behavior on wide option lists has not been documented in public benchmarks. Teams needing wide categorical sets should verify Kev against their specific label distributions.

Calibration error reveals how closely output probabilities match actual outcomes. Lower calibration numbers indicate more dependable probability scores. Both models show distinct calibration patterns under testing.

Convai Innovations reports an ECE of 0.081 for Laya in its Laya-vs-Jev comparison table. That compares favorably with Jev's published ECE of 0.246. However, Laya reaches 0.081 only after domain temperature fitting, while raw ECE is higher than Jev's.

opper.ai measured Kev-4B calibration against Jev directly. Jev maintained calibration error between 0.027 and 0.049 across tasks. Kev-4B showed wider variance, with calibration error ranging from 0.044 to 0.137.

Confident error rates provide another perspective on reliability. Kev-9B assigned at least 0.9 probability to a wrong answer on 2.4% of new-source questions. Jev produced confident errors on 3.7% of those same questions in Jared Palmer's evaluations.


Jev Compatibility and Migration

API compatibility determines migration effort across different architectures. Kev is engineered specifically as a drop-in alternative to Jev. It serves the POST /v1/systemone endpoint natively.

The TypeSafe Python SDK works against a Kev server unchanged once developers point the client at the Kev server URL. For a direct evaluation of these two systems, see our Jev vs Kev comparison.

Running Kev locally requires standard command-line tools. You clone the repository, execute uv sync --extra serve, and start the server process. The launch command is uv run python -m kev.serve --run jaredpalmer/kev-4b --port 8009.

Laya uses its own open-source API contract. Convai Innovations distributes the package via pip install laya. Serving capabilities install through pip install "laya[serve]".

Laya does not implement the POST /v1/systemone specification. Connecting a Jev-based workflow to Laya requires modifying client request schemas. Review our guide on how to use Laya for client integration patterns.

Several other community efforts implement alternative Jev designs. SemIf, NanoJev, and jevlike each test a different routing design. Teams evaluating open options can review our roundup of Jev alternatives.


Fine-Tuning Each One

Fine-tuning adapts both models to domain-specific decisions. Laya base checkpoints require fine-tuning to perform above baseline metrics on complex decisions. Kev supports fine-tuning to specialize its pointer heads for proprietary classifications.

Convai Innovations provides a public Kaggle notebook for fine-tuning Laya. Training on a custom dataset takes about 4 hours on free GPUs.

Kev's training runs in the cloud: Jared Palmer built a fine-tuning skill that runs the whole loop on Modal. Developers start the loop by executing npx skills add jaredpalmer/kev@kev-finetune.

Parameter updates differ between the Kev model tiers. On Kev-0.8B, Kev-4B, and Kev-9B, only the rank-16 LoRA adapter and pointer head train. On Kev-27B, all weights across the full 51 GB model are trained.


Conditions That Would Change the Pick

Three developments would change this pick, and each depends on numbers neither project has published.

A direct head-to-head benchmark on identical hardware would clarify relative accuracy. Currently, comparing Laya and Kev requires indirect extrapolation through Jev.

Expanding Laya's option capacity would broaden its applicability. Convai reports degradation past 20 options at the default 256-token head budget. If Convai Innovations raises that ceiling without accuracy loss, Laya would suit wide categorization tasks.

Multilingual evaluation data could shift decisions for global deployments. Convai Innovations reports laya-multilingual is usable in 48 of 51 languages on MASSIVE. Kev publishes no multilingual benchmark numbers.

If Jared Palmer releases competitive multilingual metrics, Kev would challenge Laya across international pipelines. Until then, Laya is the only one of the two with published multilingual numbers. To select between laya vs kev, audit your query option counts, target hardware, and API migration budgets.


The Verdict

Pick Laya if you need low latency on smaller hardware and your classification tasks use 20 or fewer options. Its non-autoregressive encoder architecture delivers 32.8 ms to 39.5 ms response times on a Tesla T4 GPU. Its checkpoints are 322M and 421M parameters, small enough for a single T4.

Pick Kev if you need wire compatibility with Jev or require larger parameter sizes up to 27B. Kev serves the POST /v1/systemone contract, allowing the TypeSafe Python SDK to function without client code changes. Kev-4B has published serving figures on both an Apple M5 and an NVIDIA H100.

If your workflow needs chat or creative writing, choose neither. Laya and Kev are decision classifiers that return discrete choices and scores. Neither generates text. Teams with open-ended chat requirements should remain on standard generative language models.

Sources & Disclaimer

Researched from primary vendor documentation and public regulator sources. Pricing and availability are accurate as of Oct 5, 2026 and can change — confirm current terms with each vendor before you buy.

Frequently Asked Questions

  • Neither wins outright, because nobody has benchmarked them on a shared test set. Laya fits low-latency work on a single Tesla T4 when questions have 20 or fewer options. Kev offers larger parameter sizes up to 27B and drop-in compatibility with the Jev API.
  • Nobody has timed the two on identical hardware. Convai Innovations reports Laya latency of 32.8 ms for multilingual and 39.5 ms for English on a single Tesla T4 GPU. Jared Palmer reports Kev-4B at 18.1 ms model time on an NVIDIA H100, and 721 ms for 5 questions on an Apple M5.
  • Kev can replace Jev without client code changes because it serves the standard POST /v1/systemone API contract used by the TypeSafe Python SDK. Laya uses its own API schema, meaning existing Jev integrations require updating request payloads and client code.
  • Laya supports CPU execution according to Convai Innovations, though specific CPU latency numbers are unpublished. Kev runs on Apple Silicon Macs using MLX, where Kev-4B takes 721 ms for 5 questions on new text and 136 ms when cached. Kev-27B requires 96 to 128 GB of unified memory on a Mac.
  • Yes. Both Laya and Kev are released under the Apache 2.0 open-source license. Organizations can inspect weights, fine-tune models on internal data, and deploy them in commercial production without paying software licensing fees.
  • Convai Innovations documents that Laya accuracy degrades when questions exceed 20 options at the default 256-token head budget. The Kev repository publishes no upper limit on option count per question.

Need Help Architecting Your Decision Pipeline?

Book a free 30-minute AI workflow audit. We evaluate your classification latency, hardware footprint, and API contracts to benchmark the right open decision model.

Book a Consultation