Laya Benchmarks for Accuracy and Speed on a Tesla T4
Convai Innovations' own Laya test results next to Jev, with no independent replication yet.
Laya scores 0.993 accuracy on email spam and answers one question in 32.8 milliseconds on an NVIDIA Tesla T4 graphics processing unit (GPU, multilingual checkpoint), per its maker's published benchmarks. At Layer3Labs, we build automated routing pipelines for production workflows, so we check vendor benchmark tables against the label counts and hardware a real deployment uses. These figures come directly from documentation published by Convai Innovations, covering evaluations against Jev from TypeSafe AI, standard public text sets, and multilingual language suites.
Unlike a generative large language model (LLM), Laya returns probabilities across predefined labels in a single forward pass instead of streaming tokens. Accuracy swings with the number of labels and with the checkpoint you pick. On low-cardinality tasks like AG News, Laya records 0.950 accuracy, but falls to 0.425 on the 77-label Banking77 dataset.
Every Laya benchmark reported to date represents vendor-measured results committed to the model repository on GitHub, with no independent third-party replication completed as of October 2026.
Laya Benchmarks Across Published Test Sets
Convai Innovations documents high accuracy across routine binary screening, single-topic text routing, and schema-constrained decisions, alongside a sharp drop on tasks with large option counts.
For multi-class tasks, accuracy depends on the number of labels. Convai reports that accuracy degrades past 20 options per question at the default 256-token head budget, and even 6-label DAIR Emotion tops out at 0.595.
Convai Innovations measured every Laya figure below. The Jev figures are third-party published numbers that Convai did not measure, and the latency figures come from a single NVIDIA Tesla T4 GPU.
- Email spam filtering: 0.993 accuracy (0.993 F1 score), measured by Convai Innovations.
- Phishing detection: 0.980 accuracy (0.979 F1 score), measured by Convai Innovations.
- Typed decisions benchmark: 0.766 overall accuracy across 2,000 decisions (400 test cases) for the fine-tuned checkpoint, compared to 0.727 for Jev 1.13.0.
- AG News categorization: 0.950 accuracy across 4 topic labels, compared to 0.910 for Jev 1.13.0.
- DAIR Emotion classification: 0.595 accuracy across 6 emotion labels, compared to 0.480 for Jev 1.13.0.
- Banking77 intent triage: 0.425 accuracy across 77 banking customer service categories, compared to 0.870 for Jev 1.13.0.
- Expected Calibration Error (ECE): 0.081 error score achieved after temperature fitting, compared to 0.246 for Jev 1.13.0.
- Single-question latency: 32.8 milliseconds on an Nvidia Tesla T4 GPU for the multilingual checkpoint and 39.5 milliseconds for the base English checkpoint.
Dataset Comparison Between Laya and Jev
Convai Innovations compared Laya with Jev 1.13.0 on four datasets: Laya wins on three and loses on Banking77, 0.425 to 0.870.
The comparative data indicates that Laya outperforms Jev on AG News (0.950 versus 0.910) and DAIR Emotion (0.595 versus 0.480). Convai did not measure the Jev figures itself. They are third-party published numbers.
The primary failure mode for Laya appears on Banking77, where Laya scores 0.425 against Jev's 0.870. Convai documents that Laya's accuracy degrades past 20 candidate choices at its default 256-token head budget, while Jev supports up to 255 options.
Laya is not suitable for organizations routing requests across large category taxonomies exceeding 20 choices, such as high-density enterprise support queues. Teams with those taxonomy requirements should use Jev or a larger generative language model, and our Laya vs Jev comparison covers that choice. For the wider field, see our map of AI decision models and the Laya vs Kev comparison.
- AG News (4 labels): Laya 0.950 vs Jev 0.910.
- DAIR Emotion (6 labels): Laya 0.595 vs Jev 0.480.
- Typed-decisions (2,000 decisions): Laya 0.766 vs Jev 0.727.
- Banking77 (77 labels): Laya 0.425 vs Jev 0.870.
- Calibration error (ECE): Laya 0.081 vs Jev 0.246.
- Single-question latency: Laya 32.8 ms (multilingual checkpoint, Tesla T4) vs Jev 236 to 276 ms, which Convai calls 7.8x faster.
The Typed-Decisions Benchmark and Fine-Tuning Impact
Convai says Laya's typed-decisions capability comes from fine-tuning: the base checkpoints score below the majority-class soft-accuracy baseline on the 2,000-decision test set.
Convai Innovations constructed the typed-decisions benchmark to evaluate 400 complex business cases requiring 2,000 discrete decisions across four workflows: invoice processing, security screening, customer support routing, and system observability. On this evaluation, the fine-tuned laya-typed-decisions checkpoint reaches 0.766 accuracy and 0.471 soft accuracy, ahead of Jev's 0.727 accuracy but behind Jev's 0.580 soft accuracy.
Without task-specific fine-tuning, the base models struggle with structured schema logic. The base English checkpoint (laya) scores 0.332 soft accuracy and the base multilingual checkpoint (laya-multilingual) scores 0.328, both below the 0.461 majority-class baseline. Their plain accuracy is 0.362 and 0.352.
Convai Innovations provides a Kaggle notebook that fine-tunes Laya on a custom dataset in about 4 hours on free GPUs.
- Invoice workflow accuracy: 0.804 for fine-tuned weights.
- Security workflow accuracy: 0.766 for fine-tuned weights.
- Support routing accuracy: 0.764 for fine-tuned weights.
- Observability triage accuracy: 0.730 for fine-tuned weights.
- Baseline comparison: base English (0.332 soft accuracy) and base multilingual (0.328) fall below the 0.461 majority-class soft-accuracy baseline. The fine-tuned checkpoint reaches 0.471.
- Error metrics: fine-tuned checkpoint cuts score Mean Absolute Error (MAE) from 0.694 down to 0.242 and Brier score from 0.316 down to 0.062.
Inference Speed on an Nvidia Tesla T4 GPU
On a single NVIDIA Tesla T4 GPU, Convai Innovations measured Laya at 32.8 to 39.5 milliseconds per single question and 103 to 332 questions per second batched.
The 32.8 millisecond headline latency comes specifically from the 322 million parameter laya-multilingual checkpoint built on an mmBERT-base backbone. The 421 million parameter English checkpoint built on ModernBERT-large records 39.5 milliseconds for a single question.
Batching cuts the time per question. For a 10-question evaluation, the multilingual checkpoint takes 72.3 milliseconds (7.2 milliseconds per question), while the English checkpoint takes 158.6 milliseconds (15.9 milliseconds per question).
Convai Innovations says its full docs include CPU and Apple Silicon (Metal Performance Shaders) results, but we could not verify those figures, so this page does not repeat them. Laya runs on a central processing unit (CPU), so time it on your own CPU before you plan hardware around it.
- Single-question latency: 32.8 ms for multilingual and 39.5 ms for English.
- 5-question batch latency: 40.1 ms for multilingual and 84.5 ms for English.
- 10-question batch latency: 72.3 ms (7.2 ms per question) for multilingual and 158.6 ms for English.
- 50-question batch latency: 337 ms (6.8 ms per question) for multilingual and 771 ms for English.
- Peak throughput: 103 to 332 questions per second when running fully batched evaluation passes.
- Memory efficiency: preloading weights into video memory removes cold-start overhead for downstream microservices.
Calibration Scores and Temperature Fitting
Expected Calibration Error measures the gap between predicted probability scores and actual classification accuracy, where lower numbers indicate more reliable confidence outputs.
Convai Innovations reports an Expected Calibration Error (ECE) of 0.081 for Laya against Jev's 0.246 in its Laya vs Jev table. However, Laya achieves this 0.081 score only after post-processing calibration via domain temperature fitting, and its raw uncalibrated error is higher than Jev's.
On the typed-decisions benchmark, the fine-tuned checkpoint records an ECE of 0.213 against Jev's 0.144. Jev's raw probabilities track real outcomes more closely out of the box.
When building automated routing logic, uncalibrated confidence scores can cause downstream failures if used as strict routing thresholds. Engineering teams must fit temperature parameters on a validation hold-out set before using Laya probabilities to gate critical decisions.
- Laya vs Jev table ECE: 0.081 for Laya post-temperature fitting vs 0.246 for Jev 1.13.0.
- Raw calibration behavior: uncalibrated Laya output exhibits higher error than Jev before temperature adjustments.
- Typed-decisions ECE: 0.213 for fine-tuned Laya vs 0.144 for Jev on the 2,000-decision evaluation.
- Brier score: fine-tuned Laya records 0.062 against Jev's 0.148 on typed decisions, reflecting tighter overall squared error.
- Operational consequence: confidence scores cannot be used as strict probability thresholds without prior validation on domain-specific traffic.
Multilingual Classification Performance
On international language benchmarks, the dedicated multilingual checkpoint maintains viable classification accuracy across 48 of 51 evaluated languages, while the English checkpoint clears that bar in only 23 of 51.
Evaluations on the MASSIVE intent dataset across 51 languages reveal distinct checkpoint boundaries. On native English intent routing, the 421M English checkpoint scores 0.783 accuracy, outperforming the multilingual checkpoint's 0.657.
Across 13 non-English languages in the MASSIVE corpus, the English checkpoint's accuracy falls to 0.306, while the multilingual checkpoint achieves 0.451. On the Cross-lingual Natural Language Inference (XNLI) benchmark, the multilingual model scores 0.731 across 14 non-English languages, whereas the English checkpoint drops to 0.521.
Convai Innovations defines viable cross-lingual routing as exceeding three times the random selection baseline.
- MASSIVE English accuracy: 0.783 for English checkpoint vs 0.657 for multilingual checkpoint.
- MASSIVE non-English accuracy: 0.306 for English checkpoint vs 0.451 across 13 international languages for multilingual checkpoint.
- XNLI English inference: 0.860 for English checkpoint vs 0.843 for multilingual checkpoint.
- XNLI non-English inference: 0.521 for English checkpoint vs 0.731 across 14 languages for multilingual checkpoint.
- Language coverage breadth: multilingual checkpoint exceeds three times random chance in 48 of 51 languages, compared to 23 of 51 for the English build.
- Context window capacity: multilingual checkpoint provides a 1,024-token base context extendable up to 8,192 tokens for international document classification.
Email Spam, Phishing, and Guardrail Results
Security evaluations conducted by Convai Innovations put Laya at 0.993 on spam and 0.980 on phishing, with 0.755 to 0.762 on jailbreak detection.
In email spam filtering benchmarks, Laya achieves a documented accuracy of 0.993 with an identical F1 score of 0.993.
For phishing detection, Laya records 0.980 accuracy and a 0.979 F1 score.
When deployed as an inbound security guardrail for generative chatbots, Laya achieves accuracy between 0.755 and 0.762 on LLM guardrail and jailbreak detection. That score leaves room for misses, so put a second guardrail behind Laya on sensitive systems.
- Email spam classification: 0.993 accuracy and 0.993 F1 score.
- Phishing detection: 0.980 accuracy and 0.979 F1 score.
- Generative guardrail screening: 0.755 to 0.762 accuracy on jailbreak detection.
- Pipeline latency: a single question takes 32.8 milliseconds (multilingual) or 39.5 milliseconds (English) on a Tesla T4 GPU, so screening runs before any paid API call.
Verification Status and Audit Procedures
All published Laya benchmark metrics represent internal evaluations conducted by Convai Innovations and committed to the research directory of the official repository, without independent third-party replication.
Unlike Kev from Jared Palmer, whose 4B checkpoint was tested independently by opper.ai on September 25, 2026, Laya's accuracy and latency numbers have not been replicated by outside laboratories. Furthermore, while datasets like MASSIVE, XNLI, AG News, DAIR Emotion, and Banking77 are open public standards, the 400-case typed-decisions benchmark is Convai Innovations' own corpus.
Convai Innovations did not run the Jev 1.13.0 figures side by side with Laya; they are third-party published numbers. Test Laya on a labeled hold-out set drawn from your own traffic before you deploy it.
Our answer would change if outside teams replicate Convai Innovations' tests, or if Convai ships a classifier head that handles more than 20 options without degrading. Our how to use Laya guide shows how to check the published laya benchmarks.
- Evaluation provenance: figures are committed directly within the research/results directory of the Laya code repository.
- Replication gap: no independent replication of Laya's benchmarks was found as of October 1, 2026.
- Dataset origin: typed-decisions corpus is proprietary to Convai Innovations, while general classification sets are public benchmarks.
- Comparative validity: the Jev figures are third-party published numbers that Convai did not measure itself.
- Recommended audit protocol: run labeled records from your own traffic through the fine-tuned checkpoint before you commit.
What you need to run Laya yourself
Laya is small enough to run genuinely locally — a single modern GPU, an Apple Silicon Mac, or even a mini PC handles it. You do not need a server or a cloud account; this is the tier where "run it yourself" is truly a one-box decision.
| Path | What it is | Best for | Get started |
|---|---|---|---|
| Single consumer GPU | A 16–24GB NVIDIA card runs a 4-bit quant with room to spare | One desktop you already game or work on | NVIDIA GeForce RTX 4090 |
| Apple Silicon Mac | Unified memory runs small models quietly at low power | Mac users who want always-on local AI | Apple Mac Mini (M4 Pro) |
| Mini PC | A compact, low-watt box that sits on a shelf and runs 24/7 | The cheapest always-on local option | AI mini PC |
| Call the hosted API | Skip hardware entirely and pay per token | Trying it before you commit to a box | OpenRouter |
To run Laya locally with almost no setup, use a one-click runner like Ollama or LM Studio — download the model and it is chatting in minutes. For specific hardware picks, see Best mini PCs for local AI and Local AI hardware calculator.


Frequently Asked Questions
- On a single Nvidia Tesla T4 GPU, Laya evaluates a single question in 32.8 milliseconds using the multilingual checkpoint and 39.5 milliseconds using the English checkpoint. In batches of 10 questions, that drops to 7.2 milliseconds per question for multilingual and 15.9 milliseconds for English. Batched throughput is 103 to 332 questions per second.
- Laya scores higher on narrow label sets: 0.950 on AG News (vs 0.910 for Jev), 0.595 on DAIR Emotion (vs 0.480), and 0.766 on typed decisions (vs 0.727). However, Jev scores 0.870 on 77-label Banking77 against Laya's 0.425.
- All published Laya benchmarks were measured internally by Convai Innovations and committed to the research directory of the official GitHub repository. No independent third-party research laboratory has replicated these results as of October 2026.
- Laya degrades sharply when a question presents more than 20 candidate choices, as seen in its 0.425 score on the 77-label Banking77 dataset. Additionally, base un-fine-tuned checkpoints score below the 0.461 majority-class soft-accuracy baseline on Convai's typed-decisions benchmark, and jailbreak detection accuracy remains moderate at 0.755 to 0.762.
- Yes. The 322 million parameter multilingual checkpoint supports over 100 languages, maintaining viable classification accuracy across 48 of 51 evaluated languages on the MASSIVE benchmark. In contrast, the base English checkpoint maintains viable performance in only 23 of 51 languages.
- All official latency benchmarks were recorded on a single Nvidia Tesla T4 GPU. Convai Innovations reports CPU and Apple Silicon results in its full docs, but not in a form we could verify.
- Calibration is evaluated using Expected Calibration Error, which measures the difference between confidence probabilities and actual accuracy. Laya reports an error score of 0.081, but this figure is achieved only after post-processing temperature fitting; raw uncalibrated predictions show higher error than Jev.
- Yes, Laya supports CPU execution, but Convai Innovations has not published CPU latency in a form we could verify. Time it on your own hardware against the 32.8 milliseconds Convai reports for the multilingual checkpoint on a Tesla T4 GPU.
Need help verifying AI decision models on your own data?
Layer3Labs designs low-latency routing pipelines and runs empirical benchmark audits on internal datasets. Book a consultation to see where reflex decision models outperform generative APIs.
Book a Consultation