TECH.US AI RESEARCH

Tokle-3M: The #1-Ranked Language Model Under 3 Million Parameters

Tokle-3M ranks first among models under 3 million parameters on the Open SLM Leaderboard , as of 1st Oct 2026.

Developed by Tech.us as a research model, Tokle-3M has 2.91 million parameters and uses SPAB, a guided training approach that introduces additional information about relationships between tokens during training. The work explores how training methods, not just model size, influence how efficiently a model learns and performs.

2.91M Parameters
8.91 Intelligence Index
SPAB Guided Training

BENCHMARK COMPARISON

How Tokle-3M Performs Across the Benchmark Dataset

A side-by-side view of Tokle-3M results across reasoning, commonsense, arithmetic, and the composite Intelligence Index.

Try Tokle-3M
Benchmark Tokle-3M 2.91M Pulvis-v2 2.96Mx3* Purrence-3M 2.99Mx6* Ember-2 2.96Mx2*
HellaSwag 27.93% 27.81% 27.28%
Arc-Easy 31.86% 34.55% 33.42%
Arc-Challenge 22.78% 23.63% 22.01%
Piqa 56.86% 56.64% 55.11%
ArithMark-3 37.40% 34.80% 35.90%
Intelligence Index 8.62 8.49 7.21
Tokle-3M leads Top result from another model

RESEARCH OBJECTIVE

Designed Around Learning Efficiency

Tokle-3M was developed to study how architecture and training methods influence model performance when compactness is an intentional design choice.

The research focuses on how a purpose-built model can make effective use of what it learns. Rather than relying on model size alone, Tech.us is studying how additional guidance during training can help a model identify useful relationships and retain that learning.

We designed this model to be purpose-built from the start, with guided training helping it make better use of what it learns. The focus is not on making it broader, but on making it more deliberate, efficient and useful for defined tasks.

Praveen Narra Founder & CEO, Tech.us

TRAINING APPROACH

Guided During Training. Independent at Inference.

Tokle-3M grew out of Tech.us’ work on Static Pairwise Attention Bias (SPAB), a training approach designed to give the model additional information about relationships between tokens while it learns.

The process happens in two stages.

01

Guide the Learning

SPAB uses statistical patterns in the training data to identify words or tokens that are more likely to be related.

That information is added to the model's attention process during training, helping it focus on useful relationships earlier rather than having to discover all of them on its own.

02

Retain What Was Learned

The additional SPAB information is then removed. What the model learned with that guidance is distilled into the final Tokle-3M model.

The final model runs with approximately 2.91 million parameters without requiring the additional SPAB table at inference.

BENCHMARKS

8.91 Intelligence Index

Tokle-3M was evaluated across five established benchmarks using zero-shot normalized accuracy and the Open SLM Leaderboard methodology.

The Open SLM Leaderboard's Intelligence Index combines chance-adjusted results across reasoning, commonsense and arithmetic benchmarks into a composite measure.

  • HellaSwag Commonsense evaluation
  • ARC-Easy Reasoning evaluation
  • ARC-Challenge Reasoning evaluation
  • PIQA Commonsense evaluation
  • ArithMark-3 Arithmetic evaluation

MODEL PROFILE

Tokle-3M Technical Profile

A concise view of the model’s architecture, training approach, evaluation method, and current research scope.

Tokle-3M is not instruction-tuned or safety-aligned and is currently positioned as a research model rather than a production assistant.

Because the model uses Rotary Position Embeddings (RoPE), its context window can be extended at inference time without changing the architecture.

SpecificationDetail
ParametersApproximately 2.91 million
Training DataPre-trained on English datasets
Position EmbeddingsRotary Position Embeddings (RoPE)
Training ApproachSPAB-assisted guided training
Evaluation MethodologyOpen SLM Leaderboard methodology
Benchmark MethodZero-shot normalized accuracy
Current ScopeResearch model

RESEARCH FINDINGS

What the Research Shows

Tokle-3M gives Tech.us a focused way to examine how specific training and architectural decisions affect model performance.

Guided Training

The SPAB-assisted training process allows Tech.us to study how providing additional information about token relationships influences what the model learns and how much of that learning remains after the additional guidance is removed.

Intentional Architecture

Tokle-3M was designed as a compact model from the start. This allows the research to focus on how architecture and training choices contribute to performance within a deliberately defined model design.

Purpose-Built Models

The research also supports Tech.us’ exploration of models designed around defined tasks, domains and operating requirements rather than maximum breadth.

DEPLOYMENT

More Control Over Deployment

Tokle-3M supports Tech.us’ exploration of deployment approaches that give organizations greater control over where models run and how information is handled.

On-Device Deployment

Explore deployment closer to the point of use.

Self-Hosted Environments

Explore deployment within infrastructure the organization manages.

Sensitive Information Control

Support approaches that keep more processing within controlled environments.

Could an SLM Fit Your Use Case?

Tokle-3M is one example of how Tech.us is exploring more focused approaches to AI. If your organization is considering Small Language Models for a specific workflow, deployment environment or data requirement, talk with our team about where an SLM approach could make sense.

Talk to Our AI Expert