Inference Engine
A custom high-performance C++ inference engine that runs a 30M-parameter decoder-only transformer, trained on TinyStories — tiled matmul, NEON/AVX kernels, OpenMP, and KV caching, no framework in between.
0x00Overview
What this project is and how it's put together.
This project implements a complete inference pipeline for decoder-only transformer models. The engine features optimized matrix multiplication with SIMD instructions, a GPT-2 compatible Byte Pair Encoding tokenizer, and a modular architecture that separates tensor operations, model components, and tokenization.
It's built to run one specific model well: a ~30M parameter transformer trained on the TinyStories dataset, small enough to serve entirely from CPU cache and SIMD registers rather than a GPU.
0x01Quick start
macOS, Apple Silicon.
# Step 0: Install OpenMP dependency
brew install libomp
# Step 1: Build the project
make
# Step 2: Run inference
./a.out
gpt2_vocab.json and merges.txt aren't already in the project root — see Building.0x02Platform compatibility
This project is optimized exclusively for macOS (Apple Silicon), utilizing the ARM NEON library for SIMD instructions. It is a CPU-only implementation — there is no GPU code path.
0x03Performance comparison
Mean of 5 independent runs, reproducible via benchmark.sh.
The engine was benchmarked against naive PyTorch implementations running on both Apple Silicon (MPS) and CPU.
| Metric | PyTorch (MPS) | PyTorch (CPU) | Custom C++ Engine | Custom C++ (BLAS) | Speedup vs MPS |
|---|---|---|---|---|---|
| TTFT | 910.29 ms | 13.53 ms | 14.08 ms | 8.32 ms | ≈ 109.4x |
| Avg time / token | 87.00 ms | 12.20 ms | 8.11 ms | 2.13 ms | ≈ 40.8x |
| Throughput | 11.49 tok/s | 81.96 tok/s | 120.94 tok/s | 435.29 tok/s | ≈ 37.9x |
A benchmark.sh script is provided in the root directory to reproduce these metrics and perform comparative analysis.
0x04Optimizations
Four techniques account for most of the speedup above.
1. SIMD kernels (NEON / AVX)
Utilizes hardware-level vectorization to process multiple data points in a single instruction. Implemented in operator+, operator*, LayerNorm, operator/, softmax, add_bias.
2. Tiled matrix multiplication
Implements a cache-efficient 32×32 tiling strategy to minimize cache misses and maximize memory bandwidth utilization. Implemented in operator*.
3. OpenMP parallelism
Distributes independent compute-heavy loops across multiple CPU cores, using adaptive thresholds to avoid threading overhead on small tensors. Implemented in operator* (matmul), operator+ (addition), softmax, add_bias, gelu, and multiheadattention.
4. KV caching
Implements key-value caching for O(N) incremental decoding, preventing the redundant re-computation of previous tokens during the generation phase. Implemented in attention, using the KVCache structure.
0x05Features
| Feature | Description |
|---|---|
| Multi-Head Attention | Full implementation with causal masking |
| Layer Normalization | Pre-normalization architecture support |
| BPE Tokenizer | GPT-2 compatible byte-level BPE tokenization |
| Positional Encodings | Sinusoidal positional embeddings |
| Temperature Sampling | Configurable sampling with temperature scaling |
0x06Tensor operations
The Tensor class.
| Operation | Method | Description |
|---|---|---|
| Matrix multiplication | operator* | Tiled matmul with 32×32 blocks |
| Addition | operator+ | Element-wise addition |
| Division | operator/ | Scalar division |
| Transpose | t() | Matrix transpose |
| Softmax | softmax() | Row-wise softmax with numerical stability |
| Layer normalization | LayerNorm() | Normalization with learnable parameters |
| Causal mask | mask() | Lower triangular mask for attention |
| Concatenation | concat_horizontal() | Horizontal tensor concatenation |
| Concatenation | concat_vertical() | Vertical concatenation, used for the KV cache |
| Bias addition | add_bias() | Add bias vector to each row |
0x07Weight format
Weights are stored in NumPy .npy format and loaded at runtime. The naming convention follows:
transforms.{block}.{component}.{head}.{weight|bias}.npy
| Component | Description |
|---|---|
q | Query projection |
k | Key projection |
v | Value projection |
join | Output projection (after attention) |
ffn.0 | FFN intermediate layer |
ffn.3 | FFN output layer |
norm1 | Pre-attention layer norm |
norm2 | Pre-FFN layer norm |
0x08Building
Prerequisites
- C++17 compatible compiler (
g++orclang++) - Python 3.x with the
transformerslibrary, for tokenizer export
Compile
make
The Makefile uses the following optimization flags:
| Flag | Purpose |
|---|---|
-std=c++17 | C++17 standard |
-O3 | Maximum optimization level |
-march=native | Enable CPU-specific SIMD instructions |
-ffast-math | Aggressive floating-point optimizations |
-funroll-loops | Loop unrolling for performance |
Export tokenizer
Before running inference, export the GPT-2 tokenizer vocabulary:
python tokenizer.py
This creates gpt2_vocab.json and merges.txt in the project root.
0x09Usage
Running inference
./a.out
The Runner class initializes the model with weights from the weights/ directory and performs autoregressive generation.
Customizing generation
Edit the run() function in runner.hpp to modify:
| Parameter | Default | Description |
|---|---|---|
prompt | "Hello, how are you?" | Input text for generation |
max_new_tokens | 5 | Maximum tokens to generate |
temperature | 0.8 | Sampling temperature (0 = greedy) |
seq_len | 128 | Maximum context length |
Example output
Model config: vocab=50257, d_model=256, heads=8, blocks=6
Loaded vocab: 50257 tokens
Loaded merges: 50000 BPE merge rules
Model initialized!
Tokenized: 6 tokens
Generating...
after embeddings: 6x256
after block 0: 6x256
after block 1: 6x256
after block 2: 6x256
after block 3: 6x256
after block 4: 6x256
after block 5: 6x256
logits: 6x50257
=== Generated Text ===
Hello, how are you? Once upon a time...
======================
0x0AProject structure
0x0BModel architecture
The engine supports a decoder-only transformer with the following configuration:
| Parameter | Value |
|---|---|
| Parameters | ~30M |
| d_model | 256 |
| Number of heads | 8 |
| Number of blocks | 6 |
| d_k (per head) | 32 |
| Vocabulary size | 50257 |
| Max sequence length | 128 |
| Training dataset | TinyStories |
Performance considerations
| Optimization | Impact |
|---|---|
| Tiled matrix multiplication | Reduces cache misses |
SIMD via -march=native | 2–4x speedup on modern CPUs |
-ffast-math | Enables vectorization of floating-point ops |
-funroll-loops | Reduces loop overhead |
| Const references | Avoids unnecessary copies |
| Reserve for vectors | Prevents reallocations |
0x0CError handling
- Shape mismatch detection for all tensor operations
- Try-catch blocks around attention and forward passes
- Validation of weight dimensions during initialization
- Graceful handling of missing tokenizer files
- Detailed error messages with tensor shapes
0x0DLimitations
| Limitation | Notes |
|---|---|
| Single GPU | CPU-only implementation |
| Batch size | Supports batch size of 1 |
| Precision | FP32 only (no FP16 / INT8 quantization) |
| Context | Fixed maximum context length |
0x0EDependencies
| Library | Purpose | Source |
|---|---|---|
npy.hpp | NumPy file parsing | GitHub (external) |
json.hpp | JSON parsing | GitHub (external) |
transformers | Tokenizer export | HuggingFace (Python) |
libomp | OpenMP support for parallelization | Homebrew (macOS) |
0x0FLicense
This project is for educational and research purposes.
0x10Acknowledgments
- TinyStories dataset for model training
- HuggingFace transformers for tokenizer reference
- GPT-2 architecture as the model foundation