A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
2.7kstars450forksC
avx2c99cpu-inferencedeep-learningfrom-scratchinference-enginekimi-k3linear-attentionllmllm-inferencemachine-learningmemory-efficientmixture-of-expertsmoemxfp4quantizationsimdsystems-programmingtransformerzero-dependencies
Real data pulled from GitHub this week. The author's original repo lives upstream.
View on GitHub