Swift + Metal MoE inference for Apple Silicon: Qwen 3.6 35B at 23.5–29.3 tok/s decode with 2.20× faster long-prompt prefill on a 24 GB M5; Gemma 4 26B in ~2 GB, DeepSeek-V4-Flash 284B, Inkling-Small 276B, native Mac app, CLI, and OpenAI-compatible server.
68stars2forksSwift
apple-silicondeepseekgemmallm-inferencemacosmetalmixture-of-expertsqwenswift
Real data pulled from GitHub this week. The author's original repo lives upstream.
View on GitHub