GLM-5.3-Flash (abliterated EXL3) on 2x NVIDIA DGX Spark with the TensorFold engine: 1.8x faster decode than vLLM, 4x256k concurrent threads, byte-exact speculative decoding. Work in progress.
104stars10forksPython
Real data pulled from GitHub this week. The author's original repo lives upstream.
View on GitHub