sauravsingla/MemVantaLow-memory C++20 LLM inference runtime for quantized GGUF models on CPU — mmap-backed weights, paged KV cache, Q4/Q8 kernels, and reproducible llama.cpp benchmarks.
Low-memory C++20 LLM inference runtime for quantized GGUF models on CPU — mmap-backed weights, paged KV cache, Q4/Q8 kernels, and reproducible llama.cpp benchmarks.
The only pull request from outside contributors was opened in the last 14 days, too recently to judge.
0 of 0outside PRs merged (0%)
checked
02What happened to outsiders
03Where newcomer work lands
04The evidence