Trending Repos
6 repos · Claude · AI · Dev · AWS · DevOps · Trading
- swellweb/reameNew↗
A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.
★ 93C++Created Jul 6· Updated 28d ago·Free Alternatives - poisonxa16/pxq_llamaNew↗
PXQ: PXA-native low-bit MoE quants (2/3/4-bit, E16-row scales) + fused CUDA kernels for Pascal/Volta — run a real 35B on a salvaged 12-16GB card. Fork of ik_llama.cpp.
★ 19C++Created Jul 17· Updated 24d ago·Free Alternatives - poisonxa16/pxq_llama.cppNew↗
PXQ: PXA-native low-bit MoE quants (2/3/4-bit, E16-row scales) + fused CUDA kernels for Pascal/Volta — run a real 35B on a salvaged 12-16GB card. Fork of ik_llama.cpp.
★ 9C++Created Jul 17· Updated 18d ago·Free Alternatives - mozilla-ai/llamafile↗
Distribute and run LLMs with a single file.
★ 25.6kC++Created Sep 2023· Updated 12d ago·Free Alternatives - endee-io/endee↗
Endee.io – A high-performance vector database, designed to handle up to 1B vectors on a single node, delivering significant performance gains through optimized indexing and execution. Also available in cloud https://endee.io/
★ 1.3kC++Created Jan 24· Updated 1mo ago·Free Alternatives - 1ay1/agentty↗
AI pair programming in your terminal — one static binary, sub-ms startup, any model
★ 623C++Created Apr 17· Updated 1d ago·Free Alternatives