LLMs-local — awesome platforms, tools, and resources for running LLMs locally
local-llmsawesome-listinference-enginesopen-weightsself-hosting
Abstraction: Awesome list for running LLMs locally
Key points:
- Curated GitHub awesome-list of platforms, tools, and resources for running LLMs locally/offline, organized into inference platforms, engines, UIs, models, tools, hardware, tutorials, communities.
- Inference platforms: LM Studio, Jan, LocalAI, ChatBox, lemonade. Engines: ollama, llama.cpp, vLLM, exo, BitNet (1-bit), SGLang, koboldcpp, mlx-lm (Apple silicon), distributed-llama. UIs: Open WebUI, Lobe Chat, Text generation web UI, SillyTavern.
- Open-weight model roster (dated into 2026): DeepSeek-V4, MiniMax M3 (1M context), Kimi-K2.6, Qwen3-Next, Gemma 3, gpt-oss (OpenAI open weights), Ministral 3, GLM-4.5, Granite 4.0, plus coding (Qwen3-Coder, Devstral 2), vision (Qwen3-VL), and audio (Voxtral, chatterbox) models.
- Tooling categories: agent frameworks (LangChain, AutoGen, CrewAI, llama_index), MCP servers, RAG (GraphRAG, LightRAG, Haystack), memory (mem0, letta), observability (Langfuse, garak), and AI coding agents (aider, cline, OpenHands, continue).
- Includes hardware/VRAM calculators (Kolosal, ZLUDA for CUDA-on-non-NVIDIA) and prompt/context-engineering learning resources (Anthropic, Google, DAIR.ai guides, karpathy/nanochat).
Connections: Ollama · Llama Cpp · Vllm · Lm Studio · Github · Local LLM Inference · Open Weight Models · AI Agents