Local LLM Inference
concepts · 3 notes linked
Related: Ollama · Llama Cpp · Apple · Apple Silicon · Mlx Framework · Model Inference Optimization · GPT Oss · Whisper
Notes
- I'm running a 120B local LLM on 24GB of VRAM, and now it powers my smart home — Running gpt-oss-120b locally on 24GB VRAM via MoE
- LLMs-local — awesome platforms, tools, and resources for running LLMs locally — Awesome list for running LLMs locally
- Ollama Just Got 93% Faster on Mac. Here's How to Enable It. — Ollama MLX backend delivers 93% decode speedup on Apple Silicon Macs