Model Quantization
concepts · 3 notes linked
Related: Meta · Llama · Large Language Models · Stanford · Georgi Gerganov · Open Source AI · Instruction Fine Tuning · GPT Oss
Notes
- I'm running a 120B local LLM on 24GB of VRAM, and now it powers my smart home — Running gpt-oss-120b locally on 24GB VRAM via MoE
- The Trick to Make LLaMa Fit into Your Pocket: Meet OmniQuant, an AI Method that Bridges the Efficiency and Performance of LLMs — OmniQuant learnable post-training quantization achieves low-bit LLM compression efficiently
- Why LLaMa Is A Big Deal — LLaMA enables GPT-3-class LLM inference on consumer hardware