I'm running a 120B local LLM on 24GB of VRAM, and now it powers my smart home

local-llmmixture-of-expertsself-hostinghome-assistantgpuquantization

Abstraction: Running gpt-oss-120b locally on 24GB VRAM via MoE

Key points:

Connections: Mixture Of Experts · GPT Oss · Llama Cpp · Home Assistant · Quantization

Source: https://www.xda-developers.com/running-120b-local-llm-24gb-vram-smart-home/