Spokes.wiki Search About
Collection source ↗ source url updated Sun Jun 21 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

LLMs-local (curated list)

A community awesome-list GitHub repo (maintainer 0xSojalSec) that catalogs the tooling for running LLMs on your own hardware — platforms, inference engines, models, builder tooling, hardware, tutorials, and communities. Reached the wiki via a promotional tweet from Dan Kornas (@DanKornas, 2026-06-18) that frames it as “the map” for moving from “which local stack should I try?” to a usable shortlist by grouping resources by the job they do.

It is filed here as the wiki’s first map of the on-device regime’s ecosystem — the named tools that fill out the local-llm-stack that llama-cpp and quantization describe at the mechanism level.

What it catalogs

The list organizes the local-LLM world by task rather than by vendor:

  • Inference stack, split two ways. It deliberately separates platforms — turnkey apps like LM Studio, Jan, LocalAI — from engines — the runtimes underneath: Ollama, llama.cpp, vLLM, SGLang, MLX. This platform/engine split is the substantive idea the wiki takes from it (see local-llm-stack).
  • UIs. Local chat / front-end tools: Open WebUI, Lobe Chat, text-generation-webui, SillyTavern, Page Assist.
  • Model navigation. Explorers, benchmark links, providers, and per-use-case model sections (general, coding, multimodal, image, audio).
  • Builder tooling. Agent frameworks, MCP, RAG, coding agents, browser automation, memory, testing, evaluation, observability.
  • Learning path. Tutorials grouped by models, prompt engineering, context engineering, inference, agents, RAG, and misc.

Why it’s filed, and its weight

Tier T4 — a single-maintainer curated link aggregation surfaced by a marketing tweet, with no benchmarks, no first-party measurement, and no analysis beyond the grouping itself. It does not move the spoke’s standing open question (the which-lever-bought-what benchmark decomposition); its value is coverage of the edge ecosystem, which the spoke had only as two engine pages (vllm, llama-cpp) and a synthesis claim that llama.cpp is “the core of Ollama/LM Studio.” The list names the layer above those engines for the first time. Several of its sections (agents, MCP, RAG, coding agents, evaluation) are out of this spoke’s scope — they belong to agentic-tooling-wiki and llm-providers-wiki — and are recorded here only as context, not paged. Maintainer 0xSojalSec was deferred here as a thin creator and has since been paged in speech-audio-wiki as 0xsojalsec, on recurrence: the same curator produces the same artifact for audio (free-voice-clone-list — 53 open speech/music models, same table-plus- detail construction, same absence of any measurement). Dan Kornas noted inline rather than as entity nodes, matching the how-does-vllm-work precedent (low graph signal for an inference-mechanics spoke). Being a live list, its contents are volatile; claims here trace to the snapshot read 2026-06-21.

local-llm-stack · llama-cpp · vllm · quantization · llm-inference