Recommended local models

Models are pluggable and platform-aware: roteiro model list recommends a pick per section tuned to your machine (Apple-silicon Metal builds on macOS, standard GGUF elsewhere) — treat it as the source of truth. All local models are GGUF, run through the shared llama.cpp engine. A rough guide:

Embedding — for infer / duplicates

ModelSizeDimBest for
bge-small-en-v1.5-gguf~65 MB384Any laptop — the recommended default.
bge-base-en-v1.5~200 MB768Stronger recall on a moderate machine.
bge-large-en-v1.5~640 MB1024Best quality, for a workstation.

Generative — for spec draft / serving

ModelSizeRoleBest for
qwen3-0.6b~380 MBinstructTiny, offline — the spec draft default.
qwen3-8b~4.8 GBinstructStronger drafting on a ~16 GB machine.
qwen3-32b~18 GBinstructBest drafting, for a workstation.
qwen2.5-coder-3b~1.8 GBcodingCode completion & code Q&A.
deepseek-r1-distill-qwen-1.5b~1.0 GBreasoningSmall chain-of-thought reasoning.

Vision & OCR — for image ingestion

ModelSizeBest for
ocrs-text~12 MBPure-Rust OCR — literal text in screenshots (image-ocr).
smolvlm-500m-gguf~520 MBDescribing diagrams & photos (image-vision).