Models are pluggable and platform-aware: roteiro model list recommends a
pick per section tuned to your machine (Apple-silicon Metal builds on macOS,
standard GGUF elsewhere) — treat it as the source of truth. All local models are
GGUF, run through the shared llama.cpp engine. A rough guide:
infer / duplicates| Model | Size | Dim | Best for |
|---|---|---|---|
bge-small-en-v1.5-gguf | ~65 MB | 384 | Any laptop — the recommended default. |
bge-base-en-v1.5 | ~200 MB | 768 | Stronger recall on a moderate machine. |
bge-large-en-v1.5 | ~640 MB | 1024 | Best quality, for a workstation. |
spec draft / serving| Model | Size | Role | Best for |
|---|---|---|---|
qwen3-0.6b | ~380 MB | instruct | Tiny, offline — the spec draft default. |
qwen3-8b | ~4.8 GB | instruct | Stronger drafting on a ~16 GB machine. |
qwen3-32b | ~18 GB | instruct | Best drafting, for a workstation. |
qwen2.5-coder-3b | ~1.8 GB | coding | Code completion & code Q&A. |
deepseek-r1-distill-qwen-1.5b | ~1.0 GB | reasoning | Small chain-of-thought reasoning. |
| Model | Size | Best for |
|---|---|---|
ocrs-text | ~12 MB | Pure-Rust OCR — literal text in screenshots (image-ocr). |
smolvlm-500m-gguf | ~520 MB | Describing diagrams & photos (image-vision). |