Budget VPS for Ollama & Vector Search
🏆 Our pick: Hetzner Cloud · 💰 Best value (dedicated): OVHcloud Eco (Kimsufi) · 🥈 Runner-up: Contabo
⚠ Class honesty: the editorial pick above is a fair-share (shared vCPU) plan — its specs are burst ceilings (“up to”), not guarantees. The dedicated value ranking below remains the primary guide for sustained workloads. Dedicated step-up for this workload: OVHcloud Eco (Kimsufi) at $11.65/mo (guaranteed cores).
Run quantized LLMs (7B Q4) and pgvector/Qdrant on 8–16GB RAM VPS with fast NVMe. GPU not mandatory for RAG demos.
Prices & hardware specs verified 2026-09-02
Start on dedicated — guaranteed cores for this workload (Top 5)
Our Recommendations
CX33 (4vCPU/8GB/80GB NVMe) €8.99/mo ($10.48 equivalent); scale to CX43 (8c/16GB) or CCX23 (4c/16GB dedicated) for 16GB+ — green EU.
8–16GB VPS for €8–€16 incl. VAT ($7.84–$15.68 VAT-stripped) with unmetered — cost-effective for large contexts.
GPU marketplace when you outgrow CPU inference.
Key Takeaways
- Ollama 7B Q4 needs ~5GB + OS — 8GB floor.
- Store embeddings on NVMe — SATA tanks HNSW latency.
Gotchas
- No GPU on pure VPS — latency ~5x vs A10.
- Contabo CPU steal shows under sustained inference — monitor.
Test your own workload
Adjust vCPU, RAM and bandwidth in the calculator.