knowledge📚

RAG Knowledge Base Stack

Private ChatGPT for your docs — deploys in minutes

Complete RAG pipeline: document ingestion (PDF, MD, HTML, Notion, Confluence) → embedding (local or API) → vector search (Qdrant) → chat UI (Open WebUI) → API. Private, no data leaves your server. Pre-configured for internal knowledge, support docs, onboarding.

Deploy Time5-8 minutes (scripted)
RequirementsVPS with 8GB+ RAM (16GB recommended), 50GB+ disk
Compliance✓ Paddle Compliant

What's Included

✓Open WebUI (chat interface)
✓Qdrant vector DB
✓Ollama (local embeddings) or OpenAI/Cohere API
✓Document ingestion pipeline (watch folder + API)
✓Multi-tenant collections (per team/project)
✓Citations on every answer
✓Role-based access control
✓Usage analytics + cost tracking

Choose Your Tier

One-time payment. Lifetime access. Optional maintenance.

Starter

99€one-time

Deployment: Self-serve

Ideal for getting started

  • Deploy script (Open WebUI + Qdrant + Ollama)
  • Ingestion scripts (folder watch + CLI)
  • Default embedding model (nomic-embed-text)
  • Basic RBAC
Choose Starter

Pro

299€one-time

Deployment: Assisted

Support: 30 days included

Ideal for getting started

  • Everything in Starter
  • Notion/Confluence/Drive connectors
  • Hybrid search (BM25 + vector)
  • Re-ranking (cross-encoder)
  • Custom prompt templates
Get Started

Enterprise

799€one-time

Deployment: Managed

Support: 90 days included

Ideal for getting started

  • Everything in Pro
  • Custom connector development
  • SSO/SAML/OIDC
  • Audit logging
  • SLA + priority support
Choose Enterprise

All tiers include: full source code, deploy scripts, documentation, GitHub repo access. No recurring fees unless you choose Sérénité maintenance (€149/mo).

Requirements & Deployment

Server Requirements

VPS with 8GB+ RAM (16GB recommended), 50GB+ disk

Deployment Process

  1. Purchase — Checkout via Paddle, instant access to private repo
  2. Provision — Create VPS (we recommend Contabo/Hetzner) or use existing
  3. Deploy — Run bash deploy.sh on your server (5-15 min)
  4. Configure — Add API keys via Web UI or .env file
  5. Run — Services start on boot. Access dashboards via HTTPS.

Frequently Asked Questions

Do these run on my server or yours?

100% on YOUR server (VPS, bare metal, cloud VM). You get root access, the deploy scripts, and full control. We never see your data, API keys, or logs.

What if I don't have a server?

We recommend Hetzner CX42 (8GB RAM, 160GB SSD, ~€32/mo) or similar. You provision it, give us root (for managed tiers) or run the script yourself (starter). You pay the provider directly.

Are there recurring fees to you?

No. One-time payment for the product. Optional: Sérénité maintenance at €149/mo (updates, health checks, backup audits, 1h support/mo). Cancel anytime.

How does deployment work?

Starter: you run `bash deploy.sh` on your server (5-15 min). Pro: we hop on a call, you share screen/SSH, we run it together. Enterprise: you give us temporary SSH, we deploy, harden, test, hand over with runbook.

What models do the agents use?

Default: NVIDIA NIM (free tier, 40 req/min). Optional: Ollama (local, free), OpenAI/Anthropic API (your keys, your bill), or any OpenAI-compatible endpoint. You choose at deploy time.

Is this Paddle-compliant?

Yes. Every product is a SOFTWARE PRODUCT (deployable stack, templates, scripts, config). You pay for infrastructure automation, not AI-generated output. No human services, no outbound marketing, no prohibited categories.

Can I customize after purchase?

Absolutely. You get the source: Docker Compose files, scripts, configs, workflow JSON, agent skills. Modify anything. Starter/Pro include docs for self-modification. Enterprise includes custom dev hours.

What happens after support period ends?

You keep everything forever. No license expiration. Renew Sérénité (€149/mo) for continued updates/health checks, or manage yourself. We provide upgrade guides for major versions.