Features Pricing Docs About Get started →
Now in public beta — free to self-host

The LLM gateway
your stack actually needs

One OpenAI-compatible endpoint. Every model. Local GPU or cloud — your call. Semantic memory, intelligent routing, and privacy that doesn't require a lawyer to understand.

$ helm install dream-weaver oci://ghcr.io/poser8/dream-weaver --set global.apiKey=dw-live-...
15+
model backends
<5ms
routing overhead
100%
OpenAI compatible
$0
to self-host

ZERO MIGRATION COST

Change one URL.
Get everything.

Dream-Weaver speaks OpenAI. Your existing code — Python SDK, LangChain, LlamaIndex, LiteLLM, Instructor — works without modification. Change the base URL, get intelligent routing, persistent memory, and privacy controls.

qwen3:8b codestral:22b claude-sonnet claude-opus gpt-5.4 gemini-2.5 deepseek-r1
app.py — 1 line changed
# Before — paying OpenAI for everything from openai import OpenAI client = OpenAI( api_key="sk-...", base_url="https://api.openai.com/v1" ) # After — local GPU for 80% of calls, cloud for the rest client = OpenAI( api_key="dw-live-...", base_url="https://api.dream-weaver.ai/v1" # ← only this ) # Same call. Now routes by intent: resp = client.chat.completions.create( model="dream-weaver/auto", messages=[{"role": "user", "content": msg}] )
INTELLIGENT ROUTING

The right model. Every time. Automatically.

A lightweight MoE classifier reads every request and dispatches to the optimal backend in <5ms. Free yourself from manually picking models.

routing log — dream-weaver/auto
→ "Write a merge sort in Rust" classifier: code (0.94) dispatch: 3ms backend: codestral:22b (local GPU) ✓ cost: $0.00 → "Explain the implications of Gödel's theorems" classifier: reasoning (0.89) dispatch: 2ms backend: qwen3:8b (local GPU) ✓ cost: $0.00 → "Review this 90k-token codebase for security issues" classifier: long-context (0.96) dispatch: 4ms backend: claude-opus-5 (cloud) ✓ cost: $0.18 → "Generate 500 unit tests for this module" classifier: code (0.91) dispatch: 3ms backend: codestral:22b (local GPU) ✓ cost: $0.00 ─── Session summary ───────────────────────────── Requests: 4 | Local: 3 (75%) | Cloud: 1 (25%) Saved vs all-cloud: ~$2.40 (this session)
PLATFORM

Everything in the stack.

ROUTING

MoE classifier dispatch

Lightweight intent classifier routes by content type — code, reasoning, long-context, creative, fast-lookup. <5ms overhead. Fully auditable routing log.

MEMORY

Persistent semantic memory

pgvector + mxbai-embed-large (1024 dim). Every conversation, file, and decision stored and searchable — on your hardware, not in a vendor cloud.

REASONING

LSR symbolic reasoning

Optional neuro-symbolic layer. Prolog-style rules + inductive logic programming augment LLM responses for constraint-sensitive queries.

PRIVACY

Local-first by design

Self-hosted in your k8s cluster. Sensitive traffic never leaves your network. Policy-gated cloud routing — you decide what goes where.

COMPATIBILITY

Full OpenAI API surface

Chat completions, embeddings, streaming SSE, function calling, JSON mode, vision. Any SDK that speaks OpenAI works immediately.

TOOLING

MCP + API + audit log

Native MCP tool server for Claude Desktop and IDE plugins. REST API for everything. Full Postgres audit log — cost, latency, backend, per request.

USE CASES

Built for people who run things at scale.

AI DEVELOPERS

Stop paying for tokens you could run locally

Route cheap tasks (codegen, summarization, classification) to your local GPU. Reserve expensive frontier models for the 20% of requests that actually need them. Most teams cut LLM spend 60–80%.

ENTERPRISE / REGULATED

OpenAI UX. Air-gap compliance.

Healthcare, finance, legal — where data residency isn't optional. Deploy in your VPC. Developers get the same SDK experience. InfoSec keeps their guarantees.

HPC & RESEARCH

Turn idle GPU time into internal LLM API

Running MI300X or H100 clusters? When jobs aren't running, your GPUs are wasted. Dream-Weaver turns that idle compute into a company-wide inference API. From the people who built and ran those clusters.

AGENT BUILDERS

Memory + reasoning without the plumbing

Persistent semantic memory, LSR working memory, MCP tools — all wired up. Your agents remember across sessions without you writing the infrastructure. Focus on the agent logic, not the backend.

PRICING

Free to run. Pro when you need it.

Self-host for free forever. Upgrade for auto-routing, semantic memory, and support.

OPEN CORE
$0
free forever · self-hosted
Full OpenAI-compatible gateway with local GPU inference. Community support. No credit card.
  • ✓OpenAI-compatible API
  • ✓Unlimited local inference (Ollama)
  • ✓Explicit model routing
  • ✓Streaming SSE + function calling
  • ✓REST API
  • –Auto-routing (MoE)
  • –Semantic memory
  • –LSR reasoning pod
ENTERPRISE
Custom
volume · white-glove onboarding
Air-gapped deployment, SSO, SLA, dedicated engineering support, and custom integrations.
  • ✓Everything in Pro
  • ✓Air-gapped / VPC deploy
  • ✓SSO / SAML / LDAP
  • ✓Dedicated routing cluster
  • ✓Custom model integrations
  • ✓SLA + uptime guarantee
  • ✓Onboarding engineering
  • ✓Dedicated Slack channel

Cloud API calls (Anthropic, OpenAI, Gemini) pass through at cost — no markup. You pay your providers directly.

GET STARTED

Ship it in an afternoon.

Helm chart. 15 minutes from zero to running. Your team routes smarter by end of day.

Start free → Read the docs
✓ No credit card for free tier ✓ Helm chart deploy ✓ Cancel Pro anytime