One OpenAI-compatible endpoint. Every model. Local GPU or cloud — your call. Semantic memory, intelligent routing, and privacy that doesn't require a lawyer to understand.
Dream-Weaver speaks OpenAI. Your existing code — Python SDK, LangChain, LlamaIndex, LiteLLM, Instructor — works without modification. Change the base URL, get intelligent routing, persistent memory, and privacy controls.
A lightweight MoE classifier reads every request and dispatches to the optimal backend in <5ms. Free yourself from manually picking models.
Lightweight intent classifier routes by content type — code, reasoning, long-context, creative, fast-lookup. <5ms overhead. Fully auditable routing log.
pgvector + mxbai-embed-large (1024 dim). Every conversation, file, and decision stored and searchable — on your hardware, not in a vendor cloud.
Optional neuro-symbolic layer. Prolog-style rules + inductive logic programming augment LLM responses for constraint-sensitive queries.
Self-hosted in your k8s cluster. Sensitive traffic never leaves your network. Policy-gated cloud routing — you decide what goes where.
Chat completions, embeddings, streaming SSE, function calling, JSON mode, vision. Any SDK that speaks OpenAI works immediately.
Native MCP tool server for Claude Desktop and IDE plugins. REST API for everything. Full Postgres audit log — cost, latency, backend, per request.
Route cheap tasks (codegen, summarization, classification) to your local GPU. Reserve expensive frontier models for the 20% of requests that actually need them. Most teams cut LLM spend 60–80%.
Healthcare, finance, legal — where data residency isn't optional. Deploy in your VPC. Developers get the same SDK experience. InfoSec keeps their guarantees.
Running MI300X or H100 clusters? When jobs aren't running, your GPUs are wasted. Dream-Weaver turns that idle compute into a company-wide inference API. From the people who built and ran those clusters.
Persistent semantic memory, LSR working memory, MCP tools — all wired up. Your agents remember across sessions without you writing the infrastructure. Focus on the agent logic, not the backend.
Self-host for free forever. Upgrade for auto-routing, semantic memory, and support.
Cloud API calls (Anthropic, OpenAI, Gemini) pass through at cost — no markup. You pay your providers directly.
Helm chart. 15 minutes from zero to running. Your team routes smarter by end of day.