Routing plane
Classify sensitivity and task, apply policy, then select the best eligible model across every dimension that matters, including cost, latency, quality, and fit, and dispatch synchronously on every request.
Architecture
Cardix runs as a single service within your network. Applications connect via an OpenAI-compatible endpoint; Cardix handles classification, policy, selection, and dispatch.
Multi-axis
Optimization axes
Cost, latency, quality, compliance, personalization, and custom metrics you define
Hybrid
Deployment modes
On-prem, VPC, and air-gapped
1
OpenAI-compatible endpoint
Drop-in for your existing SDK
Fail-closed
Policy enforcement
Residency rules before routing
Planes
Routing, learning, and observability work together without exposing your data or blocking your responses.
Classify sensitivity and task, apply policy, then select the best eligible model across every dimension that matters, including cost, latency, quality, and fit, and dispatch synchronously on every request.
Captures traffic signals and the custom metrics you define, then builds per-team and per-workload model-fit intelligence to improve selection over time. Runs asynchronously and never blocks the response.
Full audit trail, spend vs baseline comparison, and visibility into blocked demand.
Request journey
Five steps, milliseconds. Policy before routing, every time.
Validate the caller and attribute the request to your organization.
Determine data sensitivity and task category for the prompt.
Evaluate residency rules and produce allowed deployment locations. Fail closed if no rule matches.
Weigh every dimension you care about, including cost, latency, quality, personalization, and custom metrics, across your eligible catalog, then choose the best model for this request.
Route to the chosen backend and return the response, streaming or buffered.
Learning & personalization
Cardix does not stop at a static rule. It learns from your traffic and the custom metrics you define, then continuously improves model selection for your organization, all inside your perimeter and off the response path.
Every routing decision records its signals: the chosen model, cost and latency actuals, quality outcomes, and the custom metrics your organization defines.
Cardix builds model-fit intelligence for your organization, learning which models perform best for each team, task type, and workload, entirely on your infrastructure.
That intelligence refines how Cardix weighs each axis, so selection reflects what actually works for your traffic rather than a static, one-size-fits-all default.
The next matching query routes better. The loop runs continuously and asynchronously, off the hot path, so it never adds latency to a response.
Signals captured
Task & sensitivity class
The classification assigned to each request, so learning is grouped by the work being done.
Per-team & per-workload outcomes
How models perform for different teams and workloads across your organization.
Cost & latency actuals
The real cost and speed of each decision, measured on your infrastructure.
Quality outcomes
Automated quality signals and any application or human feedback you choose to send.
Custom organization metrics
Your own quality bars or business KPIs. You define what matters, and Cardix optimizes for it.
All signals stay inside your perimeter. The intelligence is built on your infrastructure. Your data never leaves your network, and nothing is sent back to us.
Integration
No rip-and-replace. Cardix integrates with the stack you already run.
01
Change the base URL in your existing OpenAI-compatible client. No application rewrite, and streaming is supported.
02
Define residency rules, sensitivity labels, and approved backends. Cardix enforces them before any model is considered.
03
Every request is classified, optimized, and dispatched to the best eligible model across on-prem, private cloud, or approved APIs.
Get started
Review the architecture, assess fit for your stack, and speak with our team about your deployment requirements.