Cardix

Architecture

Inside your perimeter

Cardix runs as a single service within your network. Applications connect via an OpenAI-compatible endpoint; Cardix handles classification, policy, selection, and dispatch.

YOUR PERIMETER · VPC / on-prem / air-gappedYour applicationsunchanged OpenAI-compatible SDKCardixclassify · policy · select · dispatchRouting · Learning · Observabilityapproved outbound traffic onlyPrivate inferenceon-prem + private cloudApproved external APIsrestricted by your policy

Multi-axis

Optimization axes

Cost, latency, quality, compliance, personalization, and custom metrics you define

Hybrid

Deployment modes

On-prem, VPC, and air-gapped

1

OpenAI-compatible endpoint

Drop-in for your existing SDK

Fail-closed

Policy enforcement

Residency rules before routing

Planes

Three layers, one purpose

Routing, learning, and observability work together without exposing your data or blocking your responses.

Routing plane

Classify sensitivity and task, apply policy, then select the best eligible model across every dimension that matters, including cost, latency, quality, and fit, and dispatch synchronously on every request.

Learning plane

Captures traffic signals and the custom metrics you define, then builds per-team and per-workload model-fit intelligence to improve selection over time. Runs asynchronously and never blocks the response.

Observability plane

Full audit trail, spend vs baseline comparison, and visibility into blocked demand.

Request journey

What happens on every request

Five steps, milliseconds. Policy before routing, every time.

  1. 01

    Authenticate

    Validate the caller and attribute the request to your organization.

  2. 02

    Classify

    Determine data sensitivity and task category for the prompt.

  3. 03

    Policy

    Evaluate residency rules and produce allowed deployment locations. Fail closed if no rule matches.

  4. 04

    Select

    Weigh every dimension you care about, including cost, latency, quality, personalization, and custom metrics, across your eligible catalog, then choose the best model for this request.

  5. 05

    Dispatch

    Route to the chosen backend and return the response, streaming or buffered.

Learning & personalization

Routing that gets better with every request

Cardix does not stop at a static rule. It learns from your traffic and the custom metrics you define, then continuously improves model selection for your organization, all inside your perimeter and off the response path.

01

Observe

Every routing decision records its signals: the chosen model, cost and latency actuals, quality outcomes, and the custom metrics your organization defines.

02

Learn

Cardix builds model-fit intelligence for your organization, learning which models perform best for each team, task type, and workload, entirely on your infrastructure.

03

Optimize

That intelligence refines how Cardix weighs each axis, so selection reflects what actually works for your traffic rather than a static, one-size-fits-all default.

04

Improve

The next matching query routes better. The loop runs continuously and asynchronously, off the hot path, so it never adds latency to a response.

Continuous, asynchronous, and off the hot path. The loop never adds latency to a response.

Signals captured

  • Task & sensitivity class

    The classification assigned to each request, so learning is grouped by the work being done.

  • Per-team & per-workload outcomes

    How models perform for different teams and workloads across your organization.

  • Cost & latency actuals

    The real cost and speed of each decision, measured on your infrastructure.

  • Quality outcomes

    Automated quality signals and any application or human feedback you choose to send.

  • Custom organization metrics

    Your own quality bars or business KPIs. You define what matters, and Cardix optimizes for it.

All signals stay inside your perimeter. The intelligence is built on your infrastructure. Your data never leaves your network, and nothing is sent back to us.

Integration

Three steps to intelligent routing

No rip-and-replace. Cardix integrates with the stack you already run.

01

Point your SDK

Change the base URL in your existing OpenAI-compatible client. No application rewrite, and streaming is supported.

02

Set your policy

Define residency rules, sensitivity labels, and approved backends. Cardix enforces them before any model is considered.

03

Route across your fleet

Every request is classified, optimized, and dispatched to the best eligible model across on-prem, private cloud, or approved APIs.

Get started

Deploy intelligent routing inside your perimeter

Review the architecture, assess fit for your stack, and speak with our team about your deployment requirements.