← Selected builds

Relay

Relay makes private AI workflows observable, governable—and therefore improvable.

7curated operating routes
20case evaluation baseline
$0/hrmanaged GPU compute when off

A useful answer can still come from the wrong model, the wrong context, an uneconomic route, or a process nobody can reproduce. When those operating facts are scattered across providers and interfaces, downstream quality cannot improve systematically.

Non-obvious diagnosis

You cannot improve what the system cannot attribute. Model, context, route, latency, cost, outcome, authority, and recovery must remain connected to the work they produced.

The build starts here. The feature set follows from this diagnosis.

Each run should leave enough evidence to make the next run better.

A private, single-tenant operating environment that composes established gateway, provider, and GPU infrastructure into one governed workspace: curated routes, explicit context, durable receipts, bounded files and memory, reusable workflows, and on-demand dedicated inference.

017 operating routes

Curated working bench

A focused set of routes replaces model sprawl, making task fit, capability, latency, and economics deliberate choices.

02Attribution · provenance · cost

Durable run receipts

The served model, context, tokens, latency, cost, status, and outcome remain attached to the work they produced.

03Side-by-side route evaluation

Compare before standardizing

The same task can run across two routes with durable branches, evaluation evidence, and an explicit decision note.

04Context receipts · checkpoints

Bounded context and workflows

Files, memory, reusable steps, checkpoints, and ceilings are explicit, so successful work becomes reproducible rather than anecdotal.

05$0/hr managed GPU when off

Compute only when justified

The control plane stays available while dedicated GPU capacity starts on demand, with explicit lifecycle controls and zero managed compute cost when off.

Inspect the mechanism, then inspect the evidence.

Each request becomes evidence for the next decision. The system can compare task fit, quality, latency, context, and cost; refine the route or workflow; and make better performance repeatable instead of treating every response as an isolated event.

One run becomes the next decision

Observability is valuable only when it changes what happens downstream.

Relay keeps the operating facts attached to the result, so model choice, context design, workflow quality, and infrastructure economics can improve together.

  1. 01Observe

    Model · context · route · latency · tokens · cost · status

  2. 02Compare

    Task fit · output quality · reliability · unit economics

  3. 03Decide

    Keep or change the model, prompt, context, workflow, or compute

  4. 04Improve

    Make stronger performance repeatable in the systems above it

Better task fit

Evaluation replaces reputation-led model selection.

More repeatable work

Context and output receipts make successful runs explainable.

Safer expansion

Explicit authority and bounded context let more consequential work enter the system.

Stronger economics

Hosted, local, and dedicated routes can be judged against the value of the work.

No hidden fallback.

A failed or unavailable route remains visible instead of silently changing the model—and the character of the result.

Actual product screen

The activity ledger preserves the operating facts needed to compare, diagnose, and improve later decisions.

Actual product screen

A curated working bench turns a broad provider catalogue into explicit choices about capability, route, context, and economics.

A feature matters when it improves the next decision.

  1. 01Route

    Choose a model and compute path against the task, context, and economic constraint.

  2. 02Run

    Execute with bounded files, memory, workflow steps, and explicit fallbacks.

  3. 03Receipt

    Persist the model, context, latency, tokens, cost, status, and output evidence.

  4. 04Compare

    Evaluate quality, reliability, task fit, and unit economics across consistent work.

  5. 05Improve

    Update the route, prompt, context, workflow, or compute choice and make the gain repeatable.

Architecture follows the failure mode.

The implementation is the visible surface. These choices determined whether it could solve the underlying problem.

  1. 01

    Make every run attributable

    Bind the served model, included context, latency, tokens, observed cost, status, and output evidence to the work itself—not to a disconnected infrastructure log.

  2. 02

    Curate rather than maximize choice

    Expose a focused working bench and evaluate routes against consistent tasks. More available models are useful only when the system can determine which one belongs in the workflow.

  3. 03

    Separate control from expensive compute

    Keep the operating layer available while dedicated capacity starts only when the work justifies it, with explicit ceilings, termination controls, and an independently enforced backstop.

Evidence classLive operating system
Source

Authenticated captures from the running private workspace, operating documentation, lifecycle receipts, and versioned evaluation summaries.

What it establishes

A functioning multi-model workspace, persisted run ledger, curated routes, explicit unit economics, evaluation harness, and a verified path from hosted inference to on-demand dedicated compute.

Boundary

This is a private, single-owner operating environment. The evidence establishes working behavior and design discipline—not independent security certification, multi-tenant operation, or external production scale.

Next build · Facial-expression classificationThe strongest model was selected from the error pattern—not the prestige of its architecture.Open build ↗︎