LLM cost optimization

Cut your AI bill
without cutting quality.

Run your AI workloads through varsten to reduce your API bill, and stop worrying about wether you're getting the most bang for your buck.

# Tokens processed
600k +
Overall quality degradation
0 - 1%
Avg cost reduction
> 10%
Fig 01 — Request patht = 42ms
Client applicationPOST /v1/chat/completionsmodel: gpt-4oApromptVARSTEN · PROXY NODEBroutecachetrimcompressoptimized · routedLLM provideropenai · anthropicgeminiCresponse · lower cost
A → clientB → proxyC → provider
Section 02 · Cost savers

Six cost saving mechanisms.

Traffic is sent through routing, cache, trim, downshift, and prompt compression mechanisms. Batching is a separate async workflow for non-urgent jobs. Each lever is auditable and togglable.

01●

Smart Routing

Sends each request to the most cost-effective AI model that can do the job well, deciding instantly for every prompt.

predicate → model tier
02●

Semantic Cache

Saves past AI answers and reuses them whenever a new request means the same thing, even if it's worded differently.

pgvector · TTL bound
03●

Token Trim

Cleans up prompts in real time by stripping out repeated text, extra whitespace, and unnecessary history before sending.

structural · policy gated
04●

Prompt Compression

Rewrites long system instructions into permanently shorter versions to cut token costs, requiring human review before going live.

eval/replay · hash-matched substitution
05●

Model Downshift

Uses test history to safely move routine tasks to cheaper models, testing changes on a small scale first with instant rollback.

eval-driven · canary window
06●

Batching

Groups non-urgent requests together in the background to process them at heavily discounted batch rates.

async API · off-path
Section 03 · Setup

Three ways to connect.

Three ways to connect. Each provides a different level of security and control. The SDK wrapper is recommended for production traffic, the base URL is fastest for evaluation, and Direct Monitoring keeps sensitive workloads out of the request path.

Pro · Recommended

Production SDK

The OpenAI wrapper sends healthy traffic through Varsten and falls back direct-to-provider on Varsten-origin failures before provider output starts. Your provider key stays local for fallback.

  • →Direct provider fallback
  • →Per-request metadata supported
  • →Optimizations run where configured
example.tsv0.1.0
import { VarstenOpenAI, VarstenTrace } from "@varsten/openai";

const client = new VarstenOpenAI({
  varstenApiKey: process.env.VARSTEN_API_KEY,
  openaiApiKey: process.env.OPENAI_API_KEY,
  onFallback: (event) => {
    console.warn("varsten fallback", event.reasonCode);
  },
});

const trace = new VarstenTrace();

await client.chat.completions.create(
  {
    model: "gpt-4o-mini",
    messages,
  },
  {
    varsten: trace.metadata({
      feature: "support_agent",
      taskType: "classification.intent",
      customerId: "cust_123",
    }),
  },
);
Section 04 · Pricing

Pay from savings, or audit for free.

We charge a fee on verified savings, so you only pay when we save you money. Plans are billed monthly. No commitment, cancel anytime. If you just want to audit your traffic, you can do that for free.

Fee < Savings · always
Plan · 01Free Forever

Base

Freeno credit card

Connect via Quick Eval or Direct Monitoring to audit your live traffic and map out estimated savings, with no behavior-changing optimization applied.

  • ✓Monitor AI spend
  • ✓100k requests/month
  • ✓Savings recommendations
  • ✓Quick Eval or Metadata
  • ✓No credit card required
Start a free audit
Plan · 02Billed Monthly

Pro

25%of verified savings

Unlocks the optimization engine: inline routing, cache, trim, compression, and downshift, plus async batching for eligible jobs. Pricing is capped at 25% of verified savings.

  • ✓Everything in free +
  • ✓Automated cost savings
  • ✓Unlimited requests/month
  • ✓Production-safe SDK integration
  • ✓Controls, guardrails, rollback
Request early access
Plan · 03Custom Pricing

Enterprise

For custom pricing that doesn't scale with your bill,
we negotiate a rate and fee cap.