Route AI prompts to the best model, best price.

Intelligent middleware that analyzes intent, applies semantic caching, and routes to the optimal LLM. Cut costs by up to 90%.

0% cost reduction0 AI models0 brands0% uptime
Routing pipeline · liveSaved this session $0.0000
Prompt Composer “…” cache · idle Segmenter Groq / DeepSeekcode & speed lane OpenAI / Anthropicdeep reasoning lane Google / xAIgeneral & live lane Quality Optimizer — / 100 reassemble · score · learn warming up…

1 · Prompt Composer

Normalizes the raw prompt — applies your presets, redacts PII, and checks the semantic cache before a single token is spent. Click any stage to pin its details.

Routing across 53 models from 18 brands

GoogleOpenAIAnthropicDeepSeekxAIGroqMistralQwenMetaPerplexityCohereMiniMaxAmazonMoonshotAINVIDIAByteDanceSakana AIZ.ai

Everything you need

A complete middleware solution for intelligent AI routing.

Intelligent Routing

AI-powered intent analysis routes prompts to the optimal model based on complexity, speed, and cost requirements.

Cost Arbitrage

Save up to 90% on API costs by automatically selecting the most cost-effective provider for each request.

Semantic Caching

85% similarity matching eliminates redundant API calls. Zero-cost responses for semantically identical queries.

Security First

Built-in PII redaction, threat detection, and comprehensive audit logging for enterprise compliance.

Cascading Fallbacks

Automatic failover across 53 models from 18 brands ensures 99.95% uptime for your applications.

Latency Optimization

Choose between instant inference, fast response, or deep reasoning based on your application needs.

How it works

Every request passes through a 3-tier optimization pipeline.

1

Semantic Cache Check

85% similarity matching against local buffer. Semantically identical queries return instantly at zero cost.

2

Intent Analysis

AI classifies your prompt by coding, reasoning, creative, and speed requirements to select the optimal model.

3

Smart Dispatch

Request routes to the best provider with automatic cascading fallback if primary endpoints fail.

Ready to optimize your AI costs?

Join developers and teams already saving up to 90% on LLM inference costs with intelligent routing.

Features

Everything the routing engine does on every request, plus the platform tooling built around it — wallets, budgets, workspaces, and analytics.

The routing engine

INTENTcodingrouterGroq / DeepSeekOpenAI / AnthropicGoogle / xAI

Intelligent Routing

Every prompt is classified by intent — coding, reasoning, creative, or speed-sensitive — then dispatched to the model that best fits that intent at the lowest defensible cost.

Flagship baseline$0.0000Routed by I amToken$0.0000

Cost Arbitrage (Pay-per-Saving)

We don't mark up inference. We charge a percentage of the difference between what a premium flagship model would have cost and what you actually paid through our routing.

req 1 of 4“Explain binary search”semantic cachemodel

Semantic Caching

85% similarity matching against a local buffer returns semantically identical queries instantly, at zero additional cost.

OpenAI · primaryAnthropic · fallback 1Google · fallback 2
answered anyway · 99.95% uptime

Cascading Fallbacks

If a primary provider errors or times out, the request automatically falls back across the routing pool instead of failing.

routed modelflagship escalationquality gate ≥ 85registry

Quality Layer

A quality pass scores every response against the original intent. If it misses the bar, the optimization engine escalates to a higher model and re-runs — then records actual vs predicted quality back into the routing registry.

INCOMING PROMPTMigrate our checkout API to gRPCby Q3, keep the bill under $2k.We assume traffic stays flat.Plan the rollout in stages — whichendpoints break? REST is too slow;p95 latency is 800 ms.Intent Map7 semantic dimensions · 0 ms model timegoal · migrate checkout to gRPCconstraints · ≤ $2k / monthassumptions · traffic stays flatsteps · staged rolloutquestions · breaking endpoints?claims · REST is too slowevidence · p95 = 800 msintent map → router
reading the prompt…

Logic Layer

IAM-Token extracts the goal, constraints, assumptions, steps, questions, claims, and evidence — a machine-readable map of intent before any token is spent.

Built around it

Managed or BYOK keys

Use our managed keys or bring your own OpenAI / Anthropic / Google keys. Pay-per-Saving still applies at 20% Managed or 50% BYOK.

Wallet, budgets & kill-switches

Prepaid wallet with reserve/settle/refund, auto-recharge, alerts, and a hard kill-switch before a request executes.

Webhooks

HMAC-signed outbound webhooks for usage and billing events.

Workspaces & departments

Seats, roles, shared prompts, per-department cost tracking. SSO planned — not live yet.

Lifecycle analytics

Replay any request: segmentation, intent, router decision, execution, quality scoring.

Security & auth

TOTP MFA, encrypted keys at rest, admin audit logging.

Providers

53 models across 18 brands, one routing engine. We don't lock you into a single vendor's roadmap — the router picks whichever model is actually the right fit, and falls back if a provider errors.

OpenAI

GPT-4o, GPT-4 Turbo, o4-mini, GPT-3.5 Turbo

Anthropic

Claude Opus 4, Claude Sonnet 4, Claude 3.5 Haiku

Google

Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 2.0 Flash

DeepSeek

DeepSeek R1, DeepSeek V3

xAI

Grok family

Groq

Ultra-low-latency inference on open models

Mistral

Mistral Large, Mistral Small

Qwen

Qwen large-context and coding models

Meta

Llama 3.3 70B, Llama 3.1 8B

Perplexity

Search-grounded generation

Cohere

Command R+, Command R

MiniMax

MiniMax reasoning and chat models

Amazon

Nova model family

MoonshotAI

Kimi long-context models

NVIDIA

NIM-hosted open model endpoints

ByteDance

Seed model family

Sakana AI

Evolved/merged model research releases

Z.ai

Chat and reasoning models

Managed keys

We hold the provider keys, handle rate limits and failover, and bill through your I amToken wallet.

Bring your own keys (BYOK)

Connect your own OpenAI, Anthropic, or Google keys — encrypted at rest — and pay providers directly.

Pricing & plans

Simple monthly plans, each with wallet credit included for inference. The routing engine stretches that credit further — you pay pure passthrough only on what you actually use.

Plus

For individuals & power users

$20 / month
  • $16 wallet credit included
  • Full routing engine + quality layer
  • Semantic caching
  • Chat + API access
  • Community support

Business

For teams & organizations

$499 / month
  • $399 wallet credit included
  • Everything in Pro
  • Workspace, roles & departments
  • Custom routing rules
  • 24/7 priority support + SLA

Enterprise

High volume & custom needs

Custom
  • Volume-based pricing
  • SSO & procurement
  • Custom routing rules + SLA
  • Dedicated support
Contact sales

Products

One routing engine, four ways to use it — chat in the browser, call it from your own code, roll it out to a team, or watch exactly how it decides.

Routing decisions, live

Example week · single-vendor flagship spend vs. I amToken routed spend. (Chart lives on the real Products page — design your motion over this area if you want.)

Chat

The routing engine, as a conversation.

Ask anything — every message is routed to the model that fits it best, with semantic caching so you never pay twice for the same question.

Developer API

Drop-in routing for your own app.

Generate an iam_sk_ key and call /v1/route from any codebase. Same engine, same Pay-per-Saving billing.

Workspace

One routing engine, a whole team.

Seats, roles, shared prompts, departments, per-department cost tracking. SSO planned — not live yet.

Routing analytics & lifecycle

See why a model was chosen, not just that it was.

Replay any request as a lifecycle run: segmentation, intent, router decision, execution, quality scoring.

About Us

I amToken exists because picking one AI provider and living with its pricing, its outages, and its roadmap is a bad deal for anyone building on top of language models. We built a routing gateway that sits in front of every major provider, decides which model actually fits each request, and only makes money when that decision saves you money.

That's the whole thesis: route each prompt to the best model at the best price, and charge a cut of the savings versus a premium-flagship baseline — not a markup on spend.

What we believe

Provider-neutral by design

We don't get paid more when you use a more expensive model.

We charge for savings, not spend

If routing doesn't save you money on a request, we don't make money on it either.

Nothing routed in a black box

Intent, model choice, and quality score are visible in lifecycle analytics.

Your keys, your choice

Managed keys or BYOK — same routing, caching, and analytics.

Contact Us

Fill out the form below and we'll get back to you. Or email a specific inbox if you already know who you need.

Send a message

Name, email, subject, and message.

Send

General & support

support@iam-token.com

Billing

billing@iam-token.com

Privacy

privacy@iam-token.com

Quickstart

1) Start backend and frontend

make run-gateway
cd web && npm run dev

2) Use the Chat UI

Register, then open Chat. Settings for BYOK, Keys for API keys, Dashboard for usage, Routing for quality/cost insights.

3) Call the developer API

curl -H "Authorization: Bearer iam_sk_..." \
  -H "Content-Type: application/json" \
  -d '{"prompt":"Hello"}' \
  http://localhost:8000/v1/route