Step 01 / Swap
Change one line
Point your base URL at api.infercut.com/v1 and keep your existing SDK. Prompts, parameters and response formats work unchanged. That's the whole pull request.
IFC—01 / Inference savings endpoint
InferCut is the drop-in, OpenAI-compatible endpoint that reads every call — task, complexity, context — and serves it through the cheapest strategy that still solves it. One API key for all the LLM calls in your app. You never pick a model again.
Trusted by teams shipping at
Your app doesn't need a model picker. It needs every call solved at the lowest cost that still works. InferCut sits invisibly between your code and every model — reading the task, weighing its complexity, assigning the strategy, compressing, caching — and hands the difference back to you. One endpoint. One key. Zero selection overhead. No lock-in. No asterisks. Just the cut.
— The InferCut thesis, v2.1
Saved on the same calls — up to
Added latency at p50
Uptime SLA, in writing
From signup to first saving
The router, live — pick a workload
One endpoint, five layers of savings. Hover an instrument to inspect it.
Swap one base URL and point every LLM call in your app at it. Prompts, params and response formats work unchanged.
02The router reads each task and its complexity, then assigns the model and path that solve it at the lowest cost. You never choose.
03Shrink context, keep semantics. Applied only when the task still solves without what was cut.
04Micro-batching and semantic caching, invisible to your application.
05Unoptimizable calls ride free. You never pay above list price. Ever.
Step 01 / Swap
Point your base URL at api.infercut.com/v1 and keep your existing SDK. Prompts, parameters and response formats work unchanged. That's the whole pull request.
Step 02 / Route
Each request is profiled — task, complexity, context — then assigned across routes in real time to the strategy that solves it at the lowest cost. Fallbacks are automatic, so outages never reach your users.
Step 03 / Optimize
Compression, micro-batching and semantic caching fire only when quality checks pass for that model. Anything that would degrade output is skipped, untouched.
Step 04 / Save
Identical results land in your app — the invoice lands up to 90% lighter. Itemized savings reports show exactly where every cent was cut. No lock-in, ever.
“It turned out to be a one-line pull request. Our bill dropped before lunch.”
“We stopped negotiating with providers. InferCut just cut the price.”
“Agency margins live and die on inference. This fixed the dying part.”
No subscriptions. No tiers. No minimums. Prepaid credits — every request billed for what it actually used, at the rate the router earned.
Prepaid credits
Top up once, spend as you ship. Every request deducts credits based on what it actually used — balance and per-call consumption live in your dashboard. Credits never expire.
What each call deducts is between you and your dashboard — no rate tables to memorize, no per-model spreadsheets to maintain.
The router only takes a cheaper path when the task still solves on it — every strategy is validated against a parity harness before it carries traffic. If a call can't be served at lower cost without changing the result, it runs at full strength, untouched. Identical in, identical out.
More than 90% of the teams we've onboarded save between 50% and 90% on the same API calls — identical prompts, same job done. Your exact number depends on your traffic mix: chat-heavy workloads lean on the cache layer, batch jobs lean on routing and compression.
No — that's the point. One API key and one endpoint serve every LLM call in your app. The router detects each task and its complexity and assigns the right engine automatically, call by call. Zero selection overhead, no model spreadsheets to maintain, nothing to reconfigure when the market shifts.
Zero retention of prompts and completions. No training on your data, TLS in transit, and user-controlled rotatable API keys. Cache hits are keyed by hash, never by content storage. We can't read your traffic — by design, not policy.
Zero milliseconds added at p50 — routing happens while your request is already in flight, and cache hits are often net-negative because the model never runs at all. The slowest InferCut path is the one you were already taking.
Prepaid credits: top up with a card, and every request deducts credits based on what it actually used — your balance and per-call consumption live in the dashboard. No subscriptions, no tiers, no minimums, and credits never expire. New teams start with $5 in free credits, no card required.
No card required — $5 in free credits, two-minute setup