LLM Routing

RouteIQAgentic LLM model routing & cost control

Send every query to the right-priced model — and never blow the budget.

Not every query needs your most expensive model — but most stacks send them all there anyway. RouteIQ is an agentic intelligent model router that decides, per query, which model tier to use and which model within it, then enforces budget the whole way. A routing agent maps each request to its project, scores how relevant the query is to that project's documentation, and routes on-topic, project-critical work to premium or mid-tier models while off-topic or trivial queries drop to cheap ones. It spans providers — Amazon Bedrock (Claude, Amazon Nova) and self-hosted open models on your own GPUs — and keeps humans in command through per-project policy, budget guardrails, and admin alerts.

Tiers

Premium · mid · cheap, per project

Providers

Bedrock + self-hosted open models

Budget controls

Burn-rate alerts · auto-downgrade

Every query

Logged: model · cost · outcome

Capabilities

What RouteIQ does.

Project-relevance routing

Three cost tiers, many providers

Budget guardrails & burn-rate alerts

Auditable routing decisions

How it works

5 stages, one accountable loop.

  1. 1

    Identify

    Each query is associated with the project it belongs to, along with that project's policy and budget.

  2. 2

    Score relevance

    Assess how relevant the query is to that project, using its documentation and recent context.

  3. 3

    Set the tier ceiling

    On-topic work is allowed premium or mid models; off-topic drops to cheap — and the ceiling lowers automatically if the budget is running hot.

  4. 4

    Select the model

    Within the chosen tier, pick a model according to the project's configured preference.

  5. 5

    Invoke, log & update budget

    Call the chosen model, return the answer, log the decision, and update the project's budget.

Benefits

Why teams choose it

  • Cut LLM spend by sending only the queries that need it to premium models
  • Never blow the budget — burn-rate alerts and automatic tier downgrades
  • Protect quality on project-critical work while deflecting off-topic queries to cheap models
  • Avoid lock-in — route across Bedrock and self-hosted open models together
  • Stop single-user abuse with per-user cost controls
  • Full, auditable logging of every routing decision and its cost

Use cases

Where it fits

  • Enterprise internal assistants serving many teams and projects from one gateway
  • Cost control across projects with per-team budgets and chargeback
  • Blending premium API models with cheaper self-hosted open models
  • Deflecting off-topic or trivial queries away from expensive models
  • Preventing budget overruns and single-user spend abuse
  • Standardising model access behind one policy-governed router

Integrations

Amazon BedrockClaude (Opus / Sonnet / Haiku)Amazon NovaAmazon SageMakervLLMQwen2.5 / Qwen3Llama 3.1Amazon Titan embeddingsOpenAIVirtual-key gateway

How it compares

A vendor-agnostic alternative

vs. OpenRouter / single LLM gateway

RouteIQ adds project-relevance routing and enforced per-project budgets with auto-downgrade — not just a pass-through to whichever model you name.

vs. LiteLLM / proxy routers

Beyond load-balancing and fallbacks, RouteIQ decides the cost tier from how relevant each query is to the project, governed by admin policy.

vs. RouteLLM / static cascades

Routing decisions are tuned per project rather than relying on hard-coded thresholds that drift across use cases.

vs. Manual model selection

No more hard-coding one model per app: the router picks per query and per budget state, governed by admin policy.

What is intelligent model routing?

Intelligent model routing (also called LLM routing or a model cascade) sends each request to the most cost-effective model that can still answer it well, instead of defaulting every query to one expensive model. The hard part is deciding, cheaply and reliably, which queries actually need the premium tier. RouteIQ answers that by scoring how relevant each query is to the project it belongs to, so the premium tier is reserved for the work that genuinely needs it while off-topic or trivial questions are served by cheaper models — often self-hosted open models running on your own GPUs.

Budgets and governance built in

Cost control isn't a dashboard you check after the fact — it's enforced in the routing path. Each project has a monthly budget with burn-rate detection that automatically downgrades tiers and can stop spend once the budget is exhausted, with admin alerts along the way. Every routing decision, model choice, and cost is logged for a compliance-ready audit trail.

FAQ

Common questions

How does RouteIQ decide which model to use?+

It scores how relevant each query is to the project it belongs to, then routes on-topic, project-critical queries to premium or mid-tier models and off-topic or trivial ones to cheap models. Within the chosen tier, a configurable preference picks the model — so the premium tier is reserved for the work that genuinely needs it.

Which models and providers does it support?+

RouteIQ routes across Amazon Bedrock (Claude Opus/Sonnet/Haiku and Amazon Nova) and self-hosted open models on your own GPUs, grouped into cheap, mid, and premium tiers per project — so you can blend managed APIs with self-hosting.

How does RouteIQ control cost?+

Each project has a monthly budget with burn-rate detection. As spend accelerates, RouteIQ automatically downgrades tiers and can stop spend once the budget is exhausted, with admin alerts and cost reports along the way.

Will it route private queries to third-party models?+

Routing is governed by per-project policy: admins define the allowed model pool per tier, so a project can restrict premium routing to in-account Bedrock models or self-hosted open models on SageMaker and keep sensitive traffic off external endpoints.

Is it human-in-the-loop?+

Yes. Enterprise admins set the levers — relevance thresholds, budgets, allowed models, and selection strategy — and receive alerts and cost reports as budgets burn. The agent routes within those policies; people stay in command of the rules and the spend.

Further reading

What is intelligent LLM model routing?

Read the guide

More accelerators

Part of the humaineeti platform

RouteIQ is engineered by humaineeti — and works alongside the rest of the platform.

Ready to deploy RouteIQ?

We’ll map RouteIQ to your stack, constraints, and compliance requirements — and keep humans in command.

Talk to humaineeti