io.github.InferIndex/inferindex

InferIndex

LLM API prices across 70+ providers: cheapest offer, comparisons, history and cost estimates.

1.0.0
Version
remote
Transport
8
Tools

Security review

Review passed

Reviewed 1d ago.

  • tools: 8 tools scanned
  • metadata: scanned

No findings.

Tools (8)

  • search_models

    Find the exact id of an LLM tracked by InferIndex from a name or partial name (e.g. 'deepseek', 'qwen3 max', 'claude opus'). Returns matching model ids and names, best match first.

  • cheapest

    Cheapest current API offers for one model across direct providers and aggregators, in USD per 1M tokens (input, output, blended 3:1). Stale prices, and flex/batch tiers, are excluded by default. Optional filters (context, tools, JSON, vision, region, no training on prompts, open sign-up) and usage (tokens per request, requests per day) to get an estimated cost per request and per month. Returns the winner in detail and one short line per following offer. A condition that is absent was not published by the provider, it never means "no"; a flag not listed in signals is false.

  • compare_providers

    Current offers for one model, one line per provider and source (direct or via an aggregator), cheapest first (10 by default), with price, quantization, signals and training on prompts when published; the cheapest offer comes in detail. Pass detail: "full" for context, every published condition (data regions, sign-up) and reliability from official status pages on each offer.

  • price_history

    Price history of one model: every offer tracked by InferIndex (daily or weekly min/max/last price in USD, or raw price changes), plus the official price of the model's lab over time. Give either days, or from/to (YYYY-MM-DD), or at (a date) for the prices in effect that day.

  • estimate_cost

    Estimated cost of a workload on one model at each provider: cost per request, and per month if requests_per_day is given, taking the provider's tiered pricing and prompt-cache price into account. Offers sorted by estimated cost, cheapest first: the cheapest in detail, then one short line per offer with its estimated cost.

  • list_gpus

    GPU types whose rental price InferIndex tracks (H100, A100, L40S…), with the lowest price per GPU per hour in USD for each tier (guaranteed, community, spot) and the number of GPUs of that configuration. Prices are per GPU per hour, as published by each provider, and dated. Tiers are different products and are never compared with each other.

  • gpu_rentals

    Rental offers for one GPU type: provider, price in USD per GPU per hour, number of GPUs in the published configuration, billing and region when published. One block per tier (guaranteed, community, spot), each with its own cheapest offer, also given per configuration (cheapest_by_gpu_count: the price per GPU of a single GPU and of a multi-GPU node are not comparable); tiers are never compared with each other. guaranteed: capacity the provider presents as not interrupted; community: third-party hosts; spot: interruptible capacity. Prices are as published by the provider, and dated. Get the GPU id from list_gpus.

  • self_host_or_api

    Is it cheaper to host an open-weights model yourself on rented GPUs, or to use the cheapest API offer? Returns a one-sentence verdict (headline) and the numbers behind it: the break-even volume in million tokens per day, the throughput the whole configuration must deliver for self-hosting to cost less (required_tokens_per_second: compare it with what your setup does), the break-even utilization when a throughput is known, the GPU configuration and its hourly price, and what the verdict rests on (confidence). Without a volume you still get the verdict and the break-even volume. It is an estimate: the throughput is a published measurement (measured), an estimate scaled from one (derived; estimated for a wide range), or a published minimum (lower_bound: a floor, valid for requests of up to 2,048 tokens in total, that can show self-hosting wins, never that the API is cheaper); or you give your own tokens_per_second. When no published figure decides, the headline gives the break-even volume