Skip to content
ShareQuota

Glossary

The words behind this product and every router like it, each defined once — what it means in general, then what it means here.

AI router

A program that sits between your tools and several model providers, exposes one endpoint, and decides per request which provider answers — usually with failover when one refuses.

ShareQuota is an AI router for five providers, running as one binary on your own machine.

LLM proxy

A server that receives a request meant for a language model, may alter or reroute it, and forwards it upstream — the same shape as an HTTP proxy, applied to model APIs. Often used interchangeably with AI router and LLM gateway.

The proxy listens on localhost, translates dialects, and relays the stream back unchanged in the dialect you called.

LLM gateway

An LLM proxy positioned as shared infrastructure: one entry point for a team or a product, with keys, limits and logging managed centrally.

Not this. The management API only answers on loopback; a shared gateway is what 9router or a team deployment is for.

Subscription pooling

Signing in to the consumer plans you already pay for — Claude Pro, ChatGPT Plus, a Gemini plan — and using their quota through an API instead of buying tokens separately.

The core job. The binary signs in with OAuth, on your machine, and spends the plan you hold; nothing is resold and no one else’s quota is involved.

OAuth login

An authorisation flow in which you sign in at the provider’s own page and the application receives a token, rather than you handing it a password.

Each subscription connection is an OAuth login completed in your browser; the token lives in a local SQLite file and refreshes itself.

API dialect

The request and response format one provider’s API speaks — the Anthropic messages shape, OpenAI chat completions, OpenAI Responses, Google generateContent. Tools are written against one of them.

Four dialects on one endpoint, and a translator between them, so a tool that speaks one can reach a provider that speaks another.

Dialect translation

Converting a request from one API dialect to another — and the response, including a streamed one — so a client and a provider that were never built to talk can.

Anthropic, OpenAI and Gemini shapes are translated in both directions, held to golden fixtures from a reference implementation.

Auto-failover

Moving a request to another provider or account automatically when the first one fails, refuses, or is out of quota, without the client noticing.

When a connection refuses, the router moves the request to another that serves the same model and puts the exhausted one in a cooldown.

Cooldown and backoff

After a failure, waiting before retrying the same target, and waiting longer after each further failure, so a rate-limited provider is not hammered.

Exhausted connections back off exponentially and return to rotation on their own; the constants live in the config package.

Rate limit

A provider’s cap on how many requests or tokens an account may use in a window, signalled with an HTTP 429 or an equivalent error.

A rate limit on one account is the trigger for failover to the next; it is not bypassed, it is routed around with your own other accounts.

Quota

The amount of usage a plan or key allows — messages, tokens or sessions per window — and the thing that expires unused at the end of a month.

The quota you already paid for is the whole reason the product exists: it is made reachable from your tools rather than left inside each provider’s app.

Context window

The maximum number of tokens a model can take in one request, prompt and history together. Beyond it, the request is refused or truncated.

Every model page prints the window the registry reports; the largest in the catalogue is one million tokens.

MTok

One million tokens — the unit providers price API usage in, quoted separately for input, output and cached input.

Model pages show list prices per MTok so a month of a subscription can be compared with what it would cost per token.

Free tier

Usage a provider grants at no charge, usually with lower limits — a real quota, not a trial.

Free-tier connections pool like any other; the catalogue marks the models that carry one.

Local endpoint

A URL served on your own machine, reachable as localhost and nowhere else unless you publish it.

http://localhost:20130 is the whole surface. Nothing is hosted; there is no server in the middle.

Internal key

The single credential a proxy issues to its own clients, standing in for every provider key and login behind it.

The sq- key, generated on first run. Tools only ever see it, so a leaked tool config leaks nothing that can be billed.

Model routing

The rule that maps the model field of a request to the provider that will serve it — explicit prefixes, aliases, or inference from the id.

provider/model routes explicitly; claude-, gpt-, gemini- and deepseek- prefixes are inferred; any other slash-carrying id goes to OpenRouter.

Reversible integration

Changing a tool’s configuration in a way that can be undone exactly — a backup taken first, the original values restored on disable.

The cardinal rule of the CLI integrations: never destroy user configuration. The diff after enable-then-disable is empty.

Tunnel

Publishing a local port at a public https address through a relay — a Cloudflare quick tunnel or a Tailscale Funnel — so a service that calls from its own servers can reach it.

What Cursor needs, because Cursor’s backend makes the request. Everything else on this site works on localhost.