PyxGrant / Model spend

A budget that says no before the spend.

The model proxy sits in front of your model API. It meters every call by identity, refuses once the budget is spent, and breaks a runaway loop before it drains the window. Each refusal is a signed decision in the same audit log as everything else.

one identity, one window budgetillustration
billing-agent · 24h budget$0.00 / $5.00
    Budgets

    A ceiling per identity, checked before the call is sent.

    Set a cost or token budget per principal for a window, 24 hours by default, or give each identity its own in a budgets file. Once it is spent, the next call gets a 429 and never reaches the provider. Setting "locked": true stops an identity's spend outright.

    On the record. Each refusal is written to the audit chain with the spend it prevented.
    Some models need a person. -hold-models parks every request to the models you name until someone approves it.
    Same agency counter. Each forwarded completion spends one step on the shared grant budget, the same counter a tool call draws. A spent counter returns 429 agency-exhausted. Under the hardened profile, a model with no configured price is refused with 403 model-unpriced.
    response from the model proxy429
    HTTP/1.1 429 Too Many Requests
    X-PyxGrant-Refusal: budget-exhausted
    
    {"error": {
      "message": "PyxGrant: …",
      "type": "pyxgrant_budget_exhausted"
    }}
    Loops

    Stop the retry storm before the budget does.

    A loop can burn a day's budget in minutes. Spend-rate and call-rate limits refuse a principal that spikes, marked wallet-abuse so your SIEM can tell it from an ordinary spent budget. Identical repeats can be routed to a cheaper model to break the loop instead of failing it.

    Exact repeats for free. With -cache on, a byte-identical request is answered from the cache at $0 and metered as savings.
    Output caps. -max-output-tokens trims what a single request can ask for.
    pyxgrant llmflags
    $ pyxgrant llm \
        -budget-cost 5 -budget-window 24h \
        -max-cost-per-min 0.50 \
        -max-calls-per-min 60 \
        -loop-degrade-threshold 3 \
        -loop-degrade-model a-cheaper-model \
        -cache
    Runs and fleets

    One ceiling, however many processes share it.

    Tag a workflow with a run id and every gateway on that run, on the model and tool side, is tied to it. pyxgrant runs shows what each run has spent and which sessions joined it. A freeze on the run stops all of them.

    Shared spend store. Point several proxies at one file or Postgres store and the window budget becomes one ceiling across all of them.
    No double admission. -reserve-budget holds headroom before each call, so two processes can't both pass a nearly spent budget.
    shared ceiling-spend-store
    proxy A  reserves  $0.40  admitted
    proxy B  reserves  $0.40  429  no headroom left
    proxy A  metered   $0.31  reservation released
    
    one window budget across both
    The ledger

    Who spent what, and where it is heading.

    pyxgrant economy prints the consumption ledger per identity: tokens, cost, and budget, with a forecast of who runs out first. pyxgrant quote issues a signed cost quote that a spend can be checked against later, offline.

    pyxgrant economyper identity
    identity        tokens   cost    budget
    billing-agent   …        …       $5.00
    claims-bot      …        …       $20.00
    values come from your own audit log

    Put a ceiling on agent spend this week.

    The model proxy runs from the same binary that already decides your tool calls.