1 of 158 calls
in this range have no recorded cost. That means LiteLLM could not price the
model those calls ran against — usually a model string it does not know.
Their token counts are still recorded (tokens are accounted before
pricing is attempted), so the spend is recoverable once the model is priced.
Check llm_cost_calc_failed in the server logs for the model
name, and register a price for it rather than reading these as $0.
Feedback only attaches to a question when the calling model hands back a
valid request_id. Unlinked rows are real feedback that
reached no question page — listed below.
Advice on a plain lookup must read 0 — the firing policy refuses it, so any other value is a bug rather than a dial to turn. A high offered a sharper question rate is the thing to fix first: a check that fires on everything gets ignored on everything, and once the closing lines read as boilerplate the source citation is skipped with them. Asked and never resumed counts users left holding a question nobody relayed an answer to — a client that dropped the conversation. If that climbs, switch off the probe before the ambiguity check: one is a product experiment, the other is Chandler refusing to guess between two readings.
(not classified) is a real state, not a gap: an exact cache
hit answers before the router runs. If decision_shaped runs much
above a third of classified questions the router is too generous, and the
offer will feel constant however tight the gates below it are.
Cost is the LLM completion spend LiteLLM priced for each call, summed per
request.
Chandler is running on the Codex CLI, which draws ChatGPT plan quota
rather than per-token API billing, so this figure is not money paid —
it is what the same traffic would have cost on the API, priced from the
token counts Codex reports. Read it as the size of the bill avoided.
Two known limits: embedding spend is not recorded, so
retrieval and matching cost is missing, and because the figure is one sum
per request it cannot be split by model within a request — a
request's classifier and agent calls can use different models while
llm_model holds only one name.
Only authenticated calls are logged, USAGE_LOG_ENABLED can
switch logging off, and log writes are fail-soft — this is not a complete
record of traffic. Questions are grouped into sessions: a session is
exact when the caller echoed its id back to us and otherwise inferred from
a 30-minute gap in that user’s activity, so a grouping is only as
good as the source shown on the session itself. A call with no
attributable user gets no session at all; those questions are listed
separately rather than dropped.