handleForcedSSEToJson dropped cached prompt tokens in two ways: the Responses branch summed only input_tokens, which excludes cache_read and cache_creation on cache-capable upstreams (measured 2012 reported vs ~5344 actual, 5332 from cache); and the Chat Completions branch computed usage correctly but it didn't always reach the client (an Anthropic response with cache_read_input_tokens: 11022 arrived with no usage field at all). Now folds cache counters into prompt_tokens, surfaces them via prompt_tokens_details, and re-attaches usage before serialisation. |
||
|---|---|---|
| .. | ||
| chatCore | ||
| embeddingProviders | ||
| fetch | ||
| imageProviders | ||
| search | ||
| ttsProviders | ||
| chatCore.js | ||
| embeddingsCore.js | ||
| imageGenerationCore.js | ||
| responsesHandler.js | ||
| sttCore.js | ||
| ttsCore.js | ||
| videoCore.js | ||