AI Gateway custom costs now support cache token rates
AI Gateway custom costs now support per_cache_read_token and per_cache_write_token rates via the cf-aig-custom-cost header, allowing custom cost metrics to reflect negotiated cache pricing across different providers. The system automatically handles provider differences and prevents double-counting of cache tokens.
AI Gateway custom costs now support cache-read and cache-write token rates. This lets custom cost metrics reflect negotiated cache pricing across providers.
Add per_cache_read_token or per_cache_write_token to the cf-aig-custom-cost header:
{
"per_token_in": 0.000001,
"per_token_out": 0.000002,
"per_cache_read_token": 0.0000001,
"per_cache_write_token": 0.0000005
}
Cache-token pricing activates when either cache rate is present. An omitted cache rate defaults to per_token_in. If both cache rates are omitted, AI Gateway preserves the existing input and output calculation.
Providers can include cache tokens within input tokens or report them separately. AI Gateway automatically accounts for these differences and prevents double-counting.
For more information, refer to Custom costs.
Source: original entry ↗