Input tokens estimated from characters (contract Execution, "A
request fits the model's context"). The tokenizer is unpublished and
may not be derived (MCA §2.3(c)), so the ratio is learned from what
answers reported: characters sent per usage.input_tokens, per model,
falling back to jev-axi's conservative 2.6 when nothing is known.
A request is REQUEST_TOKENS of fixed overhead plus its characters
over the ratio, so the ratio is fitted to the tokens above that
overhead, in the same form the estimate uses.
Tokens of fixed overhead in every request (Hume's measurement).
12pub const REQUEST_TOKENS: f64 = 267.0;
Attempts a request may be billed for: the first and two retries, since there is no idempotency key (contract Cost and safety).
16pub const BILLED_ATTEMPTS: f64 = 3.0;
Characters per token: learned, or the fallback.
One answered request: the characters it sent and the input tokens its answer reported.
30impl TokenRatio {
jev-axi's conservative figure, when no answer has been seen.
32 pub const FALLBACK: TokenRatio = TokenRatio(2.6);
The characters of every sample over their tokens above the request overhead, so large requests weigh as what they cost. The fallback when the samples carry no tokens or no characters above it, since nothing was learned.
38 pub fn learn(samples: impl IntoIterator<Item = Sample>) -> TokenRatio { 39 let (chars, tokens) = samples.into_iter().fold((0.0, 0.0), |(c, t), s| { 40 (c + s.chars as f64, t + (s.input_tokens as f64 - REQUEST_TOKENS).max(0.0)) 41 }); 42 if chars > 0.0 && tokens > 0.0 { TokenRatio(chars / tokens) } else { TokenRatio::FALLBACK } 43 }
Input tokens of requests requests carrying chars characters.
Dollars requests requests carrying chars characters cost at
worst: every one billed [BILLED_ATTEMPTS] times, at
price_per_mtok dollars per million input tokens.
62#[cfg(test)] 63mod tests { 64 use super::*; 65 66 #[test] 67 fn nothing_learned_is_the_fallback() { 68 assert_eq!(TokenRatio::learn([]), TokenRatio::FALLBACK); 69 // All overhead: no characters' worth of tokens to divide by. 70 assert_eq!(TokenRatio::learn([Sample { chars: 100, input_tokens: 267 }]), TokenRatio::FALLBACK); 71 assert_eq!(TokenRatio::learn([Sample { chars: 0, input_tokens: 500 }]), TokenRatio::FALLBACK); 72 } 73 74 #[test] 75 fn learns_characters_per_token_above_the_overhead() { 76 let ratio = TokenRatio::learn([ 77 Sample { chars: 400, input_tokens: 367 }, 78 Sample { chars: 800, input_tokens: 467 }, 79 ]); 80 assert_eq!(ratio.chars_per_token(), 4.0); 81 // The estimate in the same form reproduces what was reported. 82 assert_eq!(ratio.tokens(2.0, 1200.0), 834.0); 83 } 84 85 #[test] 86 fn the_fallback_estimate() { 87 assert_eq!(TokenRatio::FALLBACK.tokens(1.0, 26.0), 277.0); 88 } 89 90 #[test] 91 fn the_worst_case_bills_three_attempts() { 92 assert_eq!(TokenRatio::FALLBACK.worst_case(1.0, 26.0, 1e6), 831.0); 93 } 94}