The Workers AI models that could be the LLM, and what they cost.
3use crate::{MAX_TOKENS, Usage};
Workers AI's free allocation, shared by the whole account, reset at 00:00 UTC (developers.cloudflare.com/workers-ai/platform/pricing, 2026-10-02).
7pub const FREE_NEURONS_PER_DAY: f64 = 10_000.0;
A text model with function calling, and its price in neurons per million tokens, from the same page on the same day.
Every candidate under Haiku's price, cheapest first by a typical call (about 700 tokens in, 200 out). The eval walks this list in order.
20pub const CANDIDATES: &[Model] = &[ 21 Model { id: "@cf/ibm-granite/granite-4.0-h-micro", neurons_per_m_in: 1542.0, neurons_per_m_out: 10158.0 }, 22 Model { id: "@cf/qwen/qwen3-30b-a3b-fp8", neurons_per_m_in: 4625.0, neurons_per_m_out: 30475.0 }, 23 Model { id: "@cf/zai-org/glm-4.7-flash", neurons_per_m_in: 5500.0, neurons_per_m_out: 36400.0 }, 24 Model { id: "@cf/google/gemma-4-26b-a4b-it", neurons_per_m_in: 9091.0, neurons_per_m_out: 27273.0 }, 25 Model { id: "@cf/zai-org/glm-5.3-flash", neurons_per_m_in: 13636.0, neurons_per_m_out: 45455.0 }, 26 Model { id: "@cf/openai/gpt-oss-20b", neurons_per_m_in: 18182.0, neurons_per_m_out: 27273.0 }, 27 Model { id: "@cf/meta/llama-4-scout-17b-16e-instruct", neurons_per_m_in: 24545.0, neurons_per_m_out: 77273.0 }, 28 Model { id: "@cf/mistralai/mistral-small-3.1-24b-instruct", neurons_per_m_in: 31876.0, neurons_per_m_out: 50488.0 }, 29 Model { id: "@cf/openai/gpt-oss-120b", neurons_per_m_in: 31818.0, neurons_per_m_out: 68182.0 }, 30 Model { id: "@cf/meta/llama-3.3-70b-instruct-fp8-fast", neurons_per_m_in: 26668.0, neurons_per_m_out: 204805.0 }, 31];
What a reply cost: Cloudflare's count when it gave one, the price list otherwise.
The most a request can cost, before it is sent: the prompt at one token per two bytes (generous for English and for JSON), and a reply that runs to the token ceiling.