Why we ran this
Our free chat gives every new account its first 3 replies on Mistral Large 4 and then switches to Mistral Small 4; Pro and credit packs stay on Large 4. We wanted to know, with real tasks rather than benchmark tables, when the bigger model earns its price: at list price Large 4 costs $1.36 per million input tokens and $4.18 per million output tokens, Small 4 $0.15 and $0.60, about nine times more on input and seven times more on output.
Method: eight prompts that look like actual work (not riddles), sent once each to mistral-large-4 and mistral-small-2603 through the chat completions API with reasoning_effort set to none, temperature 0.2 and a 1,500-token output cap. Time is wall-clock from one machine in Asia, so your absolute numbers will differ; the ratio between the two models is the point. One run per prompt, so treat small differences as noise.
