Better models don’t automatically mean smaller AI bills.
Sometimes they mean you finally have a reason to run the workflow you couldn’t get working before.
That’s what makes GPT-6 Astra interesting to me.
Once a model can research, build, inspect, and refine, you start giving it more ambitious tasks. A single response becomes an entire working session.
The useful question is no longer just “what does a million tokens cost?”
It’s “what can I get done with this budget?”
That’s also why cashback on inference is worth paying attention to. On work you were already going to run, getting some of the spend back gives you a choice: keep the savings or fund another experiment.
But the goal should never be to burn more tokens.
It should be to get more verified, useful work out of the same budget.