Discussion about this post

User's avatar
Herve Cuviliez's avatar

And it will continue

New model on the block

The most intelligence per dollar.

Amazing results from @pathway_com team on ARC-AGI-1 : $0.0007 at 29.5%

~11x cheaper than GPT 5.6 Luna

Paper:

https://huggingface.co/papers/2608.09888

Rohit Yadav's avatar

The $47,000 burned by a four agent system that looped for eleven days is the number I keep coming back to here. It is a cleaner illustration than the aggregate stats of why token price and AI spend move in opposite directions, the collapse in price is exactly what let a mistake like that get so expensive before anyone noticed it.

I looked at the same problem from the model routing side. Teams that switched workloads to smaller models were not doing it for the unit economics alone, they were doing it because flat per seat pricing could not absorb a variable cost that scales with usage the way you describe with the 330x jump in Google's token volume.

The part I would push on is the caching and hard cap advice. Both are simple to state and both require someone to own the token line item as a real budget, not a shared infrastructure cost nobody is accountable for. Are you seeing engineering teams actually get that ownership, or does it still sit with whoever controls the cloud bill?

I dug into the model routing side of this cost problem here: https://theintelligencestack.substack.com/p/the-token-tax-why-enterprises-are

No posts

Ready for more?