The Camry Guest Experience
The trigger Two AI model launches, forty-eight hours apart, are the reason for this paper. On 15 July 2026, Mira Murati's Thinking Machines Lab released Inkling, a 975-billion-parameter open-weight model. The company said outright it isn't the strongest model on the market — the pitch is that enterprises would rather own and fine-tune a good model than rent a great one. A day later, China's Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight system that independent evaluators ranked close to, and on some coding benchmarks ahead of, the leading closed frontier models from the US labs. Neither event is, on its own, remarkable. Model launches happen weekly. What makes this pairing worth a Viewpoint is the price signal underneath it. Kimi K3's output pricing runs around $15 per million tokens. The frontier closed model it's being benchmarked against runs roughly $50 for the same volume. That's not a discount. That's a different market. The Lambo-and-Camry problem The trade press has taken to describing this as a Lambo-and-Camry moment — plenty of very good, very reliable options now sit on the lot beside the flagship. It's a tidy line, but it understates what's happening for an owner-operator, because a hotel guest has never once asked what car brought their room-service order up. They notice whether it arrived hot, on time, and correctly. The badge on the token is invisible to them. It is entirely visible on your P&L. The guest has never asked what model answered their question. They've only ever asked whether the answer was right. Reframing this through TCPG This is precisely the gap Token Cost Per Guest was built to expose: token consumption is not a technology metric; it's a unit-economics one, and it only means something once price is treated as a variable rather than a constant. Most owner-operators are still budgeting for AI the way they budgeted for their first PMS license — one vendor, one price, reviewed annually. That assumption breaks in an environment where the underlying compute for a comparable task can fall by an order of magnitude in a single news cycle, as it plainly has this week. The honest picture, though, is not "switch everything to the cheap model." It's that TCPG now needs to be modeled as two regimes, not one — and the next 12 months will be about learning to route between them rather than picking a single vendor and staying loyal to it out of habit. Regime Commodity tokens Judgment tokens Typical task Guest FAQ chat, review summarizing, routine drafting, translation Multi-system agentic orchestration, brand-voice output, service recovery, compliance-sensitive judgment Price trajectory Falling fast — open-weight models now price this a full order of magnitude below flagship closed models Sticky — frontier labs are holding premium pricing where the cost of a wrong answer outweighs the token saving TCPG posture Route to open-weight or mid-tier models; treat brand-name pricing here as overpaying Retain closed frontier models; the premium is buying risk reduction, not just intelligence What changes in
Share