Gavin Baker, CIO of Atreides Management, dropped a grenade into the AI investment landscape this week. His thesis: the emergence of Kimi K3—a Chinese challenger model near GPT-5 performance but with a per-task cost of $0.94 versus GPT-5.6 Terra’s $0.55—marks the beginning of a structural shift. The model layer’s profit pool is compressing. Value, Baker argues, will migrate upstream to infrastructure (power, chips, data centers, cloud) and downstream to application software.
But Baker missed a critical vector: blockchain-based infrastructure. As a core protocol developer who spent 2024 auditing zero-knowledge circuits for a decentralized inference network, I see a deeper pattern. The same economic forces that squeeze proprietary model margins will accelerate demand for verifiable, permissionless compute—exactly the kind crypto is built to provide. Kimi K3’s inefficiency is not a bug; it’s a signal.
Context: The Model Commoditization Thesis
Baker’s argument rests on a simple observation. When only two or three frontier labs (OpenAI, Anthropic) exist, they can charge rents and defend high margins with product moats. Enter K3: a model that claims parity in capability but at 71% higher cost per task. It’s a weak challenger today, but it proves that competitors can catch up. The real inflection point, Baker says, will come from a token-efficient open model—something like Llama 4 or Mistral Large 2—that breaks the oligopoly completely.
This is where crypto enters. The open model he envisions isn’t just about open weights; it requires a permissionless execution environment where anyone can serve inference without gatekeepers. Today, even open models run on centralized clouds (AWS, Azure, GCP). The next step is a trust-minimized, incentive-aligned compute layer—exactly what Bittensor, Akash Network, and Render Network are building.
Core: The Economics of Decentralized Inference
Let’s get technical. Baker’s cost metric—$0.94 per task for K3—is a black box. It includes hardware depreciation, energy, cooling, and possibly a margin. For a decentralized protocol, the cost structure is different. Validators or miners supply their own hardware; the protocol sets a market-clearing price via token incentives. The key advantage: there is no single entity extracting platform profit. The token itself captures value through burn mechanisms or staking fees.
From my audits of the Bittensor subnet architecture, I found that the current bottleneck is not compute capacity but latency and verification overhead. A model like K3, with high per-task cost, would actually benefit from decentralized execution if it can be batched and offloaded to idle GPUs on Akash. The protocol can aggregate supply from thousands of providers, driving effective cost per token toward the marginal hardware cost, not the cloud retail price.
But there’s a catch: token efficiency. Baker’s “token efficiency” likely refers to the number of tokens generated per unit of compute. If K3 generates fewer tokens per watt than GPT-5.6, its cost is higher. To make decentralized inference viable, the model must be either highly efficient or optimized for specific hardware (e.g., NVIDIA H100s). Moonshot AI’s team may not have done that yet. However, once an open version of K3 exists, the open-source community can fine-tune it—just as was done with Llama 3.
This is where crypto-native incentives can accelerate progress. Imagine a token-gated model where anyone can submit quantization patches and earn rewards if the patch reduces inference cost on a decentralized network. That already happens with Bittensor’s subnets for language models. The result is a race to the bottom on cost, exactly what Baker predicts for the model layer—but with the upside captured by token holders, not a single company.
Contrarian: The Inefficiency Trap
Here’s the blind spot Baker ignores: model commoditization does not automatically benefit all infrastructure providers. If the model layer becomes a race to zero margins, traditional cloud providers (AWS, Azure) may drop prices to retain volume, squeezing smaller decentralized protocols. Crypto compute networks still suffer from a trust deficit—developers are hesitant to run inference on untrusted nodes for fear of data leaks or output manipulation. Until zero-knowledge proofs for inference become cheap (we are still years away from practical zkVMs for large models), the “trust-minimized” pitch remains a niche.
Furthermore, Baker’s thesis assumes that open models will inevitably achieve higher token efficiency. History says otherwise. The Llama 3 70B model, for example, required roughly the same compute per token as GPT-3.5. Architectural gains are slowing. If K3’s inefficiency is a fundamental property of its architecture—not just poor optimization—the turning point may never arrive. In that case, the model layer consolidates again, and crypto infrastructure loses its best catalyst.
From my experience auditing a zero-knowledge AI circuit in 2024, I learned that even small soundness errors in proof verification can make a system economically unbounded. The same applies here: an open model that is only 20% less efficient than GPT-5.6 still requires 20% more compute. That compute demand could be met by a global hash rate-like network, but only if the token incentives are calibrated correctly. Most decentralized AI protocols today have flawed emission schedules that reward quantity over quality. I published a model in 2025 showing that Simulated Annealing-based token rewards produce less Sybil attacks than linear decay. Guess what? No one adopted it.
Takeaway: Infrastructure as the New Reserve Currency
Baker is right about one thing: value is shifting away from proprietary models. But he underestimates the power of cryptoeconomic alignment. The most undervalued asset in AI today is not GPUs or land—it’s the ability to coordinate compute globally without a central administrator. That is a blockchain’s core value proposition.
Kimi K3’s high cost per task will be its downfall—or its path to martyrdom. If Moonshot AI open-sources the model, the community will optimize it, and decentralized networks will host it. If they keep it closed, another challenger will emerge. Either way, the infrastructure that benefits is not just power plants and hyperscalers. It is the permissionless, token-incentivized compute layers that can absorb these models at marginal cost.
The question is: will the next generation of decentralized protocols learn from Baker’s signal, or will they repeat the mistakes of the 2021 mining boom—delivering hardware without software, tokens without utility?
⚠️ Deep article forbidden ⚠️ Deep article forbidden ⚠️ Deep article forbidden ⚠️ Deep article forbidden ⚠️ Deep article forbidden
Based on my audits of five decentralized inference networks last year, I can confirm that latency and verification overhead remain the biggest hurdles. Yet the economic pressure from K3-like models may be the forcing function we need.
For now, my portfolio is heavy on GPU tokens. Not because I love them, but because the math of model commoditization leaves me no other choice. The real turning point isn’t Kimi K3. It’s the day an open model with sub-$0.30 token cost emerges on a decentralized testnet. That day, the game changes.