英伟达为什么要保卫开放权重,以及它将如何重塑AI算力市场 / Why NVIDIA Is Defending Open-Weight Models—and How They Could Reshape the AI Compute Market
In mid-July, Moonshot released Kimi K3, a 2.8-trillion-parameter open-weight model that beat the flagships from OpenAI and Anthropic on the FrontierSWE coding benchmark. A few days later, Alibaba shipped a new version of Qwen. On July 20, Axios reported that the Trump administration was again pushing for a de facto ban on Chinese open-weight models; the instruments under discussion included Entity List designations, federal procurement restrictions, security advisories, liability rules, and pressure on US companies running Chinese models in production. Four days after that, on July 24, Jensen Huang used the first X post of his life to publish a policy letter titled “Open Weights and American AI Leadership.” The ask: Washington should not impose “premature restrictions” on downloadable AI models. The first version carried 25 corporate signatures, including NVIDIA, Microsoft, and Meta. Within a day it doubled to 50 — OpenAI and Google had signed; Amazon and Anthropic had not. By July 30, the official page hosted by Microsoft listed more than 230 signatories, and Amazon had appeared among them. The roster spans chips, servers, clouds, neoclouds, inference platforms, enterprise software, developer tools, and venture capital. Anthropic is the only frontier lab that has stayed away throughout.
Two ways inference demand can converge
The center of gravity in AI compute demand is shifting from training to inference. At a press Q&A at CES 2026, Huang offered a figure: one out of every four tokens generated today comes from an open model. There are two extreme shapes that inference demand could converge toward:
(1) Concentrated. Demand pools into a handful of closed frontier model companies. Enterprises rent intelligence, and tokens flow out of a few API endpoints.
(2) Diffused. Demand sinks down to tens of thousands of enterprises, inference platforms, neoclouds, sovereign AI programs, and application companies, each running its own specialized model.
Open weights are the condition that makes diffusion cheap and scalable. So this is a fight over market structure. And market structure determines whether NVIDIA faces five buyers or fifty thousand, and whether enterprises have any alternative to bring to the table when negotiating with closed APIs.
What NVIDIA wants is a fragmented buyer base
On the surface, NVIDIA should be neutral on model paradigms — whether closed labs dominate or an open ecosystem does, both sides buy GPUs. As long as AI workloads grow, NVIDIA makes money.
But NVIDIA cannot be neutral on the structure of its buyers. Its customer concentration is rising fast: in FY2026 Q3, four direct customers each accounted for more than 10% of total revenue, adding up to 61%; a year earlier it was three customers at 12% each, totaling 36%. We can only see concentration at the procurement layer, not the names of the end demand. But there is little disagreement in the market about who that procurement ultimately serves: Google, Amazon, Meta, Microsoft, OpenAI. And those same companies are the only ones capable of building an alternative to NVIDIA. Google has TPU, Amazon has Trainium, Meta has MTIA, Microsoft has Maia, and OpenAI is co-developing silicon with Broadcom. The weaknesses of these ASICs — narrow ecosystems, high CUDA migration costs — are not fatal for a hyperscale platform with a fixed internal workload, and those weaknesses are shrinking: both TPU and Trainium are now sold externally. Anthropic, for instance, uses Google TPU, AWS Trainium, and NVIDIA GPUs simultaneously, with the stated rationale of “matching the workload to the most suitable chip” — which is to say, large-scale frontier training can run on non-NVIDIA silicon. (NVIDIA is also one of Anthropic’s investors, so the two are not purely adversarial.) Which is why NVIDIA needs a world with a more fragmented buyer base. Tens of thousands of customers who can neither afford nor build their own chips are a far safer foundation than a few giants who are actively building substitutes. The Nemotron family, and the Nemotron alliance announced at GTC 2026 — with roughly $26 billion of company investment over five years — all indicate that NVIDIA is paying out of pocket to manufacture the fuel for a diffused world.
(That said, as NVIDIA’s own 10-K risk factors note, open-source AI depends on developer adoption, and if it is deployed on competitors’ platforms it could reduce demand for NVIDIA’s products. Open weights push demand downward, but which silicon it lands on once it gets there is not NVIDIA’s call.)
Sovereign AI: a new national-scale demand category
Sovereign AI refers to an AI capability stack that a country builds and controls itself: compute located within its borders, models under its own legal jurisdiction, and training data that includes its own languages, public data, and industrial knowledge. Three forces drive it — data localization (healthcare, financial, government, and defense data not flowing to foreign APIs), national security (critical systems cannot depend indefinitely on foreign companies’ closed models and terms of service), and language and culture (leading frontier models under-serve a great many non-English languages, dialects, and administrative and legal systems).
By NVIDIA’s own disclosure, sovereign AI revenue exceeded $30 billion for full-year FY2026, more than tripling year over year, with the main contributions coming from the UK, France, the Netherlands, Canada, and Singapore; FY2027 Q1 added new AI factory projects in Japan, South Korea, and Germany. The buyers of supercomputers used to be research institutions; the buyers of AI factories are finance ministries, defense ministries, national telecom operators, and sovereign wealth funds. Saudi money is betting on both tracks at once: domestically, PIF-owned HUMAIN is building gigawatt-scale compute and the Arabic model ALLAM, while Aramco Ventures led Together AI’s $800 million round in July, investing directly into the American open-weight inference layer. Nearly all of these countries are US allies, and nearly all of what they buy is NVIDIA. So what “sovereign AI” is actually doing at this stage is trading API dependence for chip dependence — and the chips still come from NVIDIA. For NVIDIA, sovereign AI is therefore a demand category that expands the number of buyers while posing no threat at all to its own moat.
For policymakers, a country that can only call foreign closed APIs does not have real sovereign AI; it needs models it can download, inspect, deploy, fine-tune, and maintain over the long term. Kimi, DeepSeek, Qwen, and GLM are mostly released under permissive licenses like MIT, which is exactly what sovereign programs need. If the US restricts its own open models, sovereign AI programs will naturally turn to Chinese ones.
Inference providers sell model utilization
Fireworks, Baseten, and Together are the most direct beneficiaries in this ecosystem. All three signed the letter. What they sell is the ability to turn a model into a production system: deployment, fine-tuning, adapters, distillation, autoscaling, GPU scheduling, and latency and cost optimization.
Capital markets placed their bets on this layer in a dense three-month stretch:
| Company | Round | Date | Amount | Valuation |
|---|---|---|---|---|
| Baseten | Series F | 2026-06-22 | $1.5B | $11B / $13B (two tranches) |
| Together AI | Series C | 2026-07-01 | $800M | $8.3B |
| Fireworks | Series D | 2026-07-16 | $1.505B | $17.5B |
What these companies really are is open-weight inference foundries. Fireworks has disclosed that more than 95% of the tokens it serves come from models specialized on customer data: fine-tunes, adapters, distillations, and models customers trained themselves and brought over to be hosted. In that last case, the inference provider takes no part in training and handles only serving — it is paid for the production system after the model goes live.
(A note: the service economics of adapters differ from traditional fine-tuning. Under the traditional logic, each customer gets its own copy of the model, and deployment costs are extremely high. Under multi-LoRA, one base model stays resident in GPU memory while hundreds of customer adapters are mounted on top of it, hot-swapped at the request level. The same fleet of GPUs serves a large number of customers’ “dedicated models,” and utilization is very high.)
On cost: Fireworks CEO Lin Qiao says that at equivalent quality, the cost is 5–10x lower than closed models; Decagon, a Together customer, says its costs after migrating fell to between one-fifth and one-seventh of the closed-model alternative.
The capability gap has also narrowed to a point where it can be negotiated with. Stanford’s 2026 AI Index puts the open–closed gap at 3.3 percentage points. When capability differs by a few points and price differs by a multiple, a large share of workloads will move to open-weight models — though the hardest ones will not.
Inference providers are not the SaaS of a new era. Ordinary SaaS can carry gross margins above 70% because marginal cost is near zero: one more user means a bit more server and bandwidth, not a linear increase in core cost. Inference providers face linear variable costs and cannot amortize fixed overhead through scale. Double the tokens served and the GPU bill roughly doubles too. Sacra estimates Fireworks’ gross margin at around 50%, against a company target of 60% — below SaaS’s 70%+, because GPU cost lands directly in COGS (cost of goods sold, i.e., the direct cost of delivering the product). Margin improvement therefore has to come from three directions: technical efficiency, moving up into customization, and moving down to lock in compute supply.
Progress on the technical efficiency side is genuinely remarkable. Per SemiAnalysis and NVIDIA, Blackwell delivers roughly 30x the tokens/sec/GPU of Hopper a year earlier on frontier inference workloads, and the cost per million tokens is falling by order-of-magnitude increments annually (EpochAI’s figure is about 10x per year). But compute prices are rising at the same time. In April 2026, spot rental for H100s had risen from about $1.70/hour last October to $2.35; the Blackwell spot index went from $2.75 to $4.08/hour in two months. Unit compute prices are rising, but throughput is rising faster, so cost per token continues to fall. This means that even as competition keeps pushing token prices down, inference providers’ gross margins need not compress — as long as costs fall faster than prices. That said, Sacra’s risk assessment of Fireworks notes that if open-source frameworks like vLLM and SGLang close the performance gap, what remains is GPU resale at 50% margins.
Circular capital, courtesy of NVIDIA
NVIDIA sells the cards and also invests in the people buying them. It is a shareholder in Fireworks, Baseten, and Together; it holds roughly 6% of CoreWeave and added a $2 billion private placement at $87.20 per share in January 2026 — and CoreWeave’s single largest expense, by order of magnitude, is buying NVIDIA GPUs. The money NVIDIA puts out comes back as demand for NVIDIA chips. So part of the demand on NVIDIA’s income statement is generated by its own capital rather than being fully exogenous market demand.
The hyperscalers’ three-layer business
Hyperscale clouds operate at three layers:
Layer one: bare GPU rental (IaaS). The customer rents an instance with eight H100s or B200s and installs the OS, drivers, inference framework, and model themselves, billed by the hour. AWS P5 and the Azure ND series sit here.
Layer two: managed model deployment (PaaS). The customer uploads or selects a model, the platform runs it and hands back an endpoint, with autoscaling, monitoring, and operations included. SageMaker and Vertex sit here.
Layer three: model-as-API (billed per token). The customer never sees a GPU and pays only for input and output tokens. Bedrock, Azure AI Foundry, the OpenAI API, the Anthropic API, and the Gemini API sit here.
Each layer up generally produces higher revenue and higher margin from the same GPU. But the clouds’ positions are complicated.
Microsoft is OpenAI’s largest partner: tied to OpenAI on one side, developing its own MAI and the open-weight Phi family on another, and simultaneously selling everyone’s models on Azure rather than betting on a single one. Google has both Gemini and Gemma; it wants to sell closed APIs and also wants Vertex to be the platform where enterprises deploy every kind of model. Amazon is the subtler case. It has no genuinely first-tier frontier model of its own, and it shut down its AGI lab on July 22. Bedrock’s marquee offering has long been Anthropic. Amazon’s cumulative actual investment in Anthropic has reached roughly $13 billion, with up to another $20 billion tied to commercial milestones, for a cap of about $33 billion; in return, Anthropic has committed to spend more than $100 billion with AWS over the next decade and gets access to up to 5GW of Trainium capacity. Widespread open weights would erode Bedrock’s Anthropic-centered differentiation — but Amazon is also an IaaS provider and an aggregator, which is why, late as it was, it signed in the end.
Fragmentation is a survival condition for neoclouds
CoreWeave, Lambda, Nebius, and Crusoe — the neoclouds — operate at layer one: buying GPU clusters at scale, typically financed with debt, then renting out compute on long-term contracts. What separates them from the big three clouds is that they have only GPUs and no full cloud product line — no databases, no decades of enterprise customer relationships, no complete set of compliance certifications. So they can only compete on the price, delivery speed, and availability of raw compute. Open weights expand the number of entities that need to run their own GPUs, which is very much to the neoclouds’ benefit.
But while fragmentation is a long-term survival condition for neoclouds, it is not the current state of affairs. CoreWeave’s FY2025 annual report shows Microsoft alone accounting for roughly 67% of revenue (62% in FY2024), with no second customer above 10%; in Q1 2026 the top two customers together came to about 65%. Its contracted backlog is heavily concentrated in OpenAI (roughly $22.4 billion in cumulative commitments) and Meta (roughly $35.2 billion). At the same time it carries tens of billions of dollars in debt and 2026 capex guidance of $31–35 billion. In that structure, a single largest customer declining to renew could leave it insolvent.
On top of that, its biggest customers are becoming its competitors. Meta, for example, is standing up a cloud business called Meta Compute to sell the surplus compute from its $115–145 billion of 2026 capex, in forms including model access through Muse Spark and raw GPU cycles. SpaceX’s Colossus site was leased in 2026 to Anthropic (roughly $45 billion through mid-2029), Google (roughly $30 billion), and Reflection ($6.3 billion). And OpenAI’s Stargate is the build-it-yourself path itself.
Neoclouds hold no model assets and have nothing to win or lose at the model layer. What they can do is find enough buyers to fill capacity they have already bought with debt — an idle GPU is pure loss. Open-weight models expand the number of long-tail buyers (enterprises, sovereign AI programs, AI application companies, vertical model companies, inference platforms, research institutions), which is the most direct route to de-concentrating their customer base.
Enterprise bargaining power
Enterprises have data concerns, but the overwhelming majority of small and mid-sized companies will not actually buy cards and run models themselves.
The reason is TCO (total cost of ownership). Self-hosting means not just buying GPUs but building and maintaining the infrastructure and employing the engineers to do it. Then there is utilization: an API is a purely variable cost — no calls, no spend — while owned GPUs are a fixed cost billed by the hour. Enterprise traffic is typically high during the day and low at night, so average utilization may be poor. A common industry rule of thumb is that self-hosting only becomes worth discussing once annual API spend reaches the seven-figure range — and that is before accounting for operational risk, hiring difficulty, and opportunity cost.
The most important value of open weights to enterprises comes down to three things:
(1) Portability. No lock-in to a single API. If a vendor changes prices, retires a model, or changes its data policy, the enterprise has a migration path.
(2) Auditability. Holding the weights yourself means the model will not be swapped out or degraded without your knowledge, and in compliance settings you can reproduce the behavior of one specific version.
(3) Negotiating leverage. Even if the enterprise ends up using a closed API anyway, having an open alternative that is good enough for most tasks puts it in a completely different negotiating position. There is no stable relationship between a closed model’s marginal inference cost and its list price; list price is set mainly by competitors, substitutes, and willingness to pay. The existence of open-weight models effectively caps what closed models can charge.
For an enterprise, if a single vendor controls the model, the pricing, the access, and the institutional knowledge the company has accumulated, then that vendor increasingly controls the company’s business.