Cheaper Per Task, Costlier In Total
An Evaluation of Four Claims About AI Infrastructure Decentralization
Executive summary
- Claim One, data centers are a bubble consumer hardware will replace: rejected. Compute bifurcates by workload. Training requires continuous multi-GPU synchronization no consumer network can meet; inference does not, and is genuinely decentralizing.
- Claim Two, decentralizing AI saves energy: not supported by direct evidence, though probably true in aggregate. Capacity-factor economics and PUE data favour centralization, but no study directly comparing energy per query on consumer versus data-center hardware was located.
- Claim Three, consumer chips keep getting cheaper: rejected. TSMC’s newest leading-edge node is the first major node transition to raise cost per transistor rather than lower it.
- Claim Four, computing history proves decentralization is inevitable: rejected as applied to AI training, partially supported in narrower form. The historical pattern holds for independent, low-interconnect workloads only. No precedent exists for a large synchronized job decentralizing.
- Conclusion: a bifurcation, not a resolution. Training stays centralized on physical grounds; inference decentralizes on economic grounds; the hope that this happens at falling cost and falling total energy is not currently supported.
1Introduction
Nvidia markets its DGX Spark, released in late 2025 at $4,699, as a “personal AI supercomputer”: a Grace Blackwell superchip, 128GB of unified memory, roughly one petaflop of AI compute at FP4 precision, built on a 3-nanometre process, small enough to sit on a desk (Nvidia, 2025). In February 2026, its price rose. The cause was not a change in ambition but a shortage: memory became scarcer and more expensive across the entire AI supply chain, and a $4,699 desktop device turned out to draw on the same constrained inputs as a data center campus costing four orders of magnitude more (IntuitionLabs, 2026).
That fact is a useful entry point because it undercuts the framing on both sides of the public argument at once. The “personal AI computer” is not a separate, insulated market whose falling costs will eventually outcompete centralized infrastructure. It buys from the same supply chain, under the same constraints, at the same time.
Nvidia DGX Spark. Repriced upward in Feb. 2026 by a memory shortage, not despite one.
Nvidia (2025); IntuitionLabs (2026)
Combined 2026 capex, five largest US cloud/AI infrastructure providers.
Futurum Group (2026a); Uncoveralpha (2026)
This paper evaluates four claims that recur, individually or in combination, whenever this subject is discussed publicly: that the data center buildout is a bubble consumer hardware will correct; that moving AI processing to personal devices would reduce energy use; that consumer chips will keep getting cheaper, carrying edge AI down in cost with them; and that computing history shows decentralization is simply how this kind of transition ends. Each is treated as a falsifiable proposition and tested against publicly available evidence. Section 2 situates these claims against prior work. Section 3 sets out the method and its limits. Section 4 presents the findings, claim by claim. Section 5 discusses what the four verdicts together imply. Section 6 states the limitations of this analysis explicitly. Section 7 concludes with operational recommendations.
2Context and prior work
Two narratives currently compete in public commentary on AI infrastructure. One holds that the scale of investment is a bubble: that consumer and edge hardware is advancing quickly enough to make centralized infrastructure unnecessary within a few years, comparable to how excess fibre-optic capacity from the late 1990s is now remembered. The other holds that the buildout is a rational response to a technology that is genuinely compute-bound, and that skepticism underestimates both the physical requirements of training frontier models and the empirical track record of scaling laws. This paper does not adopt either narrative as a starting position. It isolates the specific claims each narrative depends on and evaluates them separately, because a narrative can survive the failure of one supporting claim and fail on the strength of another, and collapsing them together obscures which is which.
Three bodies of prior work inform the evaluation that follows.
The first is the economic history of computing centralization. Grosch’s Law, the mid-twentieth-century observation that a mainframe costing three times as much delivered roughly nine times the computational throughput, describes why computing centralized during the mainframe era: bigger machines were cheaper per unit of compute (historyofcomputercommunications.info, n.d.; ethw.org, n.d.). This is the same economic logic now pulling GPU clusters together, and it did not disappear when personal computers arrived; Section 4.4 examines what actually changed.
The second is the technical literature on neural network scaling. Kaplan et al. (2020) established that model performance follows predictable power-law relationships with model size, dataset size, and compute. Hoffmann et al. (2022), publishing as Chinchilla, refined this into a compute-optimal relationship between model size and training-token count. Neither paper addresses infrastructure location directly, but both are load-bearing for Claim One, because they describe why more compute continues to buy more capability rather than hitting diminishing returns that would make further centralization uneconomical.
The third is this publisher’s own prior work. Greener Per Unit, Browner In Total (Äctvli Responsible Consulting, 2026) established a measurement-boundary method: that an efficiency metric measured within a narrow boundary can be genuinely accurate while being systematically misleading about the total system it sits inside, using EU appliance efficiency regulation and global e-waste growth as the case study. That method is the analytical spine of Claim Two in this paper, applied to AI compute rather than household appliances.
3Methodology and scope
This paper synthesizes publicly available evidence, industry disclosures, and existing technical and economic research to evaluate four specific claims about the future of AI compute infrastructure. It does not present new experimental data, benchmarks, or primary surveys, and that limitation is stated here rather than left implicit.
Research was conducted across two parallel tracks during the first week of September 2026: an infrastructure-economics track, covering hyperscaler capital expenditure, data center energy demand, power usage effectiveness, and the physical requirements of distributed model training; and an edge-compute track, covering consumer chip performance trends, local inference tooling, the historical record of computing decentralization, and consumer and enterprise AI usage data. All findings reflect a snapshot as of Q3 2026, specifically 2026-09-06, and should be understood to age accordingly. A document evaluating a fast-moving industry is a snapshot, not a permanent record, and Section 7 recommends a fixed schedule for revisiting its conclusions rather than treating them as settled.
Sourcing standard applied throughout: a claim is presented as established only where it is corroborated by a named, checkable source. Where a widely circulated figure could not be traced to a primary source, most notably a specific utilization-rate comparison between data centers and home computers that appears frequently in informal commentary, it is explicitly flagged as unverified in the relevant findings section and excluded from this paper’s conclusions rather than presented as fact. Where sources disagree, for example on individual hyperscaler capital expenditure figures, which vary by fiscal versus calendar year and by gross versus net capex accounting, the range is reported and the disagreement is noted rather than resolved by selecting the most convenient figure.
This method has a structural limitation worth stating plainly. A synthesis of existing evidence can evaluate claims that have already been made publicly; it cannot generate new primary evidence for questions the existing literature has not yet answered. Section 6 identifies where this paper’s conclusions rest on directional inference rather than direct measurement. The most significant instance is Claim Two: no study directly comparing energy consumption per inference query on consumer hardware against data center hardware could be located, and this paper does not manufacture one to fill the gap.
4Findings
4.1Claim One: Data center investment is a bubble that consumer hardware will replace
Combined hyperscaler capital expenditure and improving consumer AI hardware together mean the current data center buildout is overbuilt, soon-to-be-stranded capacity that personal and departmental AI hardware will absorb.
The evidence. Combined 2026 capital expenditure across the five largest US cloud and AI infrastructure providers, Microsoft, Amazon, Alphabet, Meta and Oracle, is running at roughly $690 to $725 billion, up from approximately $410 billion in 2025 (Futurum Group, 2026a; Uncoveralpha, 2026). Analysts covering Q2 2026 earnings project the combined figure exceeding $1 trillion in 2027, with one bank’s chief executive citing $1.3 trillion for 2027 and a possible $1.5 trillion for 2028 (Motley Fool, 2026). Amazon states it expects capacity to trail customer demand through both 2026 and 2027 and is targeting a doubling of data center power capacity by the end of 2027 (Uncoveralpha, 2026). Google Cloud’s contract backlog has roughly doubled year over year to approximately $460 billion (Uncoveralpha, 2026). Oracle exceeded its own capital expenditure guidance for the year and required a $16.3 billion financing round, with PIMCO alone anchoring $10 billion, after US banks reduced their exposure to the deal (TheNextWeb, 2026a; 2026b). OpenAI’s Stargate project targets approximately $500 billion in infrastructure investment by 2029 (Futurum Group, 2026a). Full figures by organization appear in Appendix A.
Microsoft, Amazon, Alphabet, Meta and Oracle combined, US$ billions. 2027 is the projected midpoint of a $1.0–1.3T analyst range.
This spending pattern is consistent with a physical constraint rather than a speculative one. Training a frontier model is a single, tightly synchronized computation distributed across many GPUs simultaneously. A 70-billion-parameter model requires approximately 140GB to hold its weights in mixed precision alone; including optimizer states and activations, the practical requirement runs to 400 to 600GB, necessarily spread across multiple GPUs (Together AI, n.d.). Within a server, GPUs exchange gradients over NVLink at approximately 900 GB/s; between servers, that exchange runs over InfiniBand at approximately 400 to 800 Gb/s, orders of magnitude slower, which is why cluster network topology and physical GPU placement materially affect training throughput (NADDOD, n.d.; RunPod, n.d.). Every training step requires this synchronized exchange, and network latency converts directly into idle GPU time (arXiv:2501.04266, 2025). Residential broadband sits many further orders of magnitude below InfiniBand than InfiniBand sits below NVLink, which is why no number of internet-connected consumer GPUs can substitute for a physically co-located cluster for this specific workload.
This distinction is not hypothetical, and it cuts against a common misreading of this claim worth addressing directly: the physical requirement above applies to training, not to the finished model that training produces. Once a model’s weights are fixed, they are portable. Models trained in clusters of exactly this kind already run on ordinary consumer hardware: Llama 3.1 8B and Mistral Small 3, for instance, run at practical speeds on mid-range consumer GPUs once quantized to 4-bit precision (corelab.tech, 2026), an illustrative rather than rigorously benchmarked figure, but not a disputed description of current practice. The data center is where a model is built; it is not where the finished model is required to remain.
Separately, the scaling-law literature indicates that this physical requirement will not become less binding as models improve. Kaplan et al. (2020) and Hoffmann et al. (2022) together establish that, for a fixed compute budget, model capability continues to scale with both model size and training data, rather than plateauing. Epoch AI’s 2026 analysis of compute-scaling trends estimates that the current pace of approximately 4x per year training-compute growth could plausibly continue through the decade if power, chip supply, and data availability do not become binding constraints first, implying training runs on the order of 2×1029 FLOP are feasible by 2030 (Epoch AI, 2026).
Capital allocation behaviour is also informative. All four major hyperscalers follow the same pattern: commit early to land, building shells, and power interconnects, long-lived and comparatively low-risk assets, while delaying the purchase of GPUs themselves until close to the point of deployment (Uncoveralpha, 2026). This is a textbook hedge against exactly the overbuild risk the bubble narrative describes, and its consistency across four independent organizations is evidence that the risk is understood and managed rather than ignored.
4.2Claim Two: Moving AI processing to personal devices reduces total energy consumption
If inference workloads move from data centers to personal and departmental hardware, total energy consumption for AI falls, because centralized infrastructure carries overhead that individually owned devices do not.
The evidence. The International Energy Agency’s April 2026 update reports that global data center electricity consumption grew 17% in 2025, with the AI-attributable share growing 50% in the same year, and projects total data center electricity consumption roughly doubling by 2030, from approximately 485 TWh to approximately 950 TWh, with the AI-attributable portion tripling over the same period (IEA, 2026a). The same report states that power consumption per AI task is falling at a pace the agency itself describes as historically unprecedented, while total demand continues to rise because usage volume and task complexity are growing faster than per-task efficiency gains reduce consumption (IEA, 2026b). This is structurally a Jevons paradox, and it is the same shape of finding this publisher’s companion piece documented for EU appliance efficiency regulation: a measured metric improves exactly as intended while the total system it sits inside moves the other way (Äctvli Responsible Consulting, 2026).
Terawatt-hours per year. 2030 is the IEA’s projection, not an observation.
Evidence bearing directly on the centralization question is more limited. Hyperscaler-reported power usage effectiveness (PUE) runs at approximately 1.09 to 1.2 at leading sites, compared with an industry-wide average of 1.52 in 2026 (Google, 2026; Uptime Institute, 2026). Uptime Institute’s multi-year analysis finds that larger facilities are consistently more efficient than smaller ones, attributable to newer construction, more advanced cooling, and better operational controls (Uptime Institute, 2024). Germany’s Energy Efficiency Act now mandates that new data centers achieve a PUE of 1.2 or better from July 2026, indicating that regulators do not consider current industry-average efficiency levels to be a physical ceiling. The underlying economic logic parallels the case for centralized electricity generation over distributed home generators: pooled, scheduled hardware achieves higher utilization than idle hardware distributed across millions of separate locations, and idle hardware continues to draw standby power regardless of whether it performs useful work.
No study directly measuring energy consumption per inference query on consumer hardware against energy consumption per inference query on batched data center hardware running the same model could be located in this research. This comparison, energy per token, consumer versus data center, is the specific evidence this claim requires, and it does not currently exist in a verifiable form. A statistic circulating informally, claiming data centers operate at approximately 25% utilization against 1 to 3% for distributed personal computers, could not be traced to a primary source and is excluded from this paper’s evidentiary basis.
4.3Claim Three: Consumer chip costs will keep falling, making edge AI increasingly cheap
Consumer AI hardware will continue on the historical Moore’s Law trajectory, becoming more capable and cheaper over time, progressively eroding the cost advantage of centralized infrastructure.
The evidence. On-device AI capability is improving quickly by several measures. Apple’s Neural Engine rose from 11 TOPS in the M1 (2020) to 38 TOPS in the M4 (2024), approximately 3.5 times over three generations, and the M5 generation adds dedicated Neural Accelerators within each GPU core, which Apple states delivers over four times the prior generation’s peak AI compute (Apple, 2025). Qualcomm’s Snapdragon NPU rose from 45 TOPS (X Elite, 2024) to 80 TOPS (X2 Elite, 2026) in approximately eighteen months (TechRepublic, 2026). This gain is uneven across vendors: AMD’s newest-generation flagship NPU shipped flat to marginally down against its own prior generation over a comparable period (Notebookcheck, 2026), and a widely followed hardware commentator now argues that raw TOPS is no longer the binding constraint, with memory bandwidth determining practical performance more than peak compute (LocalAIMaster, 2026). Full figures appear in Appendix C.
TOPS (trillion operations per second), vendor-published Neural Engine / NPU figures.
The capability trend, however, is not the trend this claim depends on. The claim depends on cost per unit of compute continuing to fall, and the most current evidence located in this research indicates the opposite is now occurring at the leading edge of chip manufacturing. TSMC’s newest 2-nanometre (N2) node prices wafers at approximately $30,000, more than 50% above the approximately $20,000 price of the prior 3-nanometre node, and the following node, expected in the second half of 2026, is rumoured at approximately $45,000 per wafer (EE Times, 2026). This is the first major node transition in TSMC’s history to raise cost per transistor rather than lower it. The driver is not purely commercial pricing power: at this scale, transistors require gate-all-around architecture to control current leakage, a genuine physical constraint on further scaling rather than only a business decision (EE Times, 2026).
US$ per wafer. For the first time at a major node transition, cost per transistor is rising rather than falling.
Two further findings bear on this claim specifically. Mixture-of-experts architectures are widely cited as a technique that should make edge inference more efficient, because sparse activation implies less compute per query. An empirical study testing this claim directly on real edge hardware found the theoretical efficiency gains do not consistently materialize, because edge devices are bandwidth-bound rather than compute-bound, and MoE inference cost tracks total parameter count, which must still be held in memory, rather than the smaller active-parameter count typically quoted (Alfarizy et al., 2026). And self-hosting economics, while genuinely favourable at high utilization, are sharply utilization-dependent rather than universally favourable: one industry cost analysis finds a 13-billion-parameter model reaches cost parity with premium API pricing at just 10% utilization, but at that same utilization level, the idle cost of owning the hardware raises the effective per-token cost by approximately a factor of ten, exceeding the API price it was meant to undercut (Introl, 2026).
4.4Claim Four: The history of computing shows this always ends in decentralization
Mainframe computing gave way to minicomputers, which gave way to personal computers; centralized computing paradigms are always eventually displaced by cheaper, distributed alternatives, and AI should follow the same trajectory.
The evidence. The historical record supports a narrower claim than the one usually drawn from it. Grosch’s Law described the mainframe era accurately: a machine costing three times as much delivered roughly nine times the throughput, so computing centralized because centralizing was cheaper per unit of work (historyofcomputercommunications.info, n.d.). This economic relationship was not repealed by the arrival of minicomputers or personal computers. What changed is that the absolute cost of an individually owned “good enough” machine eventually fell below the threshold at which sharing a larger one was worth the coordination overhead, specifically for independent, low-interconnect, single-user workloads: word processing, spreadsheets, and, in the present case, individual AI inference queries (ethw.org, n.d.).
What did not decentralize, in this period or since, is the other category of workload: large, tightly synchronized computational jobs. These remained in supercomputing centers, and subsequently in cloud data centers, for the same Grosch’s-Law-consistent economic reason they always had. This research found no historical case, across seventy years of computing, of a large synchronized computational workload successfully migrating to independent, consumer-owned machines. Timesharing systems in the 1960s and 1970s decentralized access to a mainframe among many users; they did not decentralize the underlying compute, which is a different claim (historyofcomputercommunications.info, n.d.). Digital Equipment Corporation’s own history offers a further data point: DEC dominated the minicomputer era but was displaced by the subsequent wave of decentralization in large part because it retained a proprietary architecture rather than adopting the open, commodity hardware standard that allowed personal computer costs to keep falling (Computer History Museum, n.d.). Today’s AI accelerator market, built substantially around Nvidia’s proprietary CUDA software stack, resembles the mainframe-era pattern of vendor lock-in more closely than it resembles the open commodity hardware transition that enabled the PC era; this paper treats that observation as an open question rather than an established finding, since no source located in this research draws the comparison directly.
5Discussion
The four verdicts converge on a single structural finding rather than four unrelated conclusions. AI compute is not moving in one direction; it is separating along a line defined by whether a workload requires synchronized coordination across many processors (training, which stays centralized on physical grounds established in Section 4.1 and consistent with the historical pattern in Section 4.4) or does not (inference, which is decentralizing on economic grounds, though not at the falling-cost trajectory Claim Three assumed, and not with a demonstrated energy saving per Claim Two).
This finding sits uncomfortably with both public narratives described in Section 2. It rejects the bubble narrative’s central prediction, that consumer hardware displaces centralized infrastructure broadly, while confirming one of its subsidiary observations, that a meaningful and growing share of AI workloads genuinely is moving to the edge. It confirms the centralization narrative’s core physical and economic argument for training, while withholding confirmation of the stronger claim, implicit in much industry commentary, that the resulting energy and cost trajectory is self-evidently sound. This publisher’s companion piece put a related asymmetry plainly: an organization can be entirely honest about what it measures and still be wrong about what that measurement proves (Äctvli Responsible Consulting, 2026). The same discipline applies to industry narratives, and it is why this paper evaluates each claim independently rather than adjudicating a single verdict on “the buildout.”
The strongest objection to this paper’s conclusions is that capital markets, not historical analogy or physical constraint, will ultimately settle the question, and capital markets have already reached a verdict by funding the buildout at the scale documented in Section 4.1. This objection has real force: the hedging pattern documented there, land and power committed early, chip purchases delayed until near deployment, is evidence that sophisticated capital allocators are aware of the overbuild risk and are pricing it in rather than ignoring it. It is not, however, evidence that the resulting scale of investment is correctly sized, only that the people spending it are managing a known risk. Whether the installed capacity is monetized quickly enough to justify its cost is a financial question this paper does not attempt to answer, and a reader should not read Claim One’s rejection of the bubble framing as an endorsement of current spending levels being optimal.
A second objection concerns Claim Three. On-device chip capability, as distinct from chip cost, continues to improve meaningfully (Section 4.3), and a reader could argue this paper understates the decentralization trend by focusing on a single leading-edge manufacturing data point. This is a fair challenge to the claim’s strength but not to its substance: capability gains sourced from more transistors, more efficient architectures, or better software do not require falling cost per transistor to continue, but the specific claim under evaluation, that edge AI becomes cheaper, does. A chip that is more capable but not cheaper supports a different and more modest claim than the one this paper set out to test.
A further observation, adjacent to the four claims this paper set out to test rather than a fifth one, is worth stating plainly because it explains why the bifurcation in Section 4.1 is likely durable rather than transitional. “Frontier” capability is not a fixed compute threshold that consumer hardware could eventually reach and settle into; it is defined relatively, as whatever the most capable model happens to be at a given time. If consumer hardware capability increases by some factor, the hardware available to an organization training a frontier model increases by the same factor, and can additionally be networked together at a scale no individual consumer device can replicate. The gap this paper documents is therefore not obviously one that improving consumer hardware closes over time; it may instead persist by construction, because the target moves with the hardware rather than staying fixed ahead of it.
This raises a question the evidence gathered here can only partially answer: whether continued frontier investment is justified by need, as opposed to competitive necessity. The scaling-law evidence reviewed in Section 4.1 (Kaplan et al., 2020; Hoffmann et al., 2022) and Epoch AI’s applied forecasting (Epoch AI, 2026) both indicate that scaling continues to improve measured model capability, with no plateau yet observed in that research. That is a distinct claim from whether most practical applications of AI require frontier capability, however, and the evidence reviewed in Section 4.3 suggests most do not, since quantized, locally-runnable models already serve a wide range of tasks adequately. The continued pursuit of frontier capability by a small number of well-resourced organizations, despite that gap, is more plausibly explained by competitive dynamics in a winner-concentrated market and a residual set of applications, long-horizon reasoning and frontier scientific research chief among them, where capability gains have not yet visibly saturated, than by broad-based demand for frontier capability itself. This paper did not set out to evaluate that explanation directly and holds no primary evidence for it either way; it is offered here as a plausible account consistent with the evidence gathered, clearly marked as interpretation rather than as a fifth tested claim.
6Limitations
This analysis has five limitations that materially affect how its conclusions should be used.
7Conclusion and recommendations
The evidence does not support a straightforward answer to “will AI infrastructure decentralize,” because the question conflates two workloads with different physics and different economics. Training remains centralized because synchronized multi-GPU computation over the necessary bandwidth is not achievable over consumer networks, a constraint that does not soften as chip performance improves. Inference is genuinely decentralizing, consistent with the historical pattern by which independent, low-interconnect workloads have always eventually moved to individually owned hardware once that hardware became cheap enough. What is not currently supported by evidence is the specific hope that this transition saves energy at the system level or that it is happening on a falling-cost trajectory: the semiconductor cost curve the “cheap edge AI” thesis depends on reversed for the first time in decades at the exact moment the thesis needs it to continue falling.
Few organizations train frontier models, and the physical argument in Section 4.1 is not relevant to their infrastructure decisions. Every organization deploying AI in production, however, faces a version of the workload-classification problem this paper has applied at industry scale, and the same discipline applies at organizational scale. Three questions do more useful work than a general stance on cloud versus on-device infrastructure:
First, does the workload require frontier-model capability, or would a smaller, locally-runnable model perform adequately? Most internal tooling, drafting, and classification tasks do not require the largest available model, and quantized local models are now a genuinely mature option for them, per the evidence reviewed in Section 4.3.
Second, is demand for the workload sustained and predictable, or bursty and occasional? The cost-parity case for self-hosting established in Section 4.3 holds only at meaningful utilization; provisioning dedicated hardware for infrequent demand is close to the worst-case cost scenario documented in this paper, not the best one.
Third, does the data carry a regulatory, contractual, or latency constraint that prevents it from leaving the device? This is the one case in which the cost comparison is not the deciding factor, because no cloud alternative satisfies the constraint at any price.
Finally, this paper’s conclusions should be revisited on a schedule rather than assumed to hold indefinitely, and two specific indicators are worth monitoring because they are the findings in this analysis most likely to change: whether TSMC’s cost-per-transistor trend at the leading edge reverses back toward its historical decline, which would materially strengthen Claim Three, and whether a direct, methodologically sound comparison of energy consumption per inference query on consumer versus data center hardware is published, which would resolve Claim Two in one direction or the other. Neither indicator had moved as of this paper’s research snapshot.
References
Äctvli Responsible Consulting (2026) Greener Per Unit, Browner In Total: The Measurement Boundary Problem. Available at: https://www.actvli.com/insights/greener-per-unit-browner-in-total (Accessed: 6 September 2026).
Alfarizy, A. et al. (2026) ‘Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study’, arXiv preprint, arXiv:2606.21428.
Apple (2025) Apple unleashes M5, the next big leap in AI performance for Apple silicon. Apple Newsroom, October 2025.
apxml (n.d.) High-Bandwidth Interconnects.
arXiv:2501.04266 (2025) Scaling LLM Training on Frontier with Low-Bandwidth Partitioning.
Computer History Museum (n.d.) What the DEC records of minicomputer giant Digital Equipment Corporation, open for research at CHM, reveal.
Congress.gov Congressional Research Service (n.d.) Data Centers and Water FAQ.
corelab.tech (2026) Best GPU for Local LLMs (Late 2026).
dasroot.net (2026) GGUF quantization quality and speed on consumer GPUs.
eciks.org (2026) CoreWeave Capex 2026 Guidance: $30–35 Billion.
EE Times (2026) TSMC price hikes end the era of cheap transistors.
Epoch AI (2026) Analysis of AI training-compute scaling trends and constraints, 2026. Epoch AI.
ethw.org (n.d.) Rise and Fall of Minicomputers. Engineering and Technology History Wiki.
Futurum Group (2026a) AI Capex 2026: The $690B Infrastructure Sprint.
Futurum Group (2026b) Snapdragon X2 Elite Pushes AI PC Performance to New Heights.
Google (2026) Data Centers: Efficiency.
historyofcomputercommunications.info (n.d.) Minicomputers, Distributed Data Processing and Microprocessors.
Hoffmann, J. et al. (2022) ‘Training Compute-Optimal Large Language Models’, arXiv preprint, arXiv:2203.15556. DeepMind.
IEA (2026a) Data centre electricity use surged in 2025, even with tightening bottlenecks driving a scramble for solutions. International Energy Agency, April 2026.
IEA (2026b) Key Questions on Energy and AI: Executive Summary. International Energy Agency.
IntuitionLabs (2026) NVIDIA DGX Spark review.
Introl (2026) Inference Unit Economics: True Cost per Million Tokens Guide.
Kaplan, J. et al. (2020) ‘Scaling Laws for Neural Language Models’, arXiv preprint, arXiv:2001.08361. OpenAI.
Lincoln Institute of Land Policy (2026) Data Drain: The Land and Water Impacts of the AI Boom. Land Lines Magazine.
LocalAIMaster (2026) Best NPU for AI 2026.
Motley Fool (2026) Hyperscalers Driving AI Capex Toward $1.3 Trillion in 2027. 6 September 2026.
NADDOD (n.d.) GPU Cluster Network Configurations for LLM Training.
Notebookcheck (2026) AMD launches Ryzen AI 400 desktop processors with up to 50 TOPS NPU and Copilot support.
Nvidia (2025) NVIDIA Announces DGX Spark and DGX Station, Personal AI Computers. Nvidia Newsroom.
Ramp / GreyJournal (2026) AI Index: How much companies spend on AI tokens in 2026.
RunPod (n.d.) Do I need InfiniBand for distributed AI training?
Stanford HAI (2026) 2026 AI Index Report: Economy. Stanford Institute for Human-Centered AI.
TechRepublic (2026) Qualcomm Snapdragon X2 Elite Extreme announcement.
TheNextWeb (2026a) Oracle Q4 FY2026 capex.
TheNextWeb (2026b) Oracle’s $16.3 billion data center financing.
Together AI (n.d.) Inside multi-node training.
Uncoveralpha (2026) Amazon, Google, Microsoft, Meta Q2 2026 earnings.
Uptime Institute (2024) Large data centers are mostly more efficient, analysis confirms. Uptime Institute Journal, 7 February 2024.
Uptime Institute (2026) 16th Annual Global Data Center Survey.
Water Foundation (2026) Water & Data Centers.
a16z (2026) Top 100 Gen AI Consumer Apps. Andreessen Horowitz, March 2026.
index.dev (2026) AI growth statistics by country.
Appendices
Appendix A: Hyperscaler 2026 capital expenditure by organization
| Organization | 2026 capex / commitment | Notable detail |
|---|---|---|
| Amazon / AWS | ~$220B | Capacity to trail demand through 2026–2027; power capacity to double by end-2027 |
| Alphabet / Google | $195–205B | Cloud contract backlog ~$460B, roughly double prior year |
| Meta | $125–145B | Guidance raised twice during 2026 |
| Microsoft | $120–190B* | *Figures diverge by source; see Section 6 |
| Oracle | $55.7B (FY26 actual) | $70B FY27 guidance; required $16.3B financing, PIMCO anchoring $10B |
| CoreWeave | $31–35B | Targeting 1.7GW active power capacity by end-2026 |
| OpenAI Stargate | ~$500B by 2029 | ~7GW planned capacity across Stargate and partner sites |
Sources: Futurum Group (2026a); Uncoveralpha (2026); TheNextWeb (2026a; 2026b); eciks.org (2026).
Appendix B: Land and water use, US data centers
| Metric | Value | Source |
|---|---|---|
| Typical hyperscale campus footprint | 200–1,000+ acres | Lincoln Institute of Land Policy (2026) |
| Typical individual site power draw | 100MW+ | Lincoln Institute of Land Policy (2026) |
| Direct US data center water use, per year | 17–19B gal | Water Foundation (2026) |
| Indirect water use via electricity generation, 2023 | ~211B gal | Congress.gov CRS (n.d.) |
Appendix C: On-device AI chip performance by vendor and generation
| Vendor | Prior generation | Current generation | Change |
|---|---|---|---|
| Apple Neural Engine | 11 TOPS (M1, 2020) | 38 TOPS (M4, 2024) | ~3.5x over 3 gens |
| Qualcomm Snapdragon NPU | 45 TOPS (X Elite, 2024) | 80 TOPS (X2 Elite, 2026) | ~1.8x in ~18 months |
| AMD Ryzen AI NPU (flagship) | ~60–75 TOPS combined | 60 TOPS (HX 475) | Flat to down |
Sources: Apple (2025); Hoxton Macs (n.d.); TechRepublic (2026); Notebookcheck (2026). Apple’s M3/M5 Neural-Engine-only figures excluded as contested; see Section 6.
