About this transcript: This is a full AI-generated transcript of The Entire AI Data Center Explained — From Electricity to ChatGPT from Leo Cui, Ph.D., CFA , published July 24, 2026. The transcript contains 7,074 words with timestamps and was generated using Whisper AI.
"Last night, sometime around 7 p.m., you pull out your phone, you type a question, maybe it was, what should I make for dinner with chicken and rice? And about two seconds later, a machine wrote you an answer, two seconds. That's what I want to do in this video. I want to slow those two seconds..."
[00:00:00] Speaker 1: Last night, sometime around 7 p.m., you pull out your phone, you type a question, maybe it was, what should I make for dinner with chicken and rice? And about two seconds later, a machine wrote you an answer, two seconds. That's what I want to do in this video. I want to slow those two seconds down, way down, because in those two seconds, your questions left your phone, traveled hundreds of miles through strands of glass thinner than human hair, and arrived at a building the size of several football fields, a building that drinks as much electricity as a small city. Inside that building, your question passed through a machine that cost as much as a house, got translated into pure math, was processed by chips running so hot, they have to be liquid cooled like a race car engine. And then the answer came back to you, letter by letter, before you had time to lower your thumb. And here's a part that should get your attention as an investor. To make those two seconds possible, the largest companies on earth are spending, this year alone, roughly $725 billion on infrastructure. And that's just four companies, Amazon, Microsoft, Google, and Meta. And that's more in one year than the inflation-adjusted cost of the entire U.S. interstate highway system. Goldman Sachs project a total buildout at $7.6 trillion between 2026 and 2031. Jensen Huang, the CEO of NVIDIA, stood on stage at Davos this January and called it, his own words, the largest infrastructure buildout in human history. So the question this whole video hangs on is simple. Where does all the money actually go? By the end of this video, you are going to be able to answer that. You understand every single step your question takes. The power plants, the cooling systems, the chips, the memory, the fiber optics, the software. You know which companies sit at every step, which companies are printing money, which companies are telling stories, and where the whole thing could crack. I'm Leo, a VC investor. This is educational purposes, not financial advice. So before we trace your question across the country, we have to answer something more basic. The internet has existed for 30 years. Google has answered trillions of questions. Why did nobody need to spend three quarters of a trillion dollar a year until now? What changed? The answer comes down to a difference between two people, a librarian and a writer. Google is a librarian. When you search chicken rice recipe, Google doesn't cook anything. It walks into a giant library it has already organized. It indexed the whole internet years ago and keeps updating it. And it hands you pages that already exist. The expensive work happen in advance. Answering you is just a lookup. Fast, cheap, done. A Google search costs a fraction of a cent. GPT is a writer. When you ask it the same question, there's no answer sitting on a shelf. There's no database entry that says, here's what to tell this person. The model can pose your answer from scratch, one word at a time, every single time. Even if a million people ask the same question today, it doesn't retrieve, it generates. And generation is expensive. A single tried GPT query can cost 10 to 100 times more compute than a Google search. Now multiply that by 900 million weekly users. And that's the entire reason this video exists. Search retrieves, AI generates. And generation is a manufacturing process. Which brings me to the analogy I'm going to use for the rest of this video. I want you to think of an AI data center as a factory. A very strange factory. Raw material goes in one side, electricity. A product comes out the other end, words. And like any factory, it has departments, a power plant, a cooling system, assembly lines, a shipping department. We are going to tour each one. But first, three terms you need. Term one, the token. A token is the product this factory makes. Language models don't actually read words, they read tokens, which are chunks of tax, roughly three quarters of a word each. Chicken and rice is about four tokens. Your question gets chopped into tokens on the way in. And the answer gets manufactured, token by token, on the way out. And here's why investors care. The token is the token. Token are the units of revenue in the AI economy. OpenAI and Anthropic literally price their product per million tokens. When you hear token, think widget comes out the assembly line. Term two, the flop. A flop is one floating point operation, one single arithmetic calculation, one multiply or one add. It's the unit of labor in this factory. Manufacturing a single token requires a model to do hundreds of billions of these calculations. Not per answer, per word. When people say a chip does a thousand trillion flops per second, they are telling you how many workers that chip has on the factory floor. Term three, and this is a big one. Training versus inference. Training is building the factory. You take a model. Think of it as a machine with over one trillion adjusted knobs, called parameters. And you show it a huge portion of the written internet, adjusting these knobs until it gets good at predicting language. This takes a month. Tens to thousands of chips running around the clock. And on the order of a hundred million dollars or more per frontier model. It happens once per model. Inference is running the factory. Every time you ask GPT anything, that's inference. The trained model manufacturing answer for you. And here's the misconception I most want to kill in this video. People assume training is where the money goes. Wrong. By 2026, roughly two-thirds of all AI compute is inference. Because training happens once, but inference happens billions of times a day. OpenAI's inference bill alone is projected around $14 billion this year. The factory was expensive to build. It's even more expensive to run. Okay, so why did all of this suddenly explode after 2022? One discovery, the most economically important discovery of this decade, and most people have never heard of it: scaling law. In 2020, researchers at OpenAI found something almost embarrassing in its simplicity. If you make the model bigger, give it more data, and spend more compute, All together, the model gets smarter. Not sometimes. Predictably. On a chart, it's nearly a straight line. If you spend 10 times more, you get a reliably better model. And stop and think about what that means for a CEO. For 50 years, better software meant hiring smarter programmers. Scaling law turned intelligence into something you could purchase. It converted AI from a research problem into a capital expenditure problem. And big companies know exactly how to compete on capital expenditures. I'll spend everyone. That is the moment software stopped being about code and started being about concrete. That's why we suddenly need factories. Because here's the closing thought for this act. For the entire history of Silicon Valley, software was escaped from the physical world. Zero marginal cost. Infinite copies. No factory needed. AI reversed that. The frontier of software is now poured in concrete, measured in megawatts, and cooled with water. Every additional smart answer requires physical machines, physical electricity, physical heat removed. Software became heavy industry. And that changes who makes money. Let's unfreeze your question. It's 7:00 p.m. You have typed, "What should I make for dinner with chicken rice?" And you thump his sand. Here's the actual complete, no-step-skips journey. Step one, the trip. Your question leaves your phone as radio waves, hitting a cell tower or your Wi-Fi router. And within a few miles becomes pulses of light inside fiber optic cable. Glass strands carrying data at two-thirds speed of light. And it gets routed to the nearest entry point of the AI company's network. Then travel, often hundreds of miles, to a data center. Total time so far? A few hundreds of seconds. Step two, the front door. Your question arrives at API Gateway. Think of it as a factory receiving desk. It tracks who you are, checks you are not sending a thousand requests a second, runs a safety screen, and stable together everything the model needs: the system instructions, your past conversation, and your new question. Step three, tokenization. That full text gets dropped into tokens. The puzzle pieces from act one. Your dinner question, plus context, maybe a few hundred tokens. And these get converted into numbers. Because from here on, everything is math. Step four, pre-fill. The model reads. Here's something almost nobody knows. The model process your question in two totally different phases. The first is called pre-fill. The model reads your entire prompt all at once, in parallel, and builds an internal understanding of it. This is a burst of raw computation, billions of calculations, and it produces something called the KV cache. Don't let the name scare you. The KV cache is simply the model's working memory of your conversation. It's notes on everything said so far, held in super fast memory right next to the chip. Ever notice TriGPT pause for a little bit before the first word appears? That pause is pre-fill. The factory is reading the work order. Step five, decode. The model arrives. Now the assembly line starts. The model generates the answer one token at a time. It looks at your answer plus everything it has written so far. Runs the entire trillion knob network. Hundreds of billions of calculations. And produces one word. Try. Then it does the whole thing again. For the next word. A. Again. One pen. Again. Every single word of every TriGPT answer on earth is manufactured this way. One at a time. Full network pass each time. When you watch the answer type itself onto your screen. That's not a design flourish. You are literally watching an assembly line run in real-time. Each word appears the moment it's manufactured. Step six. The trip home. Each token springs back through the same fiber. And two seconds later, after you hit send, you are reading the dinner ideas. One more thing happening behind the curtain. You are not alone in there. The factory will batch everything you sent, which hundreds of other people question on the same trip simultaneously. Like a delivery driver grouping orders on one route. That batching is the difference between your question causing cents, and causing dollars. Now zoom all the way out. Because here's the whole factory. In layers. This is the map for the rest of this video. 10 layers. And here's the one sentence version of this entire video. Electricity comes in one side. Flows through silicon. Becomes computation and heat. The heat gets carried away by water. The computation gets coordinated by light. And what ships out the door is words. Electrons in. Token's out. That's the factory. So let's start a tour where every factory tour starts. The power plant. Because and this surprised me the most when I first dug into this ecosystem. The story of AI in 2026 is not mainly a story about chips anymore. It's a story about electricity. Let me give you the number that framed this whole industry for me. A traditional rack of servers. The kind that ran the internet for the last 20 years. Draws about 5 to 10 kilowatt. Think of a kilowatt as 10 old-fashioned 100-watt light bulbs burning at once. Nvidia's flagship AI rack. One rack. One refrigerator-sized cabinet. Draws 120 kilowatt. And the next generation coming later this year. The very rubbing racks are projected to approach 600 kilowatt per rack. That's 60 to 100 times jump in power density in under a decade. The electrical demand of an entire neighborhood packed into a phone booth. This is called power density. And it's the root cause of nearly everything in the next two acts. Now scale up. A large AI campus today want a gigawatt or more. A gigawatt is a thousand megawatt. Roughly the output of a full-size nuclear reactor. Enough electricity for about a million homes. Individual companies are now planning multiple campuses of that size. Data centers consume about 4 to 5% of US electricity going into this boom. The credible projections put it at 9 to 17% by 2030. And here's the collision. The US electrical grid was built brilliantly decades ago. For demand that grew 1 or 2% a year. AI showed up asking for 10th of gigawatts right now. The grid physically cannot say yes. Two bottleneck numbers. And they are the most important numbers in this act. Number one, the interconnection queue. To plug a big new facility into the grid, you file a request and wait in line while utilities study whether the grid can handle you. That line is currently 4 to 5 years long. This April, there are about 410 gigawatt of large projects waiting to connect. 87% of them are data centers. That's nearly 5 times the entire Texas grid's peak demand waiting in line. Number two, the transformer. A large power transformer, the giant grid box that step voltage down, used to take about a year to order. Today, 2.5 to 4 years. With prices up nearly 80%. You can have your chip in 6 months. The grid box that powers them? 2029. So what do you do if you're Microsoft or Meta? And every month of waiting calls you the API race. You stop waiting for the grid. You go around it. The industry calls this behind the meter power, generating electricity on site or next door. So you never touch the public queue. And that decision, thousands of companies making simultaneously, is what lit a fire under an entire forgotten sector of the stock market, borrowing all the industrial power companies. Let me introduce the players. From most dramatic to most dependable. The nuclear resurrection. In 2024, Microsoft signed a deal that would have sounded like sitire a decade ago. A 20-year agreement with Constellation Energy to restart Three Mile Island, the undamaged reactor next to the one from the 1979 accident. Constellation is spending about $1.6 billion, backed by a $1 billion federal loan, to bridge the 835 megawatt unit back in the second half of 2027, and Microsoft will buy every megawatt it produced for 20 years. Constellation operates the largest nuclear fleet in America, about 22 gigawatts. And suddenly, those aging reactors became some of the most valuable energy assets on Earth. Why? Because AI factories run 24/7, and nuclear is the only carbon-free power source that also runs 24/7. Constellation stock tell the story. It's now roughly a $90 billion company. Its pair, Vistra, with a 37 gigawatt fleet, mixing nuclear and gas, rode the same wave. The mode here is beautiful in simplicity. You cannot build a new conventional nuclear plant in America this decade. Existing reactors are replaceable. The small modular reactor lottery tickets. You have heard the tickers. Oklo. Backed by Sam Altman. With over 14 gigawatt of signed pipeline. New Scale. The only SMR design actually certified by US regulators. Here's my analytical skeptic framing. And be blunt. These are pre-revenue companies whose first commercial electron arrives around 2030 at earliest. Oklo doesn't yet have final regulatory approval for its design. New Scale booked about $31 million in revenue against a $356 million loss. Both stocks are down 65% to 78% from their late 2025 peaks. That's not a business yet. That's an option on the 2030s. The fastest power in the West. If the grid takes 4 years and nuclear takes 10, what can you get in 12 to 18 months? Fuel cells. Bloom energy makes solid oxide fuel cells. Boxes that convert natural gas into electricity chemically. No combustion. And you can park them behind the meter next to a data center fast. The stock nearly quadrupled in 2025, then doubled again in the first half of this year. And Bloom announced $7.65 billion in data center contracts in a single 90-day stretch. But this July, Henderbrook, an investigative outlet whose affiliated funds short the stock it covers, so weigh the source accordingly, published a report saying that Bloom's marketed $20 billion backlog is more than 40 times its binding contract obligations versus about 2x for typical peers. And that scaling to its stated ambitions would consume nearly the entire global supply of scandium. A metal China now requires export licenses for. Bloom formally rejected claims as false and misleading. When the backlog number and SEC findings disagree by 40x, the burden of proof is unaccompanied. Next, the arms dealer selling out through 2030. My favorite business in this act is the least glamorous, GE Vinova, the power spin-off of General Electric. They make the giant gas turbines that are realistically the number one near-term power source for AI, because gas is the only thing you can build at scale before 2030. GE Vinova's turbine slots are sold off through the end of this decade. Their backlog is around $163 billion. In the first quarter of 2026 alone, they booked $2.4 billion in data center electrification orders. More than all of 2025. The stock is up so much, it's now a nearly $300 billion company. The only real global rival at scale, Siemens Energy. Between the substation and the chips, sits a layer of equipment most people never think about. Switch gears, busways, and interruptible power supplies. Essentially a giant battery that catch the load instantly if the grid blinks, because even a half-second outage can cause a training run that's been running for a month. Three companies own this layer. I remember their name because two of them show up again in the next act: Vertif, Schneider Electric, and Eaton. Eaton's electric backlog grew 48% year-over-year. Vertif's backlog more than doubled to $15 billion, and the company joined S&P 500 in March. These are the companies selling shovels to every miner, regardless of who wins. And the last line of defense: Rolls of backup generators from Caterpillar and Cummins. Diesel engines the size of school buses, idling in wait for the one hour a year the grid fails. Analysts think Caterpillar's data center generator businesses could triple by 2030. Before we move on, the uncomfortable part, because this act has one, and it's showing up in your inbox. Because data center bid for scarce power, they bid against you. In the PJM market, the grid covering 13 states from Illinois to Virginia, data center demand added over $9 billion to the latest capacity auction. Translating to residential bills rising $16 to $18 a month in parts of Ohio and Maryland. Communities are noticing. Moratoriums are being proposed. This is becoming a genuine political risk to the build-out. And any honest map of this industry has to include it. So the factory has power. 100kW are now flowing into a single rack of chips, which creates an immediate problem. Physics 101. Every one of those watts becomes heat. The factory is running a fever. A single flagship AI chip today dissipates over 1,000W of heat. A chip the size of a postcard. Putting out the heat of a full-size space heater. Now stack 72 of them into one rack. Plus their memory and networking. You have got 120kW of heat. The output of about 80 space heaters. In a cabinet you could hang. Why? Because computation is heat. Every one of those trillions of calculations pushes electrons through microscopic wires. And electrical resistance turns into warmth. The factory's raw material electricity doesn't get consumed making tokens. It gets converted almost entirely into heat. Cooling isn't a support function of AI data center. Cooling is half the job. For 30 years, the answer was air conditioning. Genuinely just fancy AC. Cold air pushed up through the floor. Hot air sucked out the back. Giant chillers and cooling towers on the roof. And the air worked fine. Up to about 30-50kW per rack. But we just passed that line. Permanently. Air physically can now carry heat away fast enough from a 120kW rack. Unit hurricane force winds through the servers. So the industry is undergoing the biggest plumbing change in history. The switch from air to liquid. Water carries heat about 3,000 times more effective than air per unit volume. The technology ladder in one breath. Rare door heat exchangers. A water-cooled radiator bolted to the back of the rack. A transitional patch. Direct-to-chip cooling. The 2026 mainstream. A metal plate with liquid channels sit directly on top of each chip. Connected by hoses to a CDU. A coolant distribution unit. Think of it as the rack's heart. Pumping coolant to every chip. And carrying the heat to the building's water loop. NVIDIA's flagship racks don't offer this as an option. They require it. And at the extreme. Immersion cooling. Literally dunking the entire servers into tanks of non-conductive fluid. Like deep-frying a computer that never burns. Two quick vocabulary items investors will encounter. PUE. Power usage effectiveness. It's the factory efficiency score. Total power in. Divided by power that actually reaches the computer. A perfect score is 1.0. Old data centers run about 2.0. A watt of cooling for every watt of computing. Modern liquid-cooled facility hit 1.1. That efficiency gap. Times a gigawatt. Times electricity prices. Is real money. And water. Many data centers cool by evaporating millions of gallons. Which is becoming a genuine permitting and political fight in dry regions. Closed-loop liquid systems help. But watch this issue. It decides where facilities get built. Who gets paid? Largely the same names as the power room. Vertif is the market leader. The rare company is selling both the power gear and the liquid cooling. A one-stop shop growing revenue 28% a year at 20% margin. And then something remarkable happened. The two electrical giants each spend billions to bear their way into liquid cooling within months of each other. Eaton paid about $9.5 billion for Boyd Thermal. Schneider Electric bought Multivare. When the electric's most disciplined industry acquires, both pay up for the same niche. They are telling you what they think every future data center looks like. Smaller pure players. Invent in cooling loops and enclosures. And private cool IT systems. The specialists whose cool plates ship inside many brand name servers. The liquid cooling market was about $5 billion in 2025. Forecast put it at $15 to $27 billion by the early 2030s. It's the single clearest picks and shovels growth lane in this entire ecosystem. Because it does not care whether NVIDIA or AMD or Google wins. Heat is heat. Alright. The factory has power. The fever is under control. It's time to walk onto the factory floor and meet the machine your dinner question actually runs on. And the $3 trillion company that built it. This is the machine your dinner question runs through. NVIDIA's GB200 NVL72. 72 GPUs wire together so tightly, they behave as a single giant computer. It weighs about a ton and a half. Draws those 120 kilowatt we discussed. And costs roughly $3 million. So let's open it up. And to keep the parts straight. Come back to the factory. Specifically, it's kitchen. The CPU is the head chef. The central processing unit ran the operating system. Takes orders. Coordinates everything. Brilliant at complex sequential tasks. But there's only a handful of them. For decades, the CPU was the star of computing. Intel's kingdom. In the AI server, it has been demoted to management. The GPUs are 10,000 line cooks. A graphic processing unit. Originally invented to draw video game graphics. Contains thousands of small simple cores. They all do the same operation simultaneously. It turns out the math inside the neural network is exactly that kind of work. Billions of identical multiply and add operations. One head chef cannot do that. 10,000 line cooks. Each chopping one onion at the same instance. Can. That accident of history. Gaming graphics and AI needing the same math. Is the foundation of NVIDIA's empire. HBM is the countertop. High bandwidth memory. Hold that thought. It gets its own act. The SSD is the pantry. The NIC, the network interface card. Is the waiter. Carrying dishes between kitchens. And the power supplies and the motherboard. Are the plumbing and wiring. Holding it all together. Now the companies. NVIDIA finished its last fiscal year with $215.9 billion in revenue. Up 65%. Of which about $194 billion was data center. It controls roughly 80 to 86% of the AI accelerator market. Its growth margin in recent quarters is about 75%. 75% on hardware. Apple. The most admired hardware company in history. Runs around $46. NVIDIA became the first $5 trillion company last October. And depending on the week. Roughly $0.07 of every dollar in the S&P 500 is NVIDIA. How is that margin possible? Everyone says best chips. And sure. But the real answer is the word will unpacked fully in Act 8. 20 years of software that every AI developer on earth was trained on. For now. The one-line version. NVIDIA doesn't only sell chips. It sells the only complete factory floor system the world's engineers already know how to operate. Buying the competitor's chip means retraining your whole workforce. The bare case. About 40% of NVIDIA's revenue come from just four customers. And all four are building their own chips to replace it. Next. The challenger. AMD. AMD's Instinct GPUs are genuinely competitive on inference. More memory per chip. And by some estimates. 25-40% better tokens per dollar. Their problem was never silicon. It's software. Their CUDA alternative. Called ROCM. Now hits 90-95% of NVIDIA's performance on standard workloads. But 90% as good. With more friction. It's a hard pitch. When your training can cost $100 million. AMD holds maybe 5-7% of this market. Watchable. But it's improving. Still a distant second. Intel. Painful to say. Is barely in this race. Still selling plenty of head chefs. But the kitchen stopped being about head chefs. Now the quiet assassin. Broadcom. Here's the plot twist most retail investors miss. Those four hyperscalers building their own chips? They cannot actually do it alone. Designing a frontier AI chip takes a decade of specialized IP. So they hire Broadcom. Which co-designs Google's TPU. Meta's training chip. And reportedly. Open AIs. Broadcom books over 60% of custom chip market. Posted AI revenue up 106% last quarter. With a $73 billion backlog. And management says it has lines of sight to $100 billion of AI revenue in 2027. In our factory analogy. NVIDIA sells finished kitchens. Broadcom helps the biggest restaurant chains build their own. And takes a cut either way. It also dominates the switch silicon in Act 6. One company. Both sides of the war. Now. Who actually built those racks? Not NVIDIA. NVIDIA designs. Supermicron integrated full liquid cooled racks faster than anyone. Revenue up 123% last quarter. Over 90% of AI. Gross margin? Between 6 and 10%. Depending on the quarter. Dell has taken over $64 billion in AI server orders. With a $43 billion backlog. A staggering number. At server segment margins under 9%. HPE rides supercomputing and sovereign AI deals. And beneath the brand see the true invisible giants. The Taiwanese ODMs. Original design manufacturers. Foxconn. Which assembles roughly 40% of the world's AI racks. Quanta. Bywin. The same rack pass through many hands. NVIDIA captures 75 point margin on the silicon. The company that physically screw it all together. Keeps 6 to 10. In hardware. The profit lives in whatever is scarce. Chips and software are scarce. Assembly is not. But I've been hiding something from you. SS72 GPUs behave as a single computer. Nobody hit that with a magic wand. Making 10,000 line cooks work as one brain. Is arguably the hardest engineering problem in the entire building. And it's where some of the best businesses in the ecosystem hide. So why can't one GPU do the job? Simple. The model doesn't fit. A frontier model has over 1 trillion parameters. Those adjacent knobs. Requiring terabytes of ultra-fast memory. The biggest GPU carries a few hundred gigabytes. So the model gets sliced across thousands of chips. And here's the consequence. To produce every single token. Those chips must exchange intermediate results. Constantly. At an unimaginable speed. Back to the kitchen. 10,000 line cooks. Preparing one dish together. Every chef needs ingredients from other chefs. Every second. If passing ingredient is slow. Your 10,000 chefs stand around waiting. And these are the most expensive chefs in history. At cluster scale. A network that's 10% slower. Can idle billions of dollars of silicon. That's why networking is roughly 40-60% of spending. For every dollar of GPU. Two words to define. Bandwidth. Is how much data moves per second. The width of the conveyor belt. Latency. Is the delay for one handoff. How long a single pass takes. AI needs both. Everywhere. At once. The wiring comes in two flavors. Scale up. Inside the rack. Is NVIDIA's preparatory NVLink. An extreme speed web that makes 72 GPU one machine. Scale out. Rack to rack. Across the building. Is where a war is being fought. For years. Serious AI clusters run on InfiniBand. A specialized ultra-low latency networking. That NVIDIA acquired in 2019. In 2019. With Melanox. Premium performance. Premium price. One vendor. Against it. Ethernet. The open universal standard. The same family of technology as your home network. Historically slower. But backed by literally everyone who is in NVIDIA. By early 2026. About 2/3 of new AI cluster networking is Ethernet. Open standards given time. Usually win. They just did. This is a proprietary versus open story. As old as tech. Who profit from the nervous system. Broadcom again. Its Tomahawk chips are the merchant silicon inside most high-end Ethernet switches. Arista networks build the switches themselves. The best-in-class boxes and software hyperscalers standardize on. Roughly 6% growth margin. Guiding to $11.5 billion this year. With the risk that its two biggest customers are 40% of revenue. Cisco. The incumbent. Still huge. Fighting to stay relevant in AI back-ends. Marwell placed the Broadcom playbook one-tiered down. Custom chips for Amazon and Microsoft. Plus leadership in the digital signal processor. Inside optical modules. And as Terra Labs. One of the most remarkable margin stories in the ecosystem. Makes tiny retimer chips. That clean up electrical signals. Degrading over mere inches of circuit board. At this speed. Boring. Invisible. 76% growth margins. Revenue up 93% last quarter. When data moves this fast. Even the space between two chips becomes a market. And then the plot literally turns to light. Copper wires can carry this speed only at a few meters before signals degrade. Fine inside a rack. Use this across a football field building. So between racks. Everything converts to light through glass fiber. The device during the conversion is the optical transceiver. A thumb-sized gadget. An electrical to light translator. And you need one at each end of every fiber link. A single large AI cluster consumes hundreds of thousands of them. At hundreds of thousands of dollars each. Replaced every upgrade cycle. It's a reservoir blade for the data centers. The names. Coherent. The market leader. Lumentum. Chinese volume champion. Inolite. Febrinite. The contract manufacturer that assemble for nearly all of them. The arms dealer's arms dealer. Corning. Which draws the glass fiber itself. And Amphenol. Whose connectors are the knuckles of the entire system. The frontier to watch is co-packaged optics. Moving the light conversion directly onto the switch chip. To slash power. The nervous system is built. 10,000 chips. Thinking as one. But there's a dirty secret on the factory floor. Most of the time. The most expensive chips in the world are waiting. Not for data from across the room. For data from 2cm away. Here's the secret. During the decode phase. The one word at a time assembly line. From Act 2. The GPU's math cores. Are often not the bottleneck. For every token. The chip must pool the model's parameters. And the conversation's working memory. That KV cache. From memory into its cores. The math is fast. The fetching is slow. Modern inference is what engineers call memory bandwidth bound. The line cooks are lightning. But the countertop cannot feed them ingredient fast enough. Which makes the countertop one of the most valuable pieces of real estate. In technology. The industry's answer is HBM. High bandwidth memory. Instead of laying memory chips flat on a board a few inches from the processor. HBM stacks them vertically. 8. 12 stories high. Drills thousands of microscopic elevator shafts. Through the silicon. And glued the whole tower directly next to the GPU on the same package. A skyscraper of memory downtown. Instead of suburbs of memory across the highway. Result is 5 to 6 times the bandwidth of conventional memory. At 5 to 6 times the price. Nvidia happily pays. Memory is now one of the biggest cost component inside every AI chip you have heard of. And only three companies on earth can make it. SK Hennix. The Korean company owns roughly 60% of HBM. And got there by out executing its giant neighbor. It bet on HBM years before it mattered. And shipped each generation first. Unlocking the lion's share of Nvidia's next generation allocation. Samsung. The largest memory company overall. Was embarrassingly late. Micron. The American champion. Went from afterthought. To selling out its entire 2026 HBM capacity in advance. Here's a stat for the act. In May 2026. All three memory makers crossed the trillion dollars in market value. Combined over 4 trillion. Roughly 16 times what they were a decade ago. Memory used to be the most brutal commodity business in tech. Boom. Bust. Bankrupt. Repeat. HBM changed the psychology. It sold out more than a year in advance. Educated like a scarce resource. Priced like a luxury good. The open question. Is whether that price survives the moment all three giant finish their capacity expansions at once. Memory cycles have broken hearts before. Now walk out the back of the factory to the warehouse. Storage. The hierarchy in one line. The closer to the chip. The faster and pricier. Cash on the chip itself. HBM beside it. Regular DRAM on the motherboard. SSDs. Flash drives. For hot data. And at the bottom. The technology everyone declared that 10 years ago. The spinning hard drive. Still unbeatable per terabyte. For cold bulk data. And AI turned out to be an AI hoarder. Training datasets. Model checkpoints saved every few hours. And the part nobody predicted. The output. Every conversation. Every log. Every generated image. Retained forever. Seagate's CEO calls it the inference inflection. AI doesn't just consume data. It produced it. Endlessly. Hard drives are literal dual play. Seagate and Western Digital. Plus flash drive players. Kyuxia. And Solidime. After a decade of decline. Both drive makers sold out their entire production into 2027. Western Digital now ships 89% of its revenue to cloud customers. And was one of the S&P 500 top performers. Consumer hard drive prices jumped 50% because AI ate the supply. A dying industry. Resurrected by the factory next door needing somewhere to put infinity. The machine is complete. Powered. Cooled. Wired. Fed. The hardware answers exactly zero questions. Something invisible has to run the place. Everything we have toured so far. You could theoretically buy. The hardware is perturbable. Where does the durable competitive advantage actually live? The stack. Briefly. Bottom to top. Linux. The free operating system running effectively every server on earth. Commercialized by Red Hat. And Canonical. Kubernetes. The invisible foreman. Open source software that schedules work across thousands of machines. Restart what crashes. And keep the factory floor humming. Then the serving layer. Then the models. Two stories in this act matter to investors more than all the rest combined. First story. CUDA. The 20-year trap. In 2006, Nvidia made a decision on Wall Street hated. It spent billions building a programming platform so scientists can use gaming chips for general math. For a decade, this looked like an expensive hobby. Then deep learning arrived. And every AI researcher on earth learned to build on CUDA. Because it was the only mature option. 20 years later, CUDA has millions of developers. Thousands of specialized libraries. And every framework optimized for it first. Understand what this means. When AMD ships a chip with better specs. And it sometimes does. The customer isn't competing chips. They are comparing chips. Plus the cost of retaining their entire engineering organization. And rewriting their code. With their 100 million training run on the line. That's why 70% gross margin survived competition. The mode was never silicon. The mode is the muscle memory of a million engineers. Second one. Why your question costs cents, not dollars? Raw and naive inference on a trading parameter model would be super expensive. The economics only work because of an unglamorous layer called the serving engine. Software like VLLM and NVIDIA's TensorRT. Doing three tricks. Batching. Grouping hundreds of users scratching through the chip at once. The delivery route trick from Act 2. Caching. Reusing what KV working memory. Instead of recomputing the conversation from scratch. For every word. And quantization. Rounding the model's numbers to lower precision. Like shipping a slightly compressed photo. Nearly identical quality. Fraction of the cost. Together a 3 to 10 times cost reduction. From software alone. When OpenAI or Anthropik cuts API price 80% in a year. It's mostly this layer. Not new chips. Token manufacturing costs are collapsing on a curve. Remember that for the finale. Because it cuts both ways. One last flaw of the stack. Your dinner question doesn't need it. But enterprise AI does. Rack. Retrieval Augmented Generation. The model is a brilliant writer with a fixed education. Rack hands it. Your company's documents. At question time. Via vector database. Search engines that find text. By meaning rather than keywords. From players like Pancon. Plus the data platform Databricks. And Snowflake. The librarian and the writer. Working together. That's the enterprise AI pitch in one sentence. And with that. The tour is over. You have seen every layer. Power. Cooling.
[00:36:32] Speaker ?: Silicon.
[00:36:32] Speaker 1: Light. Memory. Software. Now let's follow the money. All of it in one map. And then ask one question. Does any of this actually pay for itself? So here's the whole board. On the demand side. The miners. Three tiers. The hyperscalers. Microsoft. Amazon. Google. Meta. Oracle. Spending that combined $725 billion this year. The new clouds. Specialized GPU landlords. CoreView. Nibius. Lambda. Crusoe. And the AI Labs. Now valued around $850 billion. Anthropik. Which passed it. At roughly $965 billion. About $47 billion of annualized revenue. And XAI. Folded into SpaceX. Both major labs filed for IPO this June. The private market has already priced them as two of the most valuable companies on earth. The public market is about to vote. Not the costs. Bernstein estimates 1 gigawatt of AI data center. One campus. Costs about $35 billion to build. Where it goes? Roughly 39% is the chips. The single biggest line. Which is why NVIDIA's gross profit alone is estimated near 30% of total industry cost. And that $35 billion machine depreciates fast. Most operators write chips off over 4-6 years. And skeptics argue even that flatters the accounting. Since a 5-year-old GPU competes against chips 10 times better. Every year of depreciation life added or removed. Swings billions in reported profit. Watch that debate. So follow this. NVIDIA invests billions directly into CoreWeave. Into Nivea's. Into OpenAI. Those companies use the money to buy NVIDIA chips. Revenue for NVIDIA. The hyperscalers signed enormous contracts with the new clouds. Meta alone committed roughly $21 billion to CoreWeave. And reportedly $27 billion to Nivea's. Which conveniently moves data center spending of Meta's balance sheet into someone else debt. The AI Lab sends compute deals with hyperscalers. OpenAI with Microsoft and Oracle. Anthropik with Amazon and Google. Paid partly with money those same hyperscalers invested into them. Let's be fair. Because an analytical skeptic label cuts both ways. This is not fraud. And it's not new. Vendor financing build the railroads and the telephone network. The loop has exactly one opening. Where fresh money is supposed to enter. End users and enterprises. You paying $20 a month. Companies paying for APS and copilots. That is the only exit that isn't recycled capital. And today. That corner is the smallest number on the board. The two biggest labs combined analyze run rates. Now top $70 billion. Genuinely spectacular. The fastest revenue ramp in software history. But the revenue they'll actually book this calendar year. Is a fraction of that. Set against $725 billion of annual spend. With OpenAI still expected. To lose on the order of $14 billion this year. The entire $7 trillion machine. Is the bet that the small corner grows faster. Than the big circle spins. So let's watch it one more time. Same two seconds. But now you can see. Your thumb hits sand. And the question becomes light in glass fiber. Crossing state lines. It arrives at a building drawing the power of a city. Power from a restarted nuclear plant. A sold out turbine. A fuel cell parked behind the meter. Is chopped into tokens. And fed into a $3 million rack. Assembled in Taiwan. Sold at a 75 point margin. Where 10,000 line cooks fetch a trillion parameters. From memory skyscrapers. While liquid coolant carries away the heat of 80 space heaters. And light speed interconnects. Let 10,000 chips. Think one thought. The software batches you with 1,000 stringers. And the answer streams back. Token by token. Each word manufactured instance you read it. Two seconds. $700 billion a year. The largest infrastructure project our species has ever attempted. So a machine can suggest you make one pan chicken and rice. Whether that's the most important investment in history. Or the most expensive. Honestly. Nobody on earth knows yet. That's the whole map. If you find this video helpful. Please subscribe. And I'll see you in the next one.