Try Free

AMD Advancing AI 2026: Lisa Su Full Keynote

Yahoo Finance July 23, 2026 2h 11m 19,827 words
▶ Watch original video

About this transcript: This is a full AI-generated transcript of AMD Advancing AI 2026: Lisa Su Full Keynote from Yahoo Finance, published July 23, 2026. The transcript contains 19,827 words with timestamps and was generated using Whisper AI.

"Good morning. That's a pretty good good morning. How about let's try one more time? Good morning. Good morning. And welcome to Advancing AI 2026. It's so great to be back here in San Francisco and to see so many friends and partners and customers and especially all the developers that are here with"

[00:00:00] Good morning. [00:00:05] That's a pretty good good morning. [00:00:07] How about let's try one more time? [00:00:07] Good morning. [00:00:09] Good morning. [00:00:10] And welcome to Advancing AI 2026. [00:00:14] It's so great to be back here in San Francisco [00:00:16] and to see so many friends and partners and customers [00:00:19] and especially all the developers that are here with us today. [00:00:24] And I want to say a big welcome to everyone [00:00:26] who's joining us online from around the world. [00:00:29] This is my absolute favorite event of the year. [00:00:33] It's where we bring the entire AI ecosystem together [00:00:36] to show what we've been building and where we're going next. [00:00:39] And this year, this is our biggest show ever [00:00:42] because we have so much to tell you. [00:00:45] So let's go ahead and get started. [00:00:48] At AMD, our mission is to push the boundaries [00:00:51] of high performance in AI computing [00:00:53] to help solve the world's most important challenges. [00:00:56] And I often say AI is the most important technology [00:01:00] of the last 50 years. [00:01:02] And frankly, the progress that we're seeing in the industry [00:01:05] is just incredible. [00:01:07] Like every month, every few months, [00:01:09] we see something new that was far beyond our imagination. [00:01:13] And you can already see the impact across every industry. [00:01:17] In healthcare, AI is helping researchers [00:01:20] identify new drug candidates faster. [00:01:22] In science, we're solving problems [00:01:25] that we didn't think were possible a few years ago. [00:01:28] And across every industry, AI is changing the way we get work done. [00:01:33] And the most important thing is [00:01:35] we're still in the very, very early innings of what's possible. [00:01:40] Now, just take a look at some of these charts. [00:01:43] You know, we look at these every year [00:01:44] and you kind of see the rate and pace of AI growth. [00:01:49] A couple of years ago, we were just starting. [00:01:51] People were experimenting with chatbots [00:01:53] and, you know, it was really cool. [00:01:55] But if you look at today, [00:01:57] we're really seeing more than 35 quadrillion tokens [00:02:01] consumed every single month. [00:02:04] That's an increase of nearly 160 times in just two years. [00:02:09] And that curve is actually just getting steeper. [00:02:12] You're going to hear a lot about that today. [00:02:14] So just looking at, you know, where is all that demand? [00:02:17] I mean, obviously, training remains incredibly important [00:02:20] and foundational. [00:02:21] And over the last four or five years, [00:02:23] you know, we see the compute that's used for training, [00:02:26] advanced models, continuing to increase by, [00:02:28] let's call it roughly 5x every year. [00:02:31] And those models are getting better. [00:02:32] Like, we're all experiencing that. [00:02:34] We're seeing better reasoning. [00:02:35] We're seeing new capabilities. [00:02:37] We're now seeing a growing number of specialized models. [00:02:40] And one of the things that I really believe in [00:02:42] is there is no one perfect model. [00:02:45] I think we're all going to use a slew of different models, [00:02:48] depending on what task you're trying to solve [00:02:51] and which industry you're in. [00:02:53] But the bigger shift is actually in inference. [00:02:56] And we expected this. [00:02:58] We expected that inference would grow faster than training. [00:03:01] And we can say that in 2026, for the first time, [00:03:05] the world is using more AI compute to run models than to train them. [00:03:10] And we've seen that to a point where we think this year, [00:03:14] roughly 60% of the global AI compute capacity [00:03:17] will be used for inference. [00:03:19] And the reason for that is actually pretty simple. [00:03:23] As billions of people every day use AI, [00:03:27] the workloads shifts from building models to putting them to work. [00:03:30] And actually, we're going to talk a lot today about agentic AI. [00:03:35] And that's accelerating the shift even faster. [00:03:37] So if you think about all of this, [00:03:40] you know, we see agents as the next big step for AI. [00:03:44] And the most interesting thing is, [00:03:46] like, this has really just started accelerating [00:03:49] over the last, I would say, five or six months this year. [00:03:52] So when we started with LLMs, [00:03:54] they were great to answer questions. [00:03:56] But we found that to really get the power of AI, [00:04:00] you want agents to be able to really answer full questions. [00:04:05] So you give it a goal, [00:04:07] it keeps working until it's solved the problem. [00:04:10] And what that means for us in the compute industry, [00:04:13] it means that compute demand is growing at an incredible pace. [00:04:19] We're actually seeing a step change in compute demand. [00:04:22] Because when you ask the agent to do something, [00:04:25] it actually has dozens of steps. [00:04:28] And it has to reason, and it has to call tools, [00:04:30] and it has to access data, [00:04:31] and it has to keep doing it over and over [00:04:34] until it solved the problem. [00:04:35] And so you need lots of GPUs to do all that reasoning. [00:04:39] But importantly, you also need a lot of CPUs [00:04:42] to orchestrate every step around it. [00:04:45] So when you look at, you know, [00:04:47] just what that means from a market standpoint, [00:04:49] it's really hard to put this market data up [00:04:52] because I feel like every few months, [00:04:54] we're changing, you know, sort of the perspective [00:04:56] based on what our customers are saying. [00:04:58] But what we're seeing is that the shift to agentic AI [00:05:02] is certainly growing the AI accelerator market [00:05:05] very significantly. [00:05:07] It was actually last year at this event [00:05:09] that we called the AI accelerator market [00:05:12] at about 500 billion by 2028. [00:05:15] And at that time, [00:05:17] it actually felt like a very big number. [00:05:20] Didn't you think it was a big number? [00:05:22] Oh, yeah. [00:05:23] But honestly, you look at it today [00:05:25] and it looks conservative [00:05:26] because the AI demand [00:05:28] is just continuing to accelerate. [00:05:31] And you can see it, right? [00:05:32] Better models, more usage, [00:05:33] more usage requires more compute, [00:05:36] and more compute actually builds better models. [00:05:39] And so we're now expecting [00:05:40] that by 2030, [00:05:42] the AI accelerator market [00:05:44] is going to reach about 1.4 trillion by 2030. [00:05:48] So what do you think about that? [00:05:49] Is that a big number? [00:05:55] I mean, what that means is [00:05:57] by the end of the decade, [00:05:58] the AI accelerator market [00:06:00] is going to approach the size [00:06:01] of the entire semiconductor market today. [00:06:05] And although there will be [00:06:06] many different types of accelerators, [00:06:08] you know, I'm a big believer [00:06:09] in there's no one-size-fits-all [00:06:11] as it comes to chips, [00:06:13] we do expect that GPUs [00:06:15] are going to make up [00:06:16] the vast majority of that market [00:06:17] because the algorithms [00:06:19] are still very much in their infancy [00:06:21] and we're still continuing [00:06:23] to see the workloads change. [00:06:24] And that favors programmability [00:06:26] in the overall silicon ecosystem. [00:06:31] Now, perhaps the most interesting part [00:06:34] of the last six months [00:06:35] has been, you know, [00:06:36] GPUs are only part of this story. [00:06:39] What we're actually seeing [00:06:41] is Agentic AI is creating [00:06:43] an entirely new growth vector [00:06:45] for server CPUs. [00:06:47] Now, I've updated this number a lot also [00:06:49] over the last, you know, [00:06:51] six to eight months. [00:06:53] But look, we call it like we see it. [00:06:56] And, you know, we were early. [00:06:57] I mean, we saw from our largest [00:06:59] hyperscale customers, [00:07:00] and you're going to see some of them [00:07:01] here today, [00:07:02] who said, look, [00:07:03] as AI inference is going up, [00:07:06] we need more server CPUs. [00:07:07] And so we saw growth [00:07:08] in the server CPU market. [00:07:11] And we said, you know, [00:07:12] last year we thought it might grow, [00:07:13] let's call it 18 to 20% CAGR [00:07:15] to an overall market size [00:07:18] of 60 billion. [00:07:19] But frankly, [00:07:20] what we're seeing is that [00:07:21] the rate and pace [00:07:22] of agentic AI adoption [00:07:23] is much, much faster [00:07:25] than any of us thought. [00:07:27] So we're talking about agents [00:07:29] going from millions to billions. [00:07:32] Like we're able to increase [00:07:33] all of our productivity. [00:07:35] And that requires a tremendous amount [00:07:37] of CPU infrastructure. [00:07:39] And every customer conversation [00:07:41] is telling us that this buildout [00:07:43] is just beginning. [00:07:44] So based on what we're seeing today, [00:07:46] and we're going to talk more about [00:07:48] how agentic AI is evolving [00:07:49] with CPUs, [00:07:51] we now expect the server CPU market [00:07:53] to grow by over 50% [00:07:55] to over 200 billion by 2030. [00:07:59] So that's starting from today, [00:08:01] a $25 billion market [00:08:03] going to over 200 billion. [00:08:04] So there's a lot of excitement [00:08:06] about CPUs as well. [00:08:10] Now, probably one of the things [00:08:12] that is differentiating AMD [00:08:15] in how we think about the market [00:08:17] is this is not just [00:08:19] a data center opportunity. [00:08:21] As AI becomes a larger part [00:08:23] of our daily lives, [00:08:25] we want the intelligence to run [00:08:27] right where the work happens. [00:08:29] And that means cloud [00:08:31] is super important, [00:08:32] but the devices that we use every day [00:08:34] are going to be extremely important. [00:08:36] And that means that we're going to do [00:08:37] a lot more at the edge as well [00:08:39] with autonomous machines [00:08:40] that sense and act in real time. [00:08:43] And we really believe that you need AI [00:08:45] to be infused everywhere. [00:08:47] And that's exactly what we're focused on at AMD. [00:08:50] So putting all of that together, [00:08:52] we actually expect that the market [00:08:54] for our high-performance [00:08:55] and adaptive computing products [00:08:57] to grow at roughly a 40% tagger [00:08:59] over the next several years, [00:09:01] approaching a $2 trillion market [00:09:04] by 2030. [00:09:05] And frankly, [00:09:06] the only way you're going to be able [00:09:08] to service a market like that [00:09:09] is for us to work together [00:09:11] as an ecosystem. [00:09:13] There's no one company [00:09:14] that can solve it all, [00:09:15] but this is the opportunity [00:09:17] to bring the best [00:09:19] and brightest together. [00:09:20] And we have never been [00:09:22] in a better position to lead. [00:09:24] We have the broadest product portfolio. [00:09:27] We have the strongest roadmaps [00:09:28] we've ever had. [00:09:30] And the thing that I'm most proud of [00:09:31] is we have the deepest partnerships [00:09:33] with the companies [00:09:35] that are building this future. [00:09:38] So a little bit about our strategy. [00:09:40] Look, I think we've been very consistent. [00:09:43] You know, we've talked about [00:09:44] our multi-year strategy, [00:09:45] and it's really built [00:09:46] around three priorities. [00:09:48] The first is just compute leadership. [00:09:50] We're building the broadest set [00:09:51] of compute engines in the industry [00:09:53] so that we have the right compute [00:09:55] for the right workload. [00:09:57] Second, it's going to be [00:09:58] about open platforms. [00:10:00] We want everyone to come together [00:10:02] in an open ecosystem. [00:10:04] We believe an open ecosystem [00:10:05] is essential to the future of AI, [00:10:07] and that's how we get [00:10:08] the force multiplier [00:10:10] of everyone coming together. [00:10:12] And that's open on hardware [00:10:13] in terms of hardware standards, [00:10:15] as well as open on software. [00:10:17] And that's why we're investing [00:10:19] so heavily in our Rockham software. [00:10:21] And you're going to see [00:10:22] that AI has been [00:10:23] a tremendous multiplier [00:10:25] in this rate and pace [00:10:27] of progress that we're making [00:10:28] in Rockham and in software. [00:10:31] And it really means that [00:10:32] if we put all of this together, [00:10:35] we can have our customers [00:10:36] deploy on AMD hardware [00:10:39] faster than ever. [00:10:41] And the third piece is [00:10:43] we're going to get a chance [00:10:43] to talk to you [00:10:44] about powering AI everywhere. [00:10:46] So that's AI in enterprise, [00:10:48] that's AI at the PC, [00:10:50] that's AI in the physical world. [00:10:53] And that's adding new capabilities [00:10:54] across all of our products. [00:10:56] So today, you're going to see [00:10:57] all of that in action. [00:10:59] We're going to start with compute [00:11:00] for the agentic error. [00:11:02] I'm going to show you [00:11:02] a lot of hardware. [00:11:04] And then we're going to talk [00:11:05] about some software [00:11:06] and then our overall products [00:11:08] from a platform standpoint. [00:11:09] And we're honored to have [00:11:11] some really special guests [00:11:13] who are going to join us [00:11:14] to really help bring [00:11:15] the technology to life. [00:11:17] So let's start [00:11:18] with the data center. [00:11:20] Today, AMD EPYC runs [00:11:22] on the most important workloads [00:11:23] in the world. [00:11:25] We power the largest cloud providers, [00:11:27] we power the digital platforms [00:11:28] that billions of people use [00:11:29] every day and the most important [00:11:32] critical systems that are used [00:11:34] by the largest businesses, [00:11:36] including more than 60% [00:11:37] of the Fortune 100 [00:11:38] are using AMD EPYC. [00:11:40] And that momentum [00:11:41] is just building. [00:11:43] Last quarter, [00:11:44] we reached a record 46% revenue share [00:11:47] in the server market. [00:11:49] And we're continuing to see [00:11:51] every major customer [00:11:52] move more workloads to EPYC. [00:11:56] And on the GPU side, [00:11:58] our instinct adoption [00:11:59] has also accelerated. [00:12:01] We have a broad set of customers [00:12:02] across the largest AI labs, [00:12:05] cloud providers, [00:12:06] leading AI startups, [00:12:08] as well as some of the national labs [00:12:10] and sovereign AI opportunities [00:12:11] are being built on AMD instinct. [00:12:14] And we really appreciate [00:12:16] that opportunity. [00:12:19] Now, AI computes [00:12:20] becoming a lot more complicated. [00:12:22] And what we're seeing [00:12:24] with Frontier AI [00:12:24] is it's really raising the bar [00:12:27] for what infrastructure needs to do. [00:12:29] It takes more than a single chip [00:12:32] or a single server. [00:12:34] You actually have to design [00:12:36] the entire rack as one system. [00:12:39] And that requires leading CPUs, [00:12:42] that requires leading GPUs, [00:12:44] that requires high-speed networking [00:12:46] that connects everything inside the rack [00:12:48] and across the data center. [00:12:50] And just as importantly, [00:12:51] this is about making these systems [00:12:54] super easy to use. [00:12:56] And so they need to be easy to deploy, [00:12:58] easy to service, [00:12:59] and very reliable. [00:13:01] And that is exactly what we built [00:13:03] with Helios. [00:13:05] So today, [00:13:06] I'm super excited [00:13:07] to launch Helios, [00:13:08] the industry's highest-performance AI ref. [00:13:12] Now, Helios is built [00:13:19] to train and run [00:13:20] the most demanding Frontier models [00:13:22] in the world [00:13:22] at massive scale. [00:13:24] And I have a lot of show-and-tell for you. [00:13:27] So for those of you who know me, [00:13:28] you know that I love holding up chips. [00:13:33] I call myself sometimes [00:13:34] the Vanna White of chips. [00:13:37] But it turns out with Helios, [00:13:39] there are lots and lots of components. [00:13:41] And so you can see, [00:13:43] these are the chips [00:13:43] inside the Helios rack. [00:13:46] These are the Instinct. [00:13:49] Thank you. [00:13:53] So this is Instinct, [00:13:56] Epic, [00:13:57] and then the Pensando chips [00:13:59] for networking [00:13:59] that are inside each rack. [00:14:01] And every one of these chips [00:14:03] is part of the system. [00:14:05] So MI-455 is the engine. [00:14:08] It's the highest-performance GPU [00:14:09] in the industry. [00:14:11] Venice, [00:14:12] which drives it, [00:14:13] is the world's fastest CPU cores. [00:14:16] And Pensando DPUs, [00:14:20] actually, [00:14:20] and NICs, [00:14:21] actually connect them [00:14:22] with leadership programmability [00:14:23] and scale-out bandwidth. [00:14:25] So, [00:14:26] I have a few more props [00:14:27] to show you. [00:14:29] Let's kind of take a look [00:14:30] at how these things fit [00:14:31] into the overall system. [00:14:33] So first, [00:14:33] this is now [00:14:34] the MI-455 Accelerator, [00:14:37] which is mounted [00:14:38] on what we call [00:14:40] an Enhanced Accelerator Module, [00:14:42] or EAM. [00:14:43] It's a production module [00:14:45] that goes into Helios, [00:14:46] and it is much more [00:14:48] than the GPU package. [00:14:50] The EAM [00:14:50] actually integrates [00:14:52] the GPU, [00:14:53] memory, [00:14:54] power delivery, [00:14:55] high-speed interfaces, [00:14:57] system management, [00:14:59] cold plates [00:14:59] for liquid cooling, [00:15:01] all into one [00:15:02] compact, [00:15:03] serviceable module. [00:15:05] Pretty cool, right? [00:15:07] Yeah. [00:15:12] Now, [00:15:13] if I just give you [00:15:13] some of the specs, [00:15:14] we're talking about [00:15:15] 320 billion transistors. [00:15:18] This is built [00:15:19] with TSMC's [00:15:20] leading 2-nanometer [00:15:22] and 3-nanometer [00:15:22] process technology. [00:15:24] And most importantly, [00:15:25] it brings together [00:15:26] nearly a decade [00:15:27] of AMD chiplet innovation. [00:15:29] So it allows us [00:15:30] to combine [00:15:30] 12 compute [00:15:32] and I-O chiplets [00:15:33] with 432 gigabytes [00:15:35] of memory, [00:15:36] all connected [00:15:37] by leading industry [00:15:38] to 3D chip stack packaging. [00:15:41] And I can say for sure, [00:15:42] it's the highest performance [00:15:44] AI accelerator [00:15:45] in the industry. [00:15:48] OK, so this is [00:15:54] the CPU board [00:15:55] that's inside [00:15:56] the Helios compute tray. [00:15:57] And I'm going to talk [00:15:57] a lot about CPUs later. [00:15:59] But just to give you [00:16:00] a high-level view, [00:16:03] it's high-speed, [00:16:04] 96-core epic processor, [00:16:06] DDR5 memory, [00:16:07] all the I-O [00:16:08] that's needed [00:16:08] to feed the GPUs, [00:16:10] all in a single motherboard. [00:16:12] And we also have [00:16:14] our Selina DPU. [00:16:16] This is a critical part [00:16:18] of the networking [00:16:18] infrastructure. [00:16:20] And it really allows us [00:16:21] to deliver both [00:16:22] front-end [00:16:23] as well as [00:16:23] the scale-out [00:16:24] and scale-across [00:16:25] bandwidth overall. [00:16:27] And this is our [00:16:28] Vulcano AI NIC. [00:16:30] And you can see, [00:16:31] this actually allows us [00:16:33] to really have [00:16:34] up to six Vulcano NICs [00:16:36] on a board. [00:16:37] And it uses [00:16:38] the open [00:16:38] Ultra-Ethernet standard. [00:16:40] So each Helios compute tray [00:16:41] includes two of these [00:16:43] Vulcano boards. [00:16:44] And that allows us [00:16:45] to have the scale-out [00:16:47] and scale-up [00:16:48] of the scale-up [00:16:49] networking overall. [00:16:51] So you put all that [00:16:52] inside the rack [00:16:53] and we have [00:16:54] an additional [00:16:55] six dedicated [00:16:56] networking trays [00:16:57] that handle [00:16:57] the scale-up networking. [00:16:59] And those are all [00:17:00] connecting 75 GPUs [00:17:02] together with [00:17:02] UA-Link over Ethernet [00:17:04] with Silicon [00:17:05] from our partners. [00:17:06] And that's what we mean [00:17:07] by an open ecosystem. [00:17:09] So when you bring [00:17:10] this all together, [00:17:11] I want to say, [00:17:13] Helios is simply [00:17:14] the best AI rack [00:17:15] in the world. [00:17:24] Now, [00:17:25] let's take a look [00:17:25] at some numbers. [00:17:26] Clearly, [00:17:27] it's a very competitive [00:17:28] world out there. [00:17:29] So when you compare [00:17:30] Helios to the competition, [00:17:32] we're delivering [00:17:33] 15% more compute, [00:17:36] 50% more HBM4 memory [00:17:38] capacity and memory [00:17:39] bandwidth, [00:17:40] and 50% more [00:17:41] scale-out bandwidth. [00:17:43] And what that means [00:17:44] is that every Helios [00:17:46] can deliver more performance [00:17:47] for the largest models, [00:17:49] more capacity [00:17:50] for longer context, [00:17:52] and the bandwidth [00:17:52] to scale [00:17:53] across thousands [00:17:54] of racks. [00:17:56] And now, [00:17:57] we can kind of [00:17:57] come over here [00:17:58] and take a look [00:17:58] at another [00:17:59] very large piece [00:18:00] of hardware. [00:18:02] This is now [00:18:03] the production hardware [00:18:04] that customers [00:18:04] are deploying, [00:18:05] and you can see [00:18:05] more of that [00:18:06] as you go through [00:18:07] the exhibit center. [00:18:10] And each one of these [00:18:12] weighs more than [00:18:13] 160 pounds [00:18:14] and stands less [00:18:16] than 2 inches tall. [00:18:18] That's just [00:18:18] an extraordinary [00:18:19] amount of technology [00:18:21] when you look [00:18:22] at what's in [00:18:22] one of these trays. [00:18:23] And when you bring [00:18:24] these things together, [00:18:25] the numbers [00:18:26] are just incredible. [00:18:28] More than 18,000 [00:18:29] cDNA 5GPU [00:18:30] compute units, [00:18:32] over 4,600 [00:18:33] Zen6 CPU cores, [00:18:35] and 31 terabytes [00:18:37] of HBM4 memory, [00:18:38] all in a single rack. [00:18:41] That's what it takes [00:18:42] to run [00:18:43] agentic AI at scale. [00:18:46] So, [00:18:47] today, [00:18:47] I'm excited [00:18:48] to announce [00:18:48] that Helios [00:18:49] is in full production. [00:18:55] We have [00:18:59] shipments on track [00:19:00] to start [00:19:01] at the end [00:19:02] of the third quarter [00:19:02] and ramping [00:19:03] into the fourth quarter [00:19:04] and the second half [00:19:05] of the year. [00:19:06] And I can tell you, [00:19:07] customer demand [00:19:08] for Helios [00:19:08] is extremely strong. [00:19:10] And we're extremely proud [00:19:12] of the work [00:19:12] that we have done [00:19:13] across the leading [00:19:14] AI labs [00:19:15] to adopt Helios. [00:19:16] And so, [00:19:17] let's start [00:19:18] with some of our guests. [00:19:19] Today, [00:19:19] I'm excited to be joined [00:19:21] by our newest [00:19:22] Instinct partner [00:19:22] and one of the leaders [00:19:24] in Frontier AI. [00:19:25] To share more [00:19:26] about what we're [00:19:26] doing together, [00:19:27] please welcome [00:19:28] Anthropic co-founder [00:19:29] and Chief Compute Officer, [00:19:31] Tom Brown. [00:19:41] Hello, Tom. [00:19:43] It is so great [00:19:44] to have you here [00:19:45] and it is really [00:19:47] such an honor [00:19:47] to be working [00:19:48] with Anthropic. [00:19:49] We just, [00:19:49] we had a big week [00:19:50] this week. [00:19:51] We did, yeah. [00:19:52] So, look, [00:19:54] Anthropic has just [00:19:54] had incredible success [00:19:56] over the last few years. [00:19:58] You know, [00:19:58] Claude has made [00:19:59] such a big impact [00:19:59] on the overall industry. [00:20:01] Compute is such [00:20:02] an important factor [00:20:03] in the progress. [00:20:05] Can you just talk [00:20:05] a little bit [00:20:06] about your compute strategy [00:20:08] and just how things [00:20:09] have been evolving? [00:20:10] Yeah. [00:20:11] Yeah, and first, [00:20:11] thank you so much. [00:20:12] A hard act to follow [00:20:13] with Helios. [00:20:14] Amazing. [00:20:15] So, with Anthropic, [00:20:17] as you mentioned, [00:20:20] the scale [00:20:21] that we're growing [00:20:23] just as an industry [00:20:24] is enormous [00:20:25] and so we have been [00:20:27] working to make sure [00:20:29] that we have [00:20:30] the best, [00:20:31] like, different chips [00:20:33] and we can use [00:20:34] the best chips [00:20:35] for the best workloads. [00:20:37] Yeah, and look, [00:20:38] I think you've been [00:20:39] really leading in that [00:20:40] area and, you know, [00:20:42] I'm happy to say [00:20:43] that I've always [00:20:44] wanted to be [00:20:45] as part of your [00:20:46] infrastructure. [00:20:47] I remember the early [00:20:48] times we talked [00:20:48] as you were starting [00:20:49] Anthropic, [00:20:50] but we had a big [00:20:51] announcement this week. [00:20:52] We're honored [00:20:52] that you will be [00:20:53] deploying up to [00:20:54] two gigawatts [00:20:55] of Helios. [00:20:56] Why now? [00:20:57] What was the key reason? [00:20:59] Yeah, so I think [00:21:01] that, like, [00:21:02] the first thing [00:21:03] was just Helios [00:21:04] is an amazing machine. [00:21:05] It's, like, [00:21:06] it's absolutely [00:21:07] fantastic. [00:21:08] Do we like that? [00:21:09] Does that work for us? [00:21:14] And then I think [00:21:15] the thing that we [00:21:16] were thinking about [00:21:16] originally was [00:21:17] whenever we're [00:21:18] bringing up [00:21:18] a new hardware platform, [00:21:20] it's a big effort. [00:21:21] It's, like, [00:21:22] a huge thing. [00:21:23] And so as part [00:21:24] of the, [00:21:25] as we were thinking [00:21:26] about this, [00:21:27] we started doing [00:21:28] our own evaluation [00:21:30] of MI-355. [00:21:32] You guys generously [00:21:33] got us a rack [00:21:33] to start working. [00:21:35] And we expected [00:21:36] this to be [00:21:36] kind of a big process. [00:21:38] Our actual experience [00:21:39] was we had [00:21:41] one engineer [00:21:42] who started doing it. [00:21:43] They spun up, [00:21:44] they connected it [00:21:45] to Claude, [00:21:46] asked it, [00:21:47] hey, [00:21:48] bring up this machine, [00:21:49] left it going [00:21:50] over the weekend, [00:21:51] and we ended up [00:21:52] with a graph [00:21:53] of the actual [00:21:53] performance [00:21:54] of our leading [00:21:55] model on it [00:21:56] just going up [00:21:57] and up and up [00:21:58] over the weekend. [00:21:59] And so, [00:22:00] yeah. [00:22:03] And I think [00:22:04] that that's a testament [00:22:05] to the open platform [00:22:06] that you guys have made [00:22:07] where, like, [00:22:08] anyone, like, [00:22:09] human or AI [00:22:10] can now build [00:22:11] real models [00:22:12] on your platform. [00:22:15] Tom, [00:22:15] I kid you not, [00:22:16] when my team [00:22:17] told me that [00:22:17] you had an engineer [00:22:18] working on MI-355, [00:22:20] I said, [00:22:21] hey, [00:22:21] like, [00:22:22] are we, like, [00:22:22] making sure we help them? [00:22:24] And they're like, [00:22:24] they say they don't need [00:22:25] any help. [00:22:26] They got it all covered. [00:22:27] And I was like, [00:22:28] hey, [00:22:28] that's a wonderful story. [00:22:30] Thank you for that. [00:22:31] Now, [00:22:32] we're also super excited [00:22:33] about the work [00:22:34] we're doing together [00:22:35] because, [00:22:35] you know, [00:22:35] one of the most valuable [00:22:36] parts of this partnership [00:22:37] is, [00:22:38] you know, [00:22:38] Claude [00:22:39] and using Claude [00:22:40] within AMD. [00:22:41] Like, [00:22:41] we're big believers [00:22:42] in the fact that, [00:22:43] you know, [00:22:43] AI is this force multiplier [00:22:45] and the more we're able [00:22:46] to have the leading [00:22:48] foundational models [00:22:49] really familiar [00:22:50] with the AMD architecture, [00:22:51] you know, [00:22:51] the better we're going [00:22:53] to be able to service [00:22:53] the overall market. [00:22:54] So, [00:22:54] can you just, [00:22:55] you know, [00:22:55] talk a little bit [00:22:56] about how you're seeing [00:22:57] Claude really evolve [00:22:58] in, you know, [00:22:59] these types of, [00:23:00] you know, [00:23:00] highly engineering [00:23:01] specific workloads? [00:23:02] Yeah. [00:23:03] So, [00:23:03] I think this is a great place [00:23:05] for collaboration [00:23:06] where, [00:23:06] in turn, [00:23:07] like, [00:23:08] the bread and butter [00:23:08] of Claude [00:23:09] is doing software engineering [00:23:10] and now more and more [00:23:11] we're seeing it [00:23:12] help out with [00:23:13] the type of workloads [00:23:14] that you guys are doing also, [00:23:15] like design, [00:23:17] layout, [00:23:18] which is not quite [00:23:19] the normal software engineering [00:23:20] but is adjacent [00:23:21] and I do think [00:23:23] that that is the place [00:23:25] where we can work together [00:23:26] to make the next generation [00:23:28] of chips even better. [00:23:30] Yeah, [00:23:30] I think, [00:23:31] and also on the software side, [00:23:33] I think the work [00:23:33] that we've seen [00:23:34] with Claude [00:23:34] and the kernel development [00:23:36] has been just incredible. [00:23:39] So, [00:23:39] you know, [00:23:40] Tom, [00:23:40] one of the things [00:23:41] for our audience here, [00:23:42] we like to talk [00:23:43] about the present [00:23:43] but it's actually, [00:23:44] most of us [00:23:45] are working on the future. [00:23:46] so this is really [00:23:48] the beginning of, [00:23:49] you know, [00:23:49] what we believe [00:23:50] is a very strong [00:23:52] multi-year, [00:23:53] you know, [00:23:53] partnership. [00:23:54] Can you talk [00:23:55] a little bit about, [00:23:56] number one, [00:23:56] what are you most excited [00:23:57] about in the industry [00:23:58] and then, [00:24:00] number two, [00:24:00] like, [00:24:00] what can we do more together [00:24:01] to really make sure [00:24:05] that we're satisfying [00:24:06] those, [00:24:07] you know, [00:24:07] key opportunities? [00:24:08] Yeah, [00:24:09] that's a great question. [00:24:10] So, [00:24:12] I think the working together [00:24:14] for the scale-up [00:24:16] seems like [00:24:17] it's probably [00:24:18] the biggest thing [00:24:19] that I'm most excited for [00:24:20] where it does seem [00:24:21] like we see consistently [00:24:23] using more, [00:24:25] better compute [00:24:26] results in better models [00:24:28] that then add more value [00:24:30] to all the folks [00:24:30] using it downstream. [00:24:31] So, [00:24:32] I think that's [00:24:32] the biggest thing [00:24:33] and then I think [00:24:34] one thing now [00:24:35] that we're investing [00:24:36] more and more [00:24:37] with you guys too [00:24:37] and I think [00:24:38] we can continue to [00:24:40] is security [00:24:41] where that's a place [00:24:42] where now [00:24:43] as the machines [00:24:44] are getting more complex, [00:24:46] we can work together [00:24:47] to secure things [00:24:48] at the chip layer, [00:24:49] the server level, [00:24:50] the rack, [00:24:51] the whole network [00:24:52] and that'll help out [00:24:53] like, [00:24:54] not just us [00:24:55] but the entire industry [00:24:56] stay safe. [00:24:56] No, [00:24:57] I think you're [00:24:57] completely right. [00:24:58] We are super focused [00:24:59] on the idea of, [00:25:01] you know, [00:25:01] how can we use AI [00:25:02] to improve every aspect [00:25:03] of our product [00:25:04] whether it's, [00:25:05] you know, [00:25:06] performance, [00:25:06] power, [00:25:07] software, [00:25:07] and security [00:25:08] as you said. [00:25:09] So, [00:25:10] really, [00:25:10] Tom, [00:25:11] huge thank you [00:25:12] for the, [00:25:13] you know, [00:25:14] the real, [00:25:14] the opportunity [00:25:15] to work very closely [00:25:16] with you and your team. [00:25:16] Our team loves [00:25:17] working with you guys [00:25:18] and we look forward [00:25:19] to everything [00:25:20] that we're going [00:25:20] to do together. [00:25:21] Thank you so much. [00:25:22] Thank you. [00:25:23] Thank you. [00:25:23] Thank you. [00:25:23] All right. [00:25:30] So, [00:25:31] look, [00:25:31] Tom talked about [00:25:32] how important [00:25:33] performance is [00:25:34] and we have been [00:25:36] laser focused [00:25:37] on ensuring [00:25:38] that the real world [00:25:39] performance [00:25:39] of MI455 [00:25:41] and Helios [00:25:41] really comes true [00:25:43] when we look [00:25:44] at, [00:25:44] you know, [00:25:45] system results. [00:25:46] So, [00:25:46] let's take a look [00:25:47] at some of those results. [00:25:49] So, [00:25:50] here's an example. [00:25:51] We're running [00:25:51] DeepSeq V4 Flash [00:25:53] one of the latest [00:25:54] reasoning models [00:25:55] and we are actually [00:25:56] comparing MI455 [00:25:58] to MI355 [00:25:59] and, you know, [00:26:00] the key thing [00:26:01] is with every generation [00:26:02] we want to make [00:26:03] huge leaps [00:26:04] in performance. [00:26:05] So, [00:26:05] we're seeing that [00:26:06] at lower concurrency [00:26:07] that MI455 [00:26:09] delivers 4x [00:26:10] the throughput [00:26:11] of MI355 [00:26:12] but, [00:26:13] clearly, [00:26:14] what you see [00:26:14] with, [00:26:15] you know, [00:26:15] large users [00:26:16] with, [00:26:17] across the board [00:26:18] that you want [00:26:19] to increase [00:26:20] the number [00:26:20] of simultaneous users, [00:26:21] you want to increase [00:26:22] the amount of work [00:26:22] that you can do [00:26:23] at the same time. [00:26:24] And so, [00:26:25] at the highest concurrency, [00:26:26] we're seeing Helios [00:26:27] deliver up to 34 times [00:26:30] more throughput [00:26:31] than our prior generation. [00:26:33] And what that means [00:26:34] is more users, [00:26:36] much faster response, [00:26:37] and much better efficiency [00:26:39] at scale. [00:26:40] And performance [00:26:41] is just one aspect [00:26:43] of it, right? [00:26:43] The other aspect [00:26:44] and, you know, [00:26:45] what I should have mentioned [00:26:46] when Tom was on stage [00:26:47] is every one of us [00:26:48] in the enterprise [00:26:49] is seeing our AI budgets [00:26:50] go up every single month. [00:26:52] Is that right? [00:26:53] Are you seeing some of that? [00:26:55] And so, [00:26:56] customers really care [00:26:57] about cost per token. [00:26:59] And with every generation [00:27:00] of Instinct, [00:27:01] we're driving [00:27:02] the cost per token down. [00:27:04] And that's [00:27:04] with more memory, [00:27:06] more bandwidth, [00:27:06] and much more compute. [00:27:08] And we're taking [00:27:09] another major step [00:27:10] with MI455, [00:27:12] delivering up to 18 times [00:27:14] more tokens per dollar [00:27:16] so customers can serve [00:27:17] far more users [00:27:18] with the same investment. [00:27:20] And as you think about [00:27:22] how that comes together [00:27:23] in a rack, [00:27:25] many of our data centers [00:27:26] right now [00:27:27] are really limited by power. [00:27:28] So power is kind of [00:27:29] the maximum limiter. [00:27:31] And so we tested Helios [00:27:32] over a full range [00:27:34] of workloads [00:27:35] where you have [00:27:36] a fixed rack power [00:27:37] compared to the competition. [00:27:39] And what we're seeing [00:27:40] is across the highest [00:27:42] throughput workloads, [00:27:43] the most interactive applications, [00:27:44] all of these leading [00:27:45] inference workloads, [00:27:47] we're seeing Helios [00:27:48] an average of 10% to 15% [00:27:50] more performance [00:27:51] than the competition. [00:27:53] And that comes [00:27:54] with all of the system work [00:27:55] that we've done [00:27:56] in the overall system. [00:27:59] And how that translates [00:28:00] to a customer [00:28:00] is cost. [00:28:03] What we expect [00:28:03] is the Helios rack [00:28:05] delivers more performance [00:28:07] and it also delivers [00:28:08] up to 30% more tokens [00:28:11] per dollar [00:28:11] than the competition. [00:28:13] So I think that's [00:28:14] a pretty good [00:28:15] overall value proposition. [00:28:24] Now, I'm happy to say [00:28:26] that the customer demand [00:28:27] for Helios [00:28:28] is extremely strong. [00:28:29] From the largest AI labs [00:28:31] to hyperscale [00:28:32] and neocloud providers, [00:28:34] we are working [00:28:34] across the entire ecosystem [00:28:36] to enable Helios [00:28:37] together with our OEM [00:28:39] and ODM partners. [00:28:41] And one of our deepest [00:28:43] and earliest partners [00:28:44] deploying Helios [00:28:45] is OpenAI. [00:28:47] To talk about [00:28:48] where AI is headed [00:28:49] and the work [00:28:49] we're doing together, [00:28:51] please welcome [00:28:51] to the stage [00:28:52] OpenAI's [00:28:53] Head of Infrastructure, [00:28:54] Sachin Khati. [00:29:01] Thank you, Sachin Khaki. [00:29:03] Thank you, Sachin. [00:29:04] It's so great [00:29:05] to have you here [00:29:06] and most importantly, [00:29:08] we are so thankful [00:29:09] for the partnership [00:29:09] with OpenAI. [00:29:11] We've been through [00:29:11] so much [00:29:12] on this journey together. [00:29:13] So you're really [00:29:15] operating on the frontier [00:29:16] of AI. [00:29:17] Can you just tell us [00:29:18] a little bit [00:29:19] of what you're seeing [00:29:20] and just what compute [00:29:22] looks like for you? [00:29:23] Thanks. [00:29:24] Great to be here. [00:29:25] I was thinking [00:29:26] about this event [00:29:27] and over the last [00:29:28] two years, [00:29:29] it feels like [00:29:29] the attendance [00:29:30] at this event [00:29:31] is following [00:29:32] scaling loss [00:29:33] and it's been doubling [00:29:34] every year [00:29:35] over the last [00:29:36] two years. [00:29:37] But more seriously [00:29:38] to your question, [00:29:40] we've always [00:29:41] predicted that [00:29:42] models will evolve [00:29:44] where they will start [00:29:45] as chatbots, [00:29:46] then they become [00:29:47] reasoners, [00:29:48] then they become [00:29:48] agents, [00:29:49] then they become [00:29:49] interns. [00:29:50] and the trajectory [00:29:52] has kept up [00:29:53] with that. [00:29:53] So we are all [00:29:54] seeing that models [00:29:55] are becoming [00:29:56] a lot more capable, [00:29:57] more agentic [00:29:58] and able to do [00:29:59] very long running [00:30:00] tasks, [00:30:01] let them go [00:30:02] and they just [00:30:02] take care of it. [00:30:03] And so all that [00:30:04] really means is [00:30:06] we need to keep [00:30:07] scaling compute. [00:30:08] The more compute [00:30:09] we scale [00:30:10] and throw at training, [00:30:12] the more capabilities [00:30:13] emerge. [00:30:14] The more compute [00:30:14] we scale [00:30:15] and provide it [00:30:17] to distributing [00:30:18] intelligence to the world, [00:30:19] the more people [00:30:19] want to do with it. [00:30:20] and this is what [00:30:22] we see with [00:30:22] exploding token [00:30:23] budgets. [00:30:24] We've seen [00:30:25] internally, [00:30:26] for example, [00:30:26] in OpenAI [00:30:27] that not just [00:30:28] engineering, [00:30:29] but every aspect [00:30:30] of the enterprise [00:30:31] is beginning [00:30:32] to use agents [00:30:33] for every part [00:30:33] of the work. [00:30:34] This is why [00:30:35] we need platforms [00:30:36] like yours [00:30:37] that can scale [00:30:39] with how quickly [00:30:40] the capabilities [00:30:41] of these models [00:30:41] are scaling [00:30:42] as well as [00:30:43] how quickly [00:30:43] the whole world [00:30:44] is beginning [00:30:45] to embrace [00:30:46] what these things [00:30:46] can do. [00:30:47] Well, Sachin, [00:30:48] I don't think [00:30:48] I've ever [00:30:49] spoken to you [00:30:51] where you haven't [00:30:51] asked me [00:30:52] for more compute. [00:30:53] So he's actually [00:30:54] very, very consistent. [00:30:56] Look, our teams [00:30:57] have been working [00:30:58] really closely together. [00:30:59] I think it's been [00:31:00] a true journey [00:31:01] when we think [00:31:01] about just all [00:31:02] of the technology [00:31:03] that we've been doing. [00:31:04] You know, [00:31:05] last year, [00:31:05] we announced [00:31:06] a really, [00:31:07] you know, [00:31:07] landmark partnership [00:31:09] where you're going [00:31:10] to deploy [00:31:11] up to six gigawatts [00:31:13] of AMD infrastructure. [00:31:14] You're one of the first, [00:31:15] I think you're the first [00:31:16] to actually have [00:31:17] MI-455 racks [00:31:19] a few months ago. [00:31:20] So can you talk [00:31:21] a little bit [00:31:21] about the journey? [00:31:23] No, it's been [00:31:23] a phenomenal partnership. [00:31:25] As you mentioned, [00:31:27] we were first [00:31:28] to bet on AMD [00:31:29] and we are really thrilled [00:31:30] with how that bet [00:31:31] is turning out for us. [00:31:33] Got started on MI-300 [00:31:35] and expanded [00:31:36] very quickly [00:31:37] to 355. [00:31:38] And it's been [00:31:39] really fun [00:31:41] for our team [00:31:42] to see how [00:31:42] it was possible [00:31:44] to optimize [00:31:45] the software stack, [00:31:46] the networking [00:31:47] and everything [00:31:47] and deploy [00:31:48] our models [00:31:49] running on AMD [00:31:51] infrastructure. [00:31:52] So as Lisa mentioned, [00:31:54] so we just got [00:31:55] hands on the Helios [00:31:56] tracks three months ago [00:31:57] and the collaboration [00:31:58] has gone [00:31:59] to the next level. [00:32:00] Our engineers [00:32:00] are working [00:32:01] side by side [00:32:02] with AMD engineers [00:32:04] to optimize [00:32:05] the software stack [00:32:06] and run GVT class [00:32:07] workloads [00:32:08] on Helios [00:32:09] already. [00:32:10] Phil will probably [00:32:11] on stage later [00:32:12] talking about [00:32:13] how productive [00:32:14] and fast [00:32:16] that collaboration is. [00:32:17] So we are really [00:32:18] excited about [00:32:19] the capabilities [00:32:20] Helios is already [00:32:21] showing us [00:32:22] and we expect [00:32:24] that we'll be [00:32:24] deploying Helios [00:32:25] at massive scale [00:32:26] starting towards [00:32:27] the end of this year [00:32:28] and then accelerating [00:32:29] throughout 2027. [00:32:31] Like I said, [00:32:31] I know, [00:32:32] more faster. [00:32:33] We need it earlier. [00:32:36] So, [00:32:37] Sachin, [00:32:37] the other thing [00:32:38] is, you know, [00:32:39] one of the most [00:32:39] exciting parts [00:32:40] of our collaboration [00:32:41] is actually [00:32:41] some of the research [00:32:42] work. [00:32:43] So, you know, [00:32:43] very thankful [00:32:44] for the opportunity [00:32:45] to spend, [00:32:47] you know, [00:32:47] really bringing [00:32:47] our best engineers [00:32:48] together with your [00:32:49] researchers [00:32:50] to talk about [00:32:51] how AI [00:32:53] should be, [00:32:53] you know, [00:32:53] really implemented [00:32:54] and used [00:32:55] for, you know, [00:32:56] future both silicon [00:32:57] and software. [00:32:58] So can you talk [00:32:59] a little bit [00:32:59] about that work? [00:33:01] Yeah, [00:33:01] I think one of [00:33:03] our key bets [00:33:04] and I think [00:33:05] for a lot of us [00:33:06] in AI [00:33:06] is about recursion, [00:33:08] right? [00:33:09] AI recursively [00:33:11] improving the systems [00:33:12] it needs to run on [00:33:13] and eventually, [00:33:15] obviously, [00:33:15] AI doing the research [00:33:16] itself, right? [00:33:18] Today, [00:33:18] expert engineers [00:33:19] spend an enormous [00:33:20] amount of time [00:33:20] optimizing kernels, [00:33:23] figuring out [00:33:23] how to translate [00:33:24] new models [00:33:25] to run them [00:33:26] efficiently [00:33:26] on different hardware, [00:33:28] dealing with [00:33:28] all the usual things [00:33:29] you think about, [00:33:30] compilers, [00:33:30] communications, [00:33:32] libraries, [00:33:32] system configurations. [00:33:33] What we've been [00:33:35] super excited [00:33:36] and encouraged [00:33:37] by the work [00:33:38] we are doing [00:33:39] with you [00:33:39] is now [00:33:41] AI showing [00:33:42] the potential [00:33:43] and demonstrating [00:33:44] it in production [00:33:45] to automate [00:33:46] a significant portion [00:33:47] of the work [00:33:48] when anyone programs [00:33:49] to an AMD GPU. [00:33:50] So it dramatically [00:33:52] has shortened [00:33:53] the path [00:33:53] from a new model [00:33:55] coming up [00:33:55] to an efficient [00:33:57] production deployment [00:33:58] and something [00:33:59] that we can [00:34:00] much more easily [00:34:01] tune to changing [00:34:02] workload requirements, [00:34:03] right? [00:34:03] And that's [00:34:04] a big deal. [00:34:05] So one of the other [00:34:06] things that's [00:34:07] super exciting for us [00:34:08] is because AMD [00:34:09] is built around [00:34:10] an open software [00:34:10] ecosystem, [00:34:11] all these things [00:34:12] that we are now [00:34:14] inventing internally [00:34:15] for our own use, [00:34:17] we believe [00:34:18] we can also bring [00:34:19] it to the whole world [00:34:20] and allow the whole world [00:34:21] to use it [00:34:22] for all of the AI [00:34:23] models out there, [00:34:24] not just our models. [00:34:26] I think super excited [00:34:27] about that. [00:34:28] I think the idea [00:34:28] is that it's really [00:34:30] the tide [00:34:30] that lifts all boats, [00:34:31] right? [00:34:32] As we work, [00:34:33] it helps our work together, [00:34:34] but it really helps [00:34:34] the overall ecosystem. [00:34:36] So a little bit [00:34:37] of your crystal ball. [00:34:39] Looking forward, [00:34:40] I think we all know [00:34:41] that compute [00:34:41] is super critical. [00:34:43] What do you need [00:34:44] from the next generation [00:34:45] AI infrastructure [00:34:46] and what can we do [00:34:48] as really a key partner [00:34:51] to help you [00:34:52] accomplish that mission? [00:34:53] If you haven't [00:34:54] heard already, [00:34:55] I need more compute [00:34:56] more quickly. [00:34:58] At least he's consistent. [00:35:01] I think the thing [00:35:05] that the whole industry [00:35:06] is realizing [00:35:07] is it's a systems problem [00:35:08] at a data center scale. [00:35:10] So I think we obviously [00:35:12] very quickly internalized [00:35:13] that AI was a rack level [00:35:14] problem, [00:35:15] but now it is very clear [00:35:17] that it's not just [00:35:18] a rack level problem, [00:35:19] it's a rack in a data center. [00:35:20] So the whole data center [00:35:21] is a system. [00:35:23] And we have to think [00:35:24] about how we design [00:35:25] the future based on [00:35:26] what we know [00:35:27] about our workloads, [00:35:28] what's coming down the line, [00:35:29] but co-design it with you [00:35:31] from CPUs to GPUs [00:35:33] to memory, [00:35:34] networking, [00:35:34] storage, [00:35:35] power distribution [00:35:36] in a data center [00:35:37] and the cooling systems [00:35:38] that go with it. [00:35:39] These are all intertwined [00:35:40] and they all have [00:35:42] to be co-designed. [00:35:43] And one of the really fun [00:35:46] parts about what we do [00:35:47] with you and your team [00:35:49] is how easy [00:35:50] and productive it is [00:35:52] for us to work with you [00:35:53] and give you insight [00:35:54] into how we expect [00:35:56] the world to change [00:35:56] in the future [00:35:57] from what our workload [00:35:58] needs are, [00:35:59] what our models are, [00:36:00] and how quickly [00:36:01] your team responds [00:36:02] in turning those insights [00:36:04] into better chips, [00:36:06] better systems, [00:36:06] better software. [00:36:08] So that's worked well, [00:36:09] really well with MI400 [00:36:10] and left shifted [00:36:11] a lot of the engineering [00:36:14] work that we needed to do [00:36:15] to bring this up [00:36:16] online quickly. [00:36:17] And that's why [00:36:18] we are very confident [00:36:19] we can productionize [00:36:20] those systems [00:36:21] very quickly. [00:36:22] And so really looking forward [00:36:24] to accelerating that [00:36:25] with MI500 and beyond [00:36:26] and increasingly using AI [00:36:29] to actually do that work [00:36:30] for us rather than [00:36:32] being limited by humans. [00:36:34] That's fantastic, Sachin. [00:36:36] And look, [00:36:36] we really do recognize [00:36:38] that input is so, [00:36:39] so valuable. [00:36:40] So I have to say [00:36:41] thank you again [00:36:42] for just the extraordinary [00:36:43] partnership, [00:36:44] the amount of work [00:36:45] that our joint teams [00:36:46] are doing together. [00:36:47] And we couldn't be more [00:36:48] excited about the opportunities [00:36:50] in front of us. [00:36:51] Thank you, Sachin. [00:36:52] Thank you so much. [00:36:52] Yes. [00:36:58] All right. [00:36:59] So now let's turn [00:37:01] to the world of CPUs. [00:37:03] There's been a lot of talk [00:37:04] about CPUs recently, [00:37:06] and I can say [00:37:07] that this has been [00:37:07] the foundation [00:37:08] of our AMD data center [00:37:10] strategy for the longest time. [00:37:12] We launched Epic in 2017, [00:37:14] and we've been [00:37:15] on a very, very clear mission. [00:37:18] Every generation, [00:37:19] more performance, [00:37:20] more efficiency, [00:37:21] more capability [00:37:23] for a wider range [00:37:24] of workloads. [00:37:26] Now, Naples [00:37:26] was our first generation. [00:37:28] It established [00:37:28] the foundation, [00:37:30] but Rome and Milan [00:37:31] really changed [00:37:32] the economics [00:37:32] in the data center. [00:37:33] We did more [00:37:34] than what people expected [00:37:36] in terms of adding [00:37:37] core count [00:37:38] and adding throughput. [00:37:39] And then Turin set, [00:37:41] Genoa expanded [00:37:42] our performance, [00:37:43] and then Turin set [00:37:44] a new bar [00:37:45] for density, [00:37:46] throughput, [00:37:46] and total cost [00:37:47] of ownership. [00:37:48] That's really been [00:37:49] the Epic formula, [00:37:50] a clear roadmap [00:37:51] and very consistent execution [00:37:54] and leadership [00:37:55] that actually grows [00:37:57] with every generation. [00:37:58] This is why, [00:37:59] if you look today, [00:38:01] 5th Gen Epic Turin [00:38:02] is the best server CPU [00:38:03] in the world. [00:38:04] With up to 192 cores [00:38:07] and 384 threads, [00:38:09] Turin delivers [00:38:10] leadership performance [00:38:11] across a broad range [00:38:13] of cloud, [00:38:13] enterprise, [00:38:14] and HPC workloads. [00:38:16] But the thing [00:38:17] about the data center [00:38:18] is compute [00:38:19] never stands still. [00:38:21] With agentic AI, [00:38:23] it's actually creating [00:38:24] a whole new class [00:38:25] of infrastructure, [00:38:26] and the CPU matters [00:38:27] now more than ever. [00:38:29] And that's exactly [00:38:30] what we built Venice for. [00:38:33] So there's been [00:38:34] a lot of talk [00:38:35] about agentic AI, [00:38:36] but it's actually [00:38:36] a very new field. [00:38:38] And what we see [00:38:39] is that server computing [00:38:41] is actually splitting [00:38:42] across a couple [00:38:43] of different workloads. [00:38:45] So the first section, [00:38:47] you know, [00:38:47] is GPU servers, [00:38:48] and that's fairly well known. [00:38:50] That's where the GPU's job, [00:38:53] the CPU's job, [00:38:53] is basically to drive [00:38:54] the GPUs. [00:38:56] And so the job there [00:38:57] is all about speed. [00:38:59] We want the highest [00:39:00] frequency cores, [00:39:02] we want the fastest I.O. [00:39:03] and we want to keep [00:39:04] the GPUs fully fed. [00:39:06] Now the biggest middle part, [00:39:08] this is the largest growth, [00:39:09] is actually in agent servers, [00:39:11] or what we call [00:39:11] agent sandboxes. [00:39:13] And this is actually [00:39:14] a whole new class [00:39:16] of workload. [00:39:16] Here you execute code, [00:39:18] you actually call [00:39:19] a bunch of tools, [00:39:20] you actually query data [00:39:22] that's outside the model, [00:39:23] you have a whole bunch [00:39:24] of things around it. [00:39:26] And in this place, [00:39:28] the priority [00:39:28] is actually density. [00:39:30] So we want [00:39:31] the highest performing [00:39:32] cores per watt [00:39:33] to run thousands [00:39:35] of agents at once. [00:39:37] And then we have [00:39:38] our traditional [00:39:39] general purpose servers [00:39:40] that run the enterprise. [00:39:42] And here, [00:39:43] you're going to have [00:39:43] a diversity of workloads. [00:39:45] It's all about efficiency, [00:39:46] it's running the applications, [00:39:47] it's running databases, [00:39:48] it's running data services, [00:39:50] and all of that [00:39:51] is, you know, [00:39:52] what we would traditionally say [00:39:54] is, you know, [00:39:55] general purpose server. [00:39:56] And so with these [00:39:57] different workloads, [00:39:58] you actually need [00:39:59] the right CPU [00:40:00] for the right workload. [00:40:02] And Epic [00:40:03] is the only CPU portfolio [00:40:05] that leads [00:40:06] across all three. [00:40:09] So Venice [00:40:10] is our newest [00:40:11] and most advanced [00:40:12] server CPU family, [00:40:13] and it's actually designed [00:40:14] for the agentic era. [00:40:16] We're extending [00:40:17] TURN's leadership [00:40:18] across every single metric. [00:40:20] That is performance, [00:40:22] that's efficiency, [00:40:23] that's TCO for cloud, [00:40:25] that's enterprise efficiency, [00:40:27] that's HPC workloads. [00:40:29] And Venice starts [00:40:30] with our all-new [00:40:31] Zen 6 core. [00:40:33] We have higher IPC, [00:40:35] higher frequency, [00:40:36] and that delivers [00:40:37] up to 1.8 times [00:40:39] more performance [00:40:40] than TURN. [00:40:41] That is one of the largest [00:40:43] generational gains. [00:40:44] I mean, [00:40:44] we have been doing this [00:40:45] for six generations, [00:40:47] but this is one [00:40:48] of the largest [00:40:48] generational gains [00:40:49] in the history of Epic. [00:40:51] It's 203 billion [00:40:53] transistors built [00:40:54] on TSMC's newest [00:40:55] 2 nanometer process, [00:40:56] and it uses [00:40:58] a next-generation [00:40:59] chiplet that supports [00:41:00] up to 512 threads [00:41:02] per socket. [00:41:03] And this is actually [00:41:04] a really important point. [00:41:05] I'm going to show you [00:41:05] why in a few minutes. [00:41:08] This gives us [00:41:08] the highest compute density, [00:41:10] but it also allows us [00:41:11] to double both the memory [00:41:13] and the I-O bandwidth. [00:41:15] So, I have a few more [00:41:16] chips to show you, [00:41:17] if that's all right. [00:41:19] And what's clear [00:41:21] is Venice is not one chip. [00:41:25] Venice is actually [00:41:26] an entire family [00:41:27] of chips. [00:41:28] And this is where [00:41:29] our chiplet architecture [00:41:30] becomes a huge, [00:41:32] huge advantage. [00:41:33] It lets us take [00:41:35] the Zen 6 architecture [00:41:36] and build a whole portfolio [00:41:38] around it. [00:41:39] So, let's start [00:41:40] with this guy. [00:41:43] This is Venice HF. [00:41:45] This is the highest [00:41:47] performing CPU [00:41:49] that drives [00:41:50] AI host nodes. [00:41:52] It has [00:41:53] eight compute chiplets, [00:41:54] each with 12 cores, [00:41:56] running at up to 5 gigahertz [00:41:58] with the I-O [00:41:59] and memory bandwidth [00:41:59] to keep the GPUs [00:42:01] completely fed. [00:42:02] And this is the CPU [00:42:03] that I showed you [00:42:04] before shipping [00:42:05] inside Helios. [00:42:13] Now, [00:42:14] this is its [00:42:17] bigger brother. [00:42:19] It's Venice 256 core. [00:42:21] This is the chip [00:42:22] that's built [00:42:23] for EGentex sandboxes. [00:42:25] And it also has [00:42:26] eight compute chiplets, [00:42:28] but each of the chiplets [00:42:29] now has 32 cores. [00:42:31] So, it scales [00:42:32] all the way [00:42:32] to 512 threads. [00:42:34] And it is [00:42:35] the highest compute density [00:42:36] in the industry. [00:42:43] Customers get [00:42:44] more agent capacity [00:42:45] per watt, [00:42:46] per dollar, [00:42:47] per rack, [00:42:48] bar none. [00:42:52] Okay, [00:42:53] and then [00:42:53] one more member [00:42:55] of the family. [00:42:56] This one is [00:42:57] Venice 128 core. [00:42:59] And this guy [00:43:00] is optimized [00:43:00] for enterprise [00:43:01] and general purpose servers. [00:43:03] This is the same [00:43:04] leadership performance [00:43:05] per core, [00:43:06] but it's in a lower cost, [00:43:07] lower power design [00:43:08] for enterprise applications. [00:43:11] And this one [00:43:12] is really tuned [00:43:12] to give customers [00:43:13] the best performance [00:43:14] per dollar [00:43:15] across a wide range [00:43:16] of configurations. [00:43:17] So, you can see, [00:43:18] really, [00:43:19] the power of the family [00:43:21] as they come together. [00:43:23] And in addition, [00:43:24] with these guys, [00:43:25] I can tell you [00:43:25] that next year, [00:43:26] we're adding [00:43:27] a few other chips [00:43:28] to the family. [00:43:29] We're adding [00:43:29] Verano, [00:43:30] which is the next generation [00:43:31] for AI host nodes. [00:43:33] And what that adds [00:43:34] is a more power-efficient, [00:43:36] low-power memory, [00:43:37] as well as [00:43:38] faster interconnect [00:43:39] between the CPU [00:43:40] and the GPU. [00:43:41] And then we're also adding [00:43:42] for our high-performance computing, [00:43:44] our technical computing, [00:43:46] Venice X, [00:43:47] which is using [00:43:48] our 3D V-cache [00:43:49] stacking technology. [00:43:50] And we're really [00:43:51] the first [00:43:51] to stack [00:43:52] memory chiplets [00:43:53] right below [00:43:54] the compute chiplets [00:43:55] to accelerate performance [00:43:56] even for those [00:43:58] types of workloads. [00:43:59] So, you can see [00:44:00] one architecture, [00:44:01] a CPU [00:44:02] is not a CPU. [00:44:04] There are many [00:44:04] different types [00:44:05] of CPUs. [00:44:06] And for us, [00:44:07] one architecture [00:44:08] spans dozens [00:44:09] of different chips. [00:44:10] And that breadth [00:44:11] is someone [00:44:12] that no other [00:44:13] CPU vendor [00:44:14] can match. [00:44:15] And that's [00:44:15] what makes [00:44:16] 6th Gen Epic [00:44:17] the best CPU [00:44:19] in the data center. [00:44:19] Now, you're [00:44:27] going to bear [00:44:27] with me [00:44:28] because I'm [00:44:28] going to take [00:44:29] you through [00:44:29] some performance [00:44:29] because the numbers [00:44:30] are just incredible. [00:44:32] So, looking [00:44:33] at AI workloads, [00:44:34] right? [00:44:35] Again, you have [00:44:35] different use cases [00:44:37] and different workloads. [00:44:38] You know, [00:44:39] starting first [00:44:40] with GPU servers. [00:44:41] Many of the [00:44:42] largest frontier [00:44:43] models require [00:44:44] you to move [00:44:44] data between [00:44:46] the CPU [00:44:46] and the GPU. [00:44:47] And in that [00:44:48] scenario, [00:44:48] Venice moves [00:44:49] the data [00:44:49] faster than [00:44:50] the best [00:44:51] competitive [00:44:52] x86 processor [00:44:53] delivering up [00:44:54] to 1.8 times [00:44:56] more tokens [00:44:56] per second. [00:44:58] And when you [00:44:58] look at [00:44:59] the agentic [00:45:00] CPU servers [00:45:00] and sandboxes, [00:45:02] because of our [00:45:02] density, [00:45:03] Venice delivers [00:45:03] more than twice [00:45:04] the agents [00:45:05] per watt [00:45:05] for agent [00:45:06] sandboxes. [00:45:08] And for [00:45:08] general purpose [00:45:09] applications, [00:45:09] in this case, [00:45:10] we're standardizing [00:45:11] on a 100-kilowatt [00:45:12] rack, [00:45:13] Venice delivers [00:45:14] more than twice [00:45:15] the performance [00:45:16] per watt. [00:45:17] So, pretty [00:45:18] incredible. [00:45:18] results. [00:45:21] And the gap [00:45:22] is even wider [00:45:23] if you look [00:45:24] at ARM processors. [00:45:26] So, if you look [00:45:26] at the leading [00:45:27] ARM processors, [00:45:28] when we're [00:45:28] comparing at the [00:45:30] chip level, [00:45:31] Venice supports [00:45:32] up to 2.8 times [00:45:33] more agents [00:45:34] per watt. [00:45:35] And when you [00:45:36] look at the [00:45:36] rack level, [00:45:37] Venice is delivering [00:45:38] up to 3.3 times [00:45:40] more performance [00:45:41] per watt. [00:45:42] And what that [00:45:43] means is, [00:45:44] at the data [00:45:45] center scale, [00:45:45] you can just [00:45:47] get a lot [00:45:47] more agents [00:45:48] in the same [00:45:50] power envelope. [00:45:57] Now, [00:45:58] there's been a lot [00:45:59] of talk [00:46:00] about what matters [00:46:01] most for [00:46:01] agentic AI, [00:46:03] with some saying [00:46:04] that per core [00:46:05] performance under [00:46:05] load is the only [00:46:06] thing that counts. [00:46:08] that's really [00:46:09] only part of the [00:46:10] story. [00:46:11] Because agentic AI [00:46:12] actually runs as a [00:46:13] distributed platform, [00:46:15] and that is across [00:46:16] databases and vector [00:46:17] stores and [00:46:18] orchestration. [00:46:19] And all of that [00:46:20] happens has to happen [00:46:21] at once. [00:46:22] And that requires [00:46:23] not just raw speed, [00:46:25] but you absolutely [00:46:26] need density, [00:46:28] efficiency, [00:46:29] and per core [00:46:30] performance all [00:46:31] together. [00:46:32] And what I'm [00:46:33] happy to say [00:46:33] is that Venice [00:46:34] leads across [00:46:36] every one of [00:46:36] them. [00:46:37] So when you [00:46:38] compare Venice [00:46:39] against the [00:46:40] highest performing [00:46:41] ARM CPU from [00:46:42] our competition, [00:46:43] Epic delivers [00:46:44] 20% higher [00:46:46] per core [00:46:47] performance. [00:46:48] And when we... [00:46:49] I like that [00:46:55] number. [00:46:56] I like that [00:46:57] number. [00:46:58] And when we [00:46:58] pair that with [00:46:59] leadership core [00:47:00] density, [00:47:00] we're delivering [00:47:01] 2.2 times [00:47:02] more performance [00:47:03] at the socket [00:47:03] level. [00:47:04] And so the main [00:47:05] point is whether [00:47:07] a customer needs [00:47:08] per core or [00:47:09] per socket [00:47:10] performance, [00:47:11] Venice is the [00:47:12] leader. [00:47:13] And there's one [00:47:13] more point that [00:47:14] a benchmark [00:47:15] doesn't really [00:47:16] capture. [00:47:17] Because when you [00:47:18] think about what [00:47:18] agentic AI is, [00:47:20] it's really like [00:47:20] you're adding [00:47:21] thousands of [00:47:22] employees to your [00:47:23] enterprise, to [00:47:24] your sort of [00:47:24] traditional [00:47:25] workflow. [00:47:26] And that [00:47:27] software enterprises [00:47:28] have been running [00:47:29] on for years. [00:47:30] The truth [00:47:31] is, that [00:47:32] software runs [00:47:33] on x86. [00:47:35] So Venice [00:47:35] runs all of it. [00:47:37] And this is where [00:47:37] x86 has an [00:47:38] advantage. [00:47:39] You must have [00:47:40] the best [00:47:40] performance, you [00:47:41] must have the [00:47:41] best density, but [00:47:42] having that [00:47:43] software compatibility [00:47:44] is a big, big [00:47:45] plus. [00:47:46] So customers [00:47:47] scale up their [00:47:48] agents on the [00:47:49] stack that they [00:47:49] already have. [00:47:51] And now, let me [00:47:52] just show you what [00:47:53] that means at the [00:47:54] rack level. [00:47:55] So at the rack [00:47:55] level, what we [00:47:56] find is, you know, [00:47:57] density is super [00:47:58] important. [00:47:59] Everyone's trying [00:47:59] to optimize their [00:48:00] data center. [00:48:01] Everyone's trying [00:48:01] to optimize their [00:48:02] power envelope. [00:48:03] And with Venice, [00:48:05] every major [00:48:06] server OEM is [00:48:07] offering a broad [00:48:08] spectrum of racks [00:48:09] for agentic AI. [00:48:10] So what that [00:48:11] means is, we [00:48:12] have the full [00:48:12] spectrum from, [00:48:14] you know, 25,000 [00:48:15] cores to 50,000 [00:48:17] cores, and that's [00:48:18] more agents per [00:48:19] rack, and that [00:48:20] gives customers the [00:48:21] choice. [00:48:22] The key is, every [00:48:23] data center is [00:48:24] different. [00:48:25] And with that, the [00:48:26] flexibility allows [00:48:28] you to pick the [00:48:29] right cooling, the [00:48:30] right space [00:48:31] requirements for what [00:48:32] you're trying to do [00:48:33] in your data center. [00:48:34] So there you have [00:48:35] it. [00:48:36] Venice, the best [00:48:37] server CPU in [00:48:39] the industry. [00:48:46] And I'm very happy [00:48:47] to say also today [00:48:48] that Venice is in [00:48:49] full production. [00:48:51] Customer demand [00:48:51] is incredible. [00:48:53] It's the strongest [00:48:53] we've ever seen [00:48:54] for a new [00:48:55] epic generation. [00:49:01] We're seeing [00:49:02] every major [00:49:03] server OEM, [00:49:04] every major [00:49:05] cloud provider [00:49:05] on track to [00:49:07] begin rolling out [00:49:08] in the fourth [00:49:08] quarter as we [00:49:10] start with making [00:49:11] the broadest [00:49:11] epic launch [00:49:12] we've ever had. [00:49:14] So with that, [00:49:15] I want to turn [00:49:16] to my next guest [00:49:17] who runs some [00:49:18] of the largest [00:49:19] and most advanced [00:49:19] compute infrastructure [00:49:20] in the world. [00:49:21] And it's really [00:49:22] been our privilege [00:49:23] to partner with [00:49:23] them as they've [00:49:24] deployed multiple [00:49:26] generations of both [00:49:27] epic and instinct. [00:49:28] To share more [00:49:29] about our work [00:49:30] together, please [00:49:31] welcome to the [00:49:31] stage, Meta's [00:49:32] head of [00:49:33] infrastructure, [00:49:34] Santosh [00:49:34] Dinardin. [00:49:43] Santosh is a true [00:49:50] friend. [00:49:51] I have to say [00:49:52] it's been a [00:49:53] tremendous opportunity [00:49:54] over the last few [00:49:55] years. [00:49:56] My story about [00:49:57] Santosh is the first [00:49:58] time I sat in his [00:49:59] office. [00:50:00] He said to me, [00:50:00] Lisa, we're going to [00:50:02] deploy a lot of [00:50:03] CPUs. [00:50:05] please make [00:50:05] sure they [00:50:05] work. [00:50:08] And I said, [00:50:08] I will. [00:50:09] I will. [00:50:10] And you [00:50:10] lifted it up. [00:50:11] I understand. [00:50:13] I try. [00:50:13] But look, [00:50:14] Santosh, [00:50:14] Meta operates, [00:50:16] you operate some [00:50:17] of the most [00:50:18] advanced [00:50:18] infrastructure [00:50:18] in the world. [00:50:20] What's happening [00:50:21] in the data [00:50:22] center world? [00:50:22] What's happening [00:50:23] in your world? [00:50:23] Tell us a little [00:50:24] bit about [00:50:25] the architecture [00:50:26] and what you're [00:50:27] working on. [00:50:28] Sure. [00:50:28] First of all, [00:50:29] thank you. [00:50:29] It's awesome. [00:50:30] This is like a [00:50:30] hardware geek's [00:50:31] paradise, right? [00:50:33] You're talking [00:50:34] about transistors, [00:50:35] you're talking [00:50:35] about hardware, [00:50:36] you're talking [00:50:36] about cooling. [00:50:37] This is sort [00:50:38] of the gem. [00:50:38] I like it. [00:50:40] And I think [00:50:41] this is an audience [00:50:41] that's receptive [00:50:42] to it. [00:50:42] I usually talk [00:50:44] to audiences, [00:50:44] they have usually [00:50:45] either not been [00:50:46] to a data center [00:50:47] or not really [00:50:47] seen a chip ever. [00:50:49] So this is [00:50:50] refreshing, [00:50:50] I have to say. [00:50:52] The thing I'll [00:50:52] say about, [00:50:53] see, listen, [00:50:54] we are seeing [00:50:54] demand go [00:50:55] through the roof. [00:50:57] When you look [00:50:57] at inference, [00:50:58] training, [00:50:59] recommendation [00:51:00] system, [00:51:00] that's how [00:51:01] a newsfeed works. [00:51:02] content creation, [00:51:04] the demand for that [00:51:05] is just exponential [00:51:06] right now, right? [00:51:07] And when I think [00:51:08] about sort of [00:51:09] the overall, [00:51:12] Mark has this vision [00:51:13] of delivering [00:51:14] personal super [00:51:15] intelligence [00:51:15] to billions of people. [00:51:17] So if you go [00:51:18] to sort of Facebook [00:51:18] or Instagram [00:51:19] or WhatsApp [00:51:20] or whatever [00:51:20] surface you go to, [00:51:22] the idea is [00:51:22] that we'll meet [00:51:23] you there [00:51:23] and we'll deliver [00:51:24] super personalized [00:51:26] intelligence [00:51:27] right wherever you are. [00:51:29] And that's why [00:51:29] we established [00:51:30] this meta-compute [00:51:31] initiative [00:51:31] and because [00:51:32] it's a top level [00:51:33] that Mark himself [00:51:34] oversees. [00:51:35] Now, [00:51:35] what that means [00:51:36] is that [00:51:36] we are now [00:51:37] moving away [00:51:38] from an area [00:51:38] where we used [00:51:39] to take servers, [00:51:40] optimize it, [00:51:41] right? [00:51:42] This is what [00:51:42] the conversations [00:51:42] we were having [00:51:43] many years ago [00:51:44] that, hey, [00:51:44] I'm deploying a CPU [00:51:45] I really need [00:51:46] to maximize [00:51:46] performance out of it. [00:51:48] It's different now. [00:51:49] You talked about [00:51:50] how we should be [00:51:52] thinking about [00:51:52] the whole system [00:51:53] end to end. [00:51:54] We are at a point [00:51:55] where we have to [00:51:56] look about [00:51:56] the whole data center [00:51:58] and think about [00:51:58] it as one [00:51:59] integrated system. [00:52:00] Servers, [00:52:01] hardware, [00:52:01] networking, [00:52:02] cooling, [00:52:03] power, [00:52:03] all of that [00:52:04] is not optional. [00:52:05] All of that [00:52:06] has to work [00:52:06] together, [00:52:07] right? [00:52:07] And this is where [00:52:08] I think [00:52:08] there's a huge [00:52:09] opportunity [00:52:09] for collaboration [00:52:10] with partners [00:52:11] like AMD [00:52:11] because you have [00:52:13] to now co-design, [00:52:14] co-create systems, [00:52:16] not just deploy [00:52:17] sort of something [00:52:18] that comes [00:52:18] off the shelf, [00:52:19] right? [00:52:20] The other thing [00:52:20] is, by the way, [00:52:21] flexibility. [00:52:22] The thing I really [00:52:23] like about what you [00:52:23] said is that [00:52:24] this is such [00:52:25] a big opportunity [00:52:27] in the industry [00:52:28] right now. [00:52:29] All of us [00:52:29] need to lean in [00:52:30] and sort of [00:52:31] work on this [00:52:31] together. [00:52:32] So it needs [00:52:34] to be open, [00:52:34] it needs [00:52:35] to be heterogeneous [00:52:35] and I think [00:52:37] that's no one [00:52:38] company, [00:52:39] one partner [00:52:39] that works [00:52:40] for any one [00:52:41] of us. [00:52:42] We need to [00:52:42] sort of work [00:52:43] on this together [00:52:43] and that's why [00:52:44] I think [00:52:44] you're one of [00:52:45] our most important [00:52:46] strategic partners. [00:52:47] Thank you, [00:52:48] Santosh. [00:52:48] And look, [00:52:49] we completely [00:52:50] agree with [00:52:50] your philosophy. [00:52:52] Now, [00:52:52] you are actually [00:52:53] one of our [00:52:54] broadest partners [00:52:54] because if you [00:52:55] think about [00:52:55] our work [00:52:56] on CPUs, [00:52:57] GPUs, [00:52:58] you know, [00:52:58] we developed [00:52:59] the RackSkill [00:53:00] OCP systems [00:53:02] together. [00:53:03] You've deployed [00:53:04] lots and lots [00:53:04] of CPUs, [00:53:05] so millions [00:53:05] of CPUs [00:53:06] and, you know, [00:53:07] you're a lead [00:53:08] partner on [00:53:09] Venice. [00:53:09] Can you just [00:53:10] talk a little bit [00:53:10] about sort of [00:53:11] your CPU [00:53:11] and infrastructure [00:53:12] and some [00:53:12] of our work [00:53:13] together? [00:53:14] Sure. [00:53:14] It's been many [00:53:15] years and the [00:53:16] story that Lisa [00:53:17] was saying [00:53:18] was many years [00:53:18] ago. [00:53:18] I think we went [00:53:19] from Milan [00:53:20] to Bergamo [00:53:21] to Turin [00:53:22] now to Venice. [00:53:23] So at least [00:53:24] the fourth [00:53:25] generation, [00:53:26] so long-time [00:53:27] partnership, [00:53:27] obviously. [00:53:28] I actually think [00:53:30] the world is [00:53:30] changing in the [00:53:31] sense that [00:53:32] it used, [00:53:33] the world is [00:53:34] changing, [00:53:34] every year is a [00:53:34] new world these days. [00:53:35] It's like every [00:53:36] month. [00:53:36] Month, exactly. [00:53:38] So what ends up [00:53:38] happening now is [00:53:39] that it's not just [00:53:40] a GPU game [00:53:41] anymore, [00:53:41] it's CPUs [00:53:42] and GPUs. [00:53:43] If anything, [00:53:44] I think CPUs [00:53:44] are becoming [00:53:45] at least as [00:53:45] important, [00:53:46] if not more, [00:53:47] you're looking [00:53:48] at a world [00:53:48] where sort [00:53:49] of the [00:53:49] workforce, [00:53:50] the workload [00:53:51] is changing [00:53:52] because while [00:53:53] you have your [00:53:54] workhorses [00:53:55] and GPUs, [00:53:56] at the end [00:53:56] of the day, [00:53:57] there's [00:53:58] agentic workloads, [00:53:59] you have to run [00:53:59] your tools, [00:54:00] you have to run [00:54:00] your systems, [00:54:01] and all the code [00:54:02] that's being generated [00:54:03] still needs to run [00:54:04] somewhere, right? [00:54:05] So I'm super [00:54:06] excited to work [00:54:07] on Venice. [00:54:07] I think there's a lot [00:54:08] to come there. [00:54:10] And you have [00:54:12] to think again, [00:54:13] this is a theme [00:54:14] that I'm assuming [00:54:15] will go throughout [00:54:16] the presentation [00:54:16] that you have [00:54:17] to start thinking [00:54:18] about CPUs [00:54:18] and GPUs [00:54:19] as conjoint things. [00:54:21] You hand off [00:54:21] workloads, [00:54:22] depending on the workload [00:54:23] you employ [00:54:23] the right hardware. [00:54:24] Yes, [00:54:25] no, [00:54:25] absolutely right. [00:54:26] And we've gotten [00:54:27] a lot of feedback [00:54:28] from your technical team. [00:54:29] I think that's [00:54:29] what I really appreciate [00:54:30] about the partnership [00:54:32] with Meta. [00:54:34] Now, [00:54:35] clearly, [00:54:35] you're deploying [00:54:35] lots of GPUs too, [00:54:37] a lot of accelerators, [00:54:38] and you're also [00:54:39] one of our [00:54:40] deepest partners [00:54:41] on the accelerator side [00:54:42] starting with MI300 [00:54:43] and then we had [00:54:44] a very large [00:54:46] strategic partnership [00:54:48] announced around [00:54:49] up to 6 gigawatts [00:54:50] starting with MI450. [00:54:52] And what we're doing [00:54:53] is quite unique [00:54:53] with Meta. [00:54:54] So can you talk [00:54:55] a little bit about [00:54:55] sort of the evolution [00:54:57] and the MI450 plans? [00:55:01] Again, [00:55:01] multi-year collaboration. [00:55:03] I think we started [00:55:03] with MI300. [00:55:05] We did a bunch. [00:55:06] It had really nice [00:55:07] memory capacity. [00:55:07] I remember talking [00:55:08] to you about that, [00:55:09] right? [00:55:09] and then [00:55:10] there's a bunch [00:55:11] of pipe cleaning [00:55:11] with it. [00:55:12] It was important [00:55:13] to deploy that [00:55:14] at production scale, [00:55:16] sort of get hands-on [00:55:17] experience on both sides. [00:55:19] 350 is the first time [00:55:20] we deployed this [00:55:21] on our ranking [00:55:21] and recommendation systems. [00:55:23] So that's now [00:55:24] getting into [00:55:25] sort of some good numbers. [00:55:27] And 450 is, [00:55:28] I think, [00:55:28] I'm super excited about it. [00:55:29] We just got some [00:55:30] of the racks. [00:55:31] And we are going [00:55:32] to deploy it [00:55:33] across the board [00:55:33] because it gives us [00:55:35] the opportunity [00:55:35] to go and sort [00:55:37] really collaborate. [00:55:38] I think the difference, [00:55:39] big difference [00:55:40] between sort of [00:55:40] the 300, [00:55:41] 350 and 450 [00:55:42] is that we have had [00:55:43] an engineer sitting [00:55:44] in the same room [00:55:45] talking to each other, [00:55:46] co-designing, [00:55:47] co-creating, [00:55:48] like power cooling. [00:55:50] Like I was saying, [00:55:50] it's not just a single thing [00:55:52] and we are not shy [00:55:54] in feedback. [00:55:55] They're not shy, [00:55:56] for sure. [00:55:57] But you're very receptive. [00:55:59] So this is true partnership, [00:56:00] right? [00:56:00] This is the point [00:56:01] about co-designing [00:56:02] and co-creating [00:56:03] that I think. [00:56:04] And like I was saying, [00:56:06] that there's a huge [00:56:07] sort of deployment [00:56:08] coming in. [00:56:09] The more we deploy, [00:56:10] the earlier we co-design, [00:56:12] the better we are. [00:56:13] Yeah, absolutely. [00:56:14] Really, really appreciate [00:56:15] the effort together [00:56:16] on MI450. [00:56:18] Now, I'm going to ask you [00:56:19] also about your crystal ball. [00:56:21] So you look out [00:56:21] into the future [00:56:22] and you see what the next [00:56:24] few years really means [00:56:26] as you push forward [00:56:27] in the AI industry. [00:56:29] What are the biggest challenges [00:56:30] and really opportunities [00:56:32] for us? [00:56:32] This is the industry [00:56:33] ecosystem here. [00:56:34] So what should we be [00:56:35] focused on as an industry? [00:56:37] Listen, all of you [00:56:38] know about the bottlenecks. [00:56:40] You know about the power, [00:56:41] the data centers, [00:56:42] the silicon. [00:56:43] All of these, I think, [00:56:44] are choke points [00:56:45] that I think the industry [00:56:46] is sort of waiting on. [00:56:48] But those are things [00:56:49] I'm pretty confident [00:56:50] will sort of work [00:56:51] our way through. [00:56:51] At the end of the day, [00:56:53] I think that demand [00:56:54] is immense. [00:56:54] People are responding. [00:56:56] But the thing [00:56:57] that I really want [00:56:58] to make sure [00:56:59] all of us realize [00:56:59] is that we should not [00:57:01] think about systems [00:57:02] in isolations. [00:57:04] Like I was saying, [00:57:05] CPUs and GPUs [00:57:06] are one conjoint system. [00:57:07] You should start [00:57:07] thinking about it that way. [00:57:09] We really need [00:57:10] to co-design things early. [00:57:13] See, data centers [00:57:14] takes years to build. [00:57:15] One of my favorite stories [00:57:17] is Mark comes to me [00:57:18] and says, [00:57:18] hey, I want a gig [00:57:19] worth of data centers. [00:57:21] Well, you should have [00:57:21] talked to me two years ago. [00:57:23] It takes time. [00:57:24] It's the same with silicon. [00:57:26] It just takes time. [00:57:27] So we need to be starting [00:57:29] and sitting down [00:57:30] in a room, [00:57:31] co-designing today [00:57:32] for what we need [00:57:33] to deploy in 27 and 28. [00:57:35] That, I think, [00:57:36] is when we truly unlock [00:57:37] the powers of the system. [00:57:39] The systems are amazing. [00:57:40] The specs are amazing, [00:57:41] right? [00:57:41] But think about [00:57:42] how much more [00:57:43] you can get out of it [00:57:44] if you sit in a room [00:57:45] and just co-design it today. [00:57:48] And then in 28, [00:57:49] we'll be having [00:57:49] much better graphs out there. [00:57:52] Fantastic. [00:57:52] Santoj, [00:57:53] I completely agree with you. [00:57:54] I want to say again, [00:57:55] thank you. [00:57:56] We are so, so happy [00:57:57] with the amazing partnership [00:57:59] that we have, [00:58:00] you know, [00:58:00] across the board [00:58:01] and most importantly, [00:58:02] you know, [00:58:02] with your engineering teams. [00:58:03] And we truly are excited [00:58:04] about what we're going [00:58:05] to do in the future. [00:58:06] Thank you so much. [00:58:07] Thanks a lot. [00:58:07] Cheers. [00:58:07] Thank you. [00:58:08] Thank you. [00:58:10] So, you heard Santoj [00:58:15] talk about just [00:58:16] the diversity of workloads [00:58:17] and really the massive scale [00:58:20] that you require with AI. [00:58:22] Now, I want to turn [00:58:23] to another part [00:58:24] of the inference market [00:58:25] where the requirements [00:58:26] are actually very different. [00:58:27] So, lots and lots [00:58:29] of workloads [00:58:29] and as inference is moving [00:58:31] into more products [00:58:32] and services, [00:58:33] we're actually seeing [00:58:34] it segment [00:58:34] into different workloads [00:58:35] depending on what [00:58:36] you're trying to do. [00:58:37] Some of these are, [00:58:39] let's call it less [00:58:39] interactive, [00:58:40] so they're really tuned [00:58:41] for, let's call it [00:58:42] maximum throughput [00:58:43] or the lowest cost. [00:58:45] You know, [00:58:45] other of these applications [00:58:46] require more balanced throughput, [00:58:48] so you have to balance [00:58:49] throughput and responsiveness. [00:58:52] And then there's [00:58:52] this new class [00:58:53] of applications [00:58:54] that actually want [00:58:55] very, very fast results. [00:58:57] You have extremely [00:58:58] smart engineers [00:58:59] and they don't like to wait. [00:59:01] And that is where [00:59:02] ultra-low latency [00:59:03] or every millisecond [00:59:04] actually matters. [00:59:06] Each one of these [00:59:07] types of workloads [00:59:08] requires a different [00:59:09] type of compute. [00:59:11] And one of the best ways [00:59:12] to reach ultra-low latency [00:59:14] today is disaggregated inference. [00:59:17] So, if you think about, [00:59:18] you know, [00:59:18] the different pieces, [00:59:19] both pre-fill [00:59:20] and decode [00:59:21] are two different jobs. [00:59:23] So, pre-fill [00:59:24] tends to need more compute. [00:59:26] Decode tends to need [00:59:27] more memory bandwidth. [00:59:29] And with Instinct, [00:59:30] we have a very balanced machine [00:59:31] so that we deliver [00:59:32] both great compute [00:59:33] and great memory bandwidth. [00:59:35] But you can actually [00:59:36] take this a step further. [00:59:37] If you know what workloads [00:59:38] you're trying to run, [00:59:40] you can actually let customers [00:59:41] tune each of these pieces [00:59:43] independently. [00:59:44] And to address [00:59:45] this market opportunity, [00:59:47] we've been working [00:59:48] with Cerebrus. [00:59:49] So, to talk more [00:59:50] about what we're [00:59:50] building together, [00:59:51] please welcome to the stage [00:59:53] Cerebrus co-founder [00:59:54] and CEO, [00:59:55] Andrew Feldman. [00:59:56] Andrew, it's great [01:00:06] to have you here. [01:00:06] It's been an exciting [01:00:07] few months for you. [01:00:09] I know that [01:00:09] for those who don't know [01:00:11] exactly what you've [01:00:13] been working on, [01:00:13] can you talk a little bit [01:00:14] about Cerebrus [01:00:15] and what problem [01:00:16] you've been trying to solve? [01:00:17] Sure. [01:00:18] Great to be here [01:00:19] and great to be [01:00:20] sort of among people [01:00:22] who love hardware. [01:00:24] It's nice. [01:00:25] Look, I'm through [01:00:26] and thrilled to be [01:00:26] here with you [01:00:27] and to announce [01:00:29] our cool new partnership. [01:00:32] At Cerebrus, [01:00:33] we build the world's [01:00:35] largest and fastest chip. [01:00:37] It's a full wafer. [01:00:39] And we package this [01:00:40] into a system [01:00:41] and we package [01:00:42] two systems, right, [01:00:44] into racks [01:00:45] and then racks [01:00:46] into clusters. [01:00:48] And we deliver them [01:00:49] both on-premise [01:00:50] and via the cloud. [01:00:52] And you've had [01:00:53] some of our customers [01:00:54] up here already today [01:00:55] on the frontier labs. [01:00:57] We serve customers [01:00:58] like OpenAI [01:00:59] and in the hyperscalers [01:01:01] like AWS [01:01:02] and the small agentic [01:01:04] and the coding space, [01:01:06] leaders like Cognition. [01:01:08] And they do this [01:01:10] because we're blisteringly fast. [01:01:14] There's no question, [01:01:15] there's no question, [01:01:15] there's no question, [01:01:15] Andrew, [01:01:16] that you have [01:01:16] some tremendous innovation [01:01:17] with what you've been working on. [01:01:20] And our teams [01:01:21] have been collaborating [01:01:21] really closely [01:01:22] over the last few years [01:01:24] to really have Helios [01:01:26] together with the wafer scale engine [01:01:27] for ultra-low latency inference. [01:01:29] Can you just kind of just [01:01:31] educate the audience [01:01:32] a little bit? [01:01:32] What are we trying to do [01:01:33] and what does that mean [01:01:34] for customers? [01:01:35] Sure. [01:01:35] I think what's happened [01:01:37] and you described it previously [01:01:39] is that AI has moved [01:01:41] from being a novelty [01:01:42] to being useful [01:01:44] and then in some domains [01:01:45] a necessity. [01:01:47] And when something's [01:01:48] a necessity, [01:01:49] people want to use it [01:01:50] and they want to use it quickly. [01:01:52] And to serve this market, [01:01:54] this segment [01:01:55] of ultra-low latency, [01:01:57] we sought a partnership [01:01:58] that could extend [01:02:00] our capabilities [01:02:02] and our footprint. [01:02:04] And there was no better answer [01:02:05] than AMD. [01:02:06] We were already [01:02:07] an AMD customer [01:02:08] as we use AMD CPUs [01:02:10] to surround our systems. [01:02:13] And we were so excited [01:02:14] at the opportunity [01:02:15] to build a disaggregated solution [01:02:17] that combined AMD CPUs, [01:02:21] the cool new Helios rack [01:02:23] and the Cerebrus [01:02:25] wafer scale engine. [01:02:28] Yeah, look, [01:02:29] the technology [01:02:30] is pretty cool. [01:02:31] So just talk a little bit [01:02:32] about how it comes together. [01:02:33] Sure. [01:02:34] So as Lisa described, [01:02:36] what you have [01:02:37] with Instinct [01:02:38] and the Helios rack [01:02:39] is you have the leader [01:02:41] in performance [01:02:42] and memory capacity. [01:02:45] And you marry that [01:02:46] with our wafer scale engine, [01:02:48] which is the leader [01:02:49] in sort of in SRAM [01:02:51] and in memory bandwidth. [01:02:53] And that combination [01:02:55] allows us to deliver [01:02:56] a solution [01:02:57] that is unmatched [01:02:58] in the industry. [01:02:59] So very, very exciting. [01:03:00] I know that customers [01:03:01] are excited [01:03:02] about what we're going [01:03:02] to do together as well. [01:03:03] So when are we going [01:03:05] to have it in market? [01:03:07] Later this year. [01:03:09] And it will be first [01:03:10] in the Cerebrus cloud. [01:03:14] and it will be later [01:03:16] at a store near you. [01:03:18] I think the key thing [01:03:21] is we're going to give [01:03:22] customers a choice [01:03:23] to really put together [01:03:25] what is the solution [01:03:26] that they want. [01:03:27] And I think we're very, [01:03:28] very happy to be able [01:03:29] to do this together with you. [01:03:30] I think it's a huge step up [01:03:32] in terms of what we can do [01:03:33] for this ultra low latency segment. [01:03:35] I think that's right. [01:03:36] I think until recently [01:03:37] customers sort of, [01:03:39] they could have [01:03:40] sort of high throughput [01:03:42] or they could have [01:03:44] extraordinary speed. [01:03:45] And by bringing together [01:03:47] the Helios rack [01:03:49] with the Cerebrus [01:03:50] wafer scale engine, [01:03:51] we give you five times [01:03:53] the throughput [01:03:54] while continuing [01:03:56] to deliver [01:03:57] this extraordinary speed. [01:03:58] it's really something amazing. [01:04:02] Fantastic. [01:04:02] Well, look, Andrew, [01:04:03] huge congratulations [01:04:04] to what you [01:04:05] and the Cerebrus team [01:04:06] have done. [01:04:07] I mean, I think [01:04:07] it's been incredible. [01:04:08] I think you've been clear [01:04:10] on what you're trying [01:04:11] to accomplish. [01:04:12] And it's really nice [01:04:13] to see not only it [01:04:14] come together, [01:04:15] but us come together [01:04:16] to offer a solution [01:04:17] that will be very, [01:04:19] very compelling [01:04:19] to customers. [01:04:20] So we can't wait [01:04:20] to get this [01:04:21] into customers' hands. [01:04:22] Soon. [01:04:23] Thank you. [01:04:24] Good to see you. [01:04:24] Thank you so much. [01:04:25] Thank you, Andrew. [01:04:28] All right. [01:04:31] So look, [01:04:32] I think we've shown you [01:04:33] a lot of hardware [01:04:34] this morning. [01:04:35] But as we all know, [01:04:37] hardware is only part [01:04:38] of the story. [01:04:38] It's actually software [01:04:40] that turns all of this [01:04:41] into a platform [01:04:41] that developers [01:04:43] can really build on. [01:04:44] So to share the progress [01:04:45] we're making with Rockham, [01:04:47] please welcome [01:04:47] AMD's Senior Vice President [01:04:49] of AI, [01:04:50] Vamsi Bopana, [01:04:50] to the stage. [01:04:58] Thank you, Risa. [01:04:59] Good morning, everyone. [01:05:01] Morning. [01:05:02] It's great to be back [01:05:03] to talk about software. [01:05:05] Just a few years ago, [01:05:06] programming AMD GPUs [01:05:08] required deep expertise [01:05:09] and significant [01:05:10] engineering effort. [01:05:12] We've come a long, long way [01:05:13] in a remarkably short [01:05:15] period of time. [01:05:16] Through our open source [01:05:18] community collaboration [01:05:19] and sustained investment [01:05:21] in Rockham, [01:05:21] our software stack, [01:05:23] AMD Instinct GPUs [01:05:24] are powering some [01:05:25] of the most important [01:05:26] and consequential [01:05:28] workloads on the planet. [01:05:30] But the biggest change [01:05:32] is still ahead of us. [01:05:34] AI is transforming [01:05:35] how software is getting built [01:05:37] and we are bringing [01:05:38] that transformation [01:05:39] to Rockham. [01:05:41] And that's what I'm excited [01:05:42] to share with you guys today. [01:05:44] Our software teams [01:05:45] have been moving fast [01:05:46] with relentless focus [01:05:47] on developers. [01:05:49] We've invested [01:05:50] at every level of the stack [01:05:52] in build and test infrastructure [01:05:53] so we ship faster. [01:05:56] Rockham releases now go out. [01:05:57] But every six weeks, [01:05:59] not every four months, [01:06:01] we've continued [01:06:02] to expand the work [01:06:03] we do with our AI ecosystem [01:06:04] partners. [01:06:05] And the result? [01:06:07] More features, [01:06:09] faster performance, [01:06:10] and a significantly better [01:06:11] out-of-the-box experience. [01:06:14] Our strategy [01:06:15] that has gotten us here [01:06:16] has been remarkably consistent. [01:06:19] It's built [01:06:19] on two core principles. [01:06:21] Partner deeply [01:06:22] with the open source ecosystem [01:06:24] and build [01:06:26] with the right layers [01:06:27] of abstraction [01:06:27] to enable developer productivity. [01:06:30] Open source gives us velocity [01:06:32] and scale. [01:06:34] And abstraction [01:06:34] makes developers [01:06:36] more productive. [01:06:39] Our deep, deep commitment [01:06:40] to open source [01:06:41] has resonated [01:06:42] with the community [01:06:43] that has truly embraced AMD. [01:06:45] From frameworks [01:06:46] and compiler stacks [01:06:48] to inference engines [01:06:49] and model hubs, [01:06:50] AMD is now becoming [01:06:51] part of default enablement [01:06:53] for the most important [01:06:54] AI communities, [01:06:56] such as Hugging Face, [01:06:57] PyTorch, [01:06:59] Jax, [01:07:00] VLM, [01:07:01] and SGLang. [01:07:02] This is why new models [01:07:03] get day zero support [01:07:04] and why the world's [01:07:06] most important AI workloads [01:07:07] run on AMD today. [01:07:09] That is a big shift. [01:07:12] Inspired by engineers [01:07:13] seeking productivity, [01:07:16] abstractions have evolved [01:07:17] to enable work [01:07:18] at the right level of detail, [01:07:20] hiding complexity [01:07:21] when they want speed [01:07:22] and exposing control [01:07:24] when they need performance. [01:07:26] From low-level programming [01:07:27] in assembly and C [01:07:29] to block programming abstractions [01:07:30] like Triton [01:07:31] to Pythonic frameworks, [01:07:33] we've invested [01:07:34] in key abstractions. [01:07:36] Ater, [01:07:37] our kernel library, [01:07:38] and Atom, [01:07:39] our serving engine, [01:07:41] get you peak performance [01:07:42] without writing [01:07:43] the kernels yourself. [01:07:45] Fly DSL [01:07:46] is a brand new [01:07:48] Pythonic domain-specific language [01:07:49] that gives you [01:07:50] low-level control [01:07:51] with the performance [01:07:52] of hand-tuned C++. [01:07:55] And Mori, [01:07:56] our communications library, [01:07:57] helps deliver [01:07:58] leadership performance [01:08:00] on the latest models [01:08:01] like minimax. [01:08:04] But look, [01:08:05] something very, [01:08:06] very exciting [01:08:07] is happening. [01:08:08] Over the past year, [01:08:09] I've seen something [01:08:10] remarkable inside AMD. [01:08:12] our engineers [01:08:12] are using AI models [01:08:13] to generate GPU kernels, [01:08:16] optimize code, [01:08:17] debug issues, [01:08:18] and improve performance. [01:08:20] And in some cases, [01:08:21] these AI-generated kernels [01:08:22] are shockingly good, [01:08:24] better than what we expected, [01:08:26] sometimes better than [01:08:27] the most manually-tuned versions. [01:08:30] The first time you see this, [01:08:32] you're a little bit skeptical. [01:08:34] you run more tests, [01:08:35] you try to break it, [01:08:36] you look for what went wrong, [01:08:38] and then you realize [01:08:39] this is real. [01:08:41] What has happened [01:08:42] in general software development [01:08:44] is coming [01:08:45] to significantly transform [01:08:47] GPU programming. [01:08:49] It's going to reduce [01:08:50] the time it takes [01:08:51] to bring up workloads, [01:08:52] make optimization automated, [01:08:55] and remove any remaining barriers [01:08:57] to broad adoption. [01:08:59] We want to put [01:09:00] that capability [01:09:01] in the hands [01:09:03] of every developer. [01:09:05] So today, [01:09:06] I am excited [01:09:07] to introduce [01:09:08] Rockham.ai. [01:09:09] It's an agentic AI platform [01:09:19] that brings the capabilities [01:09:20] of AI-assisted GPU programming [01:09:22] to developers. [01:09:23] It starts with the simple idea [01:09:25] of giving developers access [01:09:27] to the power of an AMD ecosystem [01:09:29] through the AI coding agents [01:09:31] they already use. [01:09:33] Whether you're using cursor, [01:09:35] cloud, codex, [01:09:36] Rockham.ai helps those agents [01:09:38] understand AMD platforms, [01:09:40] understand Rockham, [01:09:41] and helps you build [01:09:43] and optimize your workloads. [01:09:45] In other words, [01:09:47] we are making [01:09:47] those popular coding agents [01:09:49] into Rockham superusers. [01:09:52] You should be able [01:09:52] to describe the workload [01:09:53] you want to run, [01:09:54] the performance target [01:09:55] you want to hit, [01:09:56] and then let the agents [01:09:57] help you get there. [01:10:00] For years, [01:10:01] we have been building [01:10:01] the open software foundation. [01:10:03] Now, [01:10:04] we are adding [01:10:05] an AI-assisted layer [01:10:06] that helps developers [01:10:08] use that foundation [01:10:09] faster and more effectively. [01:10:11] So let me show you [01:10:12] what is inside Rockham.ai. [01:10:15] It's built [01:10:15] on the foundational capabilities [01:10:17] of Rockham, [01:10:18] our core software stack. [01:10:19] Runtime, [01:10:21] libraries, [01:10:22] compilers, [01:10:23] tools, [01:10:23] framework integration, [01:10:25] all of the core software [01:10:26] that makes the stack work. [01:10:28] On top of that, [01:10:29] we built a layer [01:10:30] of AI-assisted optimization [01:10:32] called Hyperloom. [01:10:35] Using Hyperloom, [01:10:36] the system can analyze [01:10:37] the workload, [01:10:38] tune configurations, [01:10:40] select and tune kernels, [01:10:41] adjust parallelism strategies, [01:10:43] and iterate [01:10:44] towards performance goals. [01:10:45] to give you an example [01:10:47] of the power of Hyperloom. [01:10:49] I was talking [01:10:50] to the lead developer [01:10:51] last week. [01:10:52] He told me [01:10:53] they just pushed through [01:10:54] a suite of 14,000 models [01:10:56] through Hyperloom, [01:10:58] optimizing them [01:10:59] and creating valuable insights. [01:11:01] This would have been [01:11:02] impossible to imagine, [01:11:04] even with a large team [01:11:06] of engineers before. [01:11:07] That's the power [01:11:08] of Hyperloom. [01:11:08] And finally, [01:11:10] we provide an AI-native interface [01:11:12] for our developers, [01:11:14] AI skills that allow coding agents [01:11:16] to understand Rokam natively. [01:11:18] I'm incredibly excited [01:11:19] by what we have built [01:11:20] with Rokam.ai. [01:11:22] It's going to profoundly transform [01:11:24] the way developers interact [01:11:25] with our platforms. [01:11:27] So let me make this concrete [01:11:28] for you with an example. [01:11:30] Imagine you are a developer [01:11:31] who wants to optimize [01:11:33] Minimax M3 [01:11:34] with VLLM [01:11:35] on MI-355s. [01:11:37] Today, [01:11:38] there are many things [01:11:39] you need to get right. [01:11:40] You need the right [01:11:41] VLLM optimizations, [01:11:42] the right model recipes, [01:11:44] the right environment variables. [01:11:46] You may need to write [01:11:47] new kernels [01:11:47] or optimize kernels. [01:11:48] And then you still need [01:11:49] to tune all of this. [01:11:52] With Rokam.ai, [01:11:53] this interaction [01:11:54] becomes much simpler. [01:11:57] So let's see it. [01:11:58] You can say, [01:11:59] optimize Minimax M3 [01:12:01] with Hyperloom. [01:12:03] Enter. [01:12:05] And behind the scenes, [01:12:07] the agent does [01:12:08] what an expert would do. [01:12:10] Pulls a known good recipe, [01:12:12] runs and profiles the workload, [01:12:13] it creates and tests [01:12:14] a bunch of kernel [01:12:15] and runtime configurations, [01:12:17] and it keeps iterating [01:12:18] towards target. [01:12:20] You can actually watch [01:12:22] on screen here, [01:12:23] Rokam.ai write a GPU kernel. [01:12:25] It sees an opportunity [01:12:27] write a more optimized [01:12:28] MOE group gem. [01:12:30] And the result, [01:12:32] Rokam.ai delivers [01:12:33] a 38% more tokens [01:12:36] per second improvement [01:12:37] on this specific example. [01:12:40] This is the experience [01:12:41] we want for developers. [01:12:43] Simple to start. [01:12:47] That's pretty cool, huh? [01:12:52] Simple, simple to start [01:12:53] and powerful to optimize. [01:12:55] We've continued [01:12:56] our relentless focus [01:12:57] and performance. [01:12:58] With every release, [01:13:00] we've delivered [01:13:00] significant performance gains. [01:13:02] On leading models [01:13:03] like DeepSeq, [01:13:04] Rokam.ai delivers [01:13:05] 3.3 times speed up [01:13:07] over Rokam 7. [01:13:09] This has been made possible [01:13:10] by innovations [01:13:11] at every level [01:13:12] of the stack. [01:13:13] Techniques, [01:13:14] such as the use [01:13:15] of block-scaled, [01:13:16] fused MOE kernels, [01:13:18] KB cache quantization, [01:13:19] KB cache quantization, [01:13:20] expert parallelism. [01:13:21] But look, [01:13:22] the bigger story [01:13:23] is how AI agents [01:13:24] actually profiled the workload, [01:13:26] proposed new kernels, [01:13:27] tested configurations, [01:13:28] and actually validated [01:13:29] all of the results. [01:13:32] That exact same focus [01:13:33] is also on delivering [01:13:35] performance gains [01:13:36] for training. [01:13:37] On the same hardware, [01:13:39] Rokam.ai improves [01:13:40] training performance [01:13:41] by an average [01:13:41] of 2.4 times. [01:13:44] Through optimized kernels, [01:13:45] such as fused flash attention, [01:13:47] more efficient checkpointing, [01:13:49] and advanced parallelism strategies. [01:13:51] Our engineers [01:13:52] have done amazing work [01:13:54] to unlock the capabilities [01:13:55] of MI455 [01:13:56] through Rokam.ai. [01:13:58] So I'm delighted [01:13:59] to share some incredible results [01:14:01] on hardware. [01:14:02] On running real workloads, [01:14:04] we are able to see [01:14:05] the system deliver [01:14:06] 20 terabytes [01:14:08] of memory bandwidth [01:14:09] and 20 petaflops [01:14:11] of delivered compute. [01:14:13] These are the highest [01:14:14] demonstrated [01:14:15] compute plus memory [01:14:16] capabilities [01:14:17] of any accelerated platform [01:14:19] in the industry today. [01:14:21] It's not theoretical, [01:14:23] there's measured. [01:14:30] Super exciting [01:14:31] that we were able [01:14:31] to share this on day one. [01:14:33] I'm also thrilled [01:14:34] to share [01:14:34] that the leading [01:14:35] AI ecosystem partners [01:14:37] who have had their hands [01:14:38] on MI455 [01:14:39] have had a delightful experience. [01:14:41] You can see here [01:14:42] what they've had to say [01:14:43] from end-to-end [01:14:45] PyTorch testing [01:14:46] to ensuring [01:14:47] hugging face models [01:14:48] are getting validated [01:14:49] to VLLM [01:14:50] and SGLang running [01:14:51] well on MI455. [01:14:53] We are incredibly encouraged [01:14:55] and excited [01:14:56] by what the community [01:14:57] is about to unlock [01:14:58] with MI455. [01:14:59] This matters [01:15:00] because day zero readiness [01:15:02] is one of the most [01:15:03] important objectives [01:15:04] we strive for [01:15:05] and that's what [01:15:06] rockm.ai enables. [01:15:09] So now let's look [01:15:10] at rockm.ai in action [01:15:11] on Helios. [01:15:13] This time [01:15:13] we are going to use [01:15:14] a codex front-end. [01:15:16] Let's deploy [01:15:16] DeepSeq V4 Pro [01:15:18] a 1.6 trillion parameter [01:15:20] frontier model [01:15:21] on Helios. [01:15:23] Now watch what happens. [01:15:26] rockm.ai pulls [01:15:27] the right recipes, [01:15:29] it configures [01:15:30] the runtime, [01:15:30] it implements [01:15:32] the model graph [01:15:32] onto hardware [01:15:33] and finally [01:15:35] serves up the model. [01:15:38] There you go. [01:15:39] The model is actually [01:15:40] up and running. [01:15:42] And let's see the model [01:15:43] write a simple poem now [01:15:44] about advancing AI. [01:15:47] Let's hit enter. [01:15:50] Isn't that a pretty cool poem? [01:15:57] This is the experience [01:15:58] we want to give our users. [01:15:59] The developers [01:16:00] should not have [01:16:01] to manually reason [01:16:02] through every layer [01:16:03] of model architecture, [01:16:04] kernel selection, [01:16:06] comms strategy, [01:16:07] and rack-level deployment. [01:16:09] AI is transforming [01:16:10] this experience [01:16:11] and that is a very, [01:16:13] very big shift [01:16:13] for GPU hardware [01:16:14] and software. [01:16:17] Now, [01:16:17] there are few organizations [01:16:19] shaping the frontier of AI [01:16:20] as much as open AI. [01:16:22] You heard Lisa [01:16:23] and Sachin talk [01:16:24] about our partnership. [01:16:25] It's been truly wonderful [01:16:26] collaborating closely [01:16:28] with the open AI [01:16:29] technical team [01:16:30] to make rapid progress [01:16:31] with Rockham. [01:16:32] To talk about how [01:16:34] they are advancing [01:16:35] the frontier of software [01:16:36] development [01:16:36] and our collaboration, [01:16:37] I'm delighted to welcome [01:16:39] to stage [01:16:40] Philippe Tillet [01:16:41] from OpenAI. [01:16:50] Thank you for joining us [01:16:52] here, Philippe. [01:16:53] So just so you all know, [01:16:55] Philippe is the creator [01:16:56] of Triton, [01:16:57] one of the most important [01:16:58] software innovations [01:16:59] in modern AI infrastructure. [01:17:01] It's a huge privilege [01:17:01] having you here. [01:17:03] Thank you. [01:17:07] OpenAI has always pushed [01:17:09] the limits of infrastructure. [01:17:11] So, you know, [01:17:12] can you tell the team [01:17:12] a little bit here today [01:17:14] as to how your team [01:17:15] actually develops [01:17:15] GPU software [01:17:16] and how have technologies [01:17:17] like Triton evolved? [01:17:19] Yeah, for sure. [01:17:20] Well, first of all, [01:17:21] thanks so much [01:17:21] for having me. [01:17:22] I was here two years ago. [01:17:23] It's my great pleasure [01:17:24] to be back now. [01:17:27] A way of looking at things [01:17:28] is there's no one way [01:17:30] to fit all applications [01:17:32] and we really try [01:17:33] to meet users [01:17:34] where they are. [01:17:35] So for that reason, [01:17:36] we've developed [01:17:36] multiple solutions. [01:17:38] So for workflows [01:17:40] where speed of iterations [01:17:41] is more important, [01:17:42] we've developed Triton, [01:17:44] as you mentioned. [01:17:45] So this is typically [01:17:45] what researchers will use. [01:17:48] But more and more, [01:17:49] we've been in a case [01:17:50] where we've had to [01:17:51] hyper-optimize performance [01:17:52] and Triton didn't give us [01:17:53] the level of control [01:17:54] that we wanted. [01:17:55] So we've developed [01:17:56] Gluon for that, [01:17:57] which is a lower-level language. [01:18:00] And this way, [01:18:01] we've been really able [01:18:02] to optimize our kernels [01:18:03] for AMD, [01:18:04] among other things. [01:18:06] And, yeah, [01:18:08] like our agents [01:18:09] have become extremely good [01:18:10] at both of them [01:18:11] and we've worked with AMD [01:18:13] on both of them as well. [01:18:16] Yeah, and, you know, [01:18:17] more and more [01:18:17] what we're seeing [01:18:18] is agents [01:18:19] slowly taking over, [01:18:21] like, yeah, [01:18:23] they're getting just very good, [01:18:24] as you just mentioned, [01:18:25] at writing this kind of code. [01:18:26] Amazing. [01:18:27] So, look, [01:18:28] I think the work [01:18:28] you have done with Triton [01:18:29] and now Gluon, right, [01:18:30] it's been really, really [01:18:31] moving the ball forward [01:18:33] for the entire industry. [01:18:34] Now, we began our work [01:18:35] together with 300 [01:18:36] and, you know, [01:18:37] extended at 350, [01:18:38] but now we're collaborating [01:18:40] at a different scale [01:18:41] with 450 and Helios. [01:18:42] So I would love [01:18:44] to hear your thoughts [01:18:45] on how that collaboration [01:18:46] has been going [01:18:46] and how has that experience been? [01:18:48] Yeah, super well. [01:18:49] It feels like ages. [01:18:51] I think we started years ago [01:18:53] on hardware co-design, [01:18:55] you know, [01:18:55] when MI450 [01:18:56] was, like, [01:18:57] in the very early [01:18:58] prototyping stages. [01:19:00] I think we're able [01:19:01] to work really well together [01:19:02] and it's so nice [01:19:03] to see the outcome [01:19:04] of that [01:19:04] and the chip. [01:19:07] And more recently, [01:19:08] we've been working [01:19:09] very, very deeply [01:19:10] on software enablement [01:19:12] for MI450, [01:19:13] both for our models [01:19:14] but also [01:19:15] for the broader ecosystem. [01:19:16] So a lot of this development [01:19:17] has been done [01:19:18] open source. [01:19:20] And, yeah, [01:19:21] we're collaborating [01:19:21] on the full stack now [01:19:23] down to the very low-level [01:19:25] LVM co-generation, [01:19:27] you know, [01:19:28] so that, like, [01:19:28] all the nice [01:19:30] hardware advances [01:19:31] that came with MI450 [01:19:32] can be properly leveraged, [01:19:34] you know, [01:19:34] in end-to-end applications. [01:19:36] Yeah. [01:19:37] And, you know, [01:19:38] you got your Helios racks, [01:19:40] you know, [01:19:40] so tell us a little bit [01:19:40] about that, yeah. [01:19:41] Yeah, it worked. [01:19:42] Like, [01:19:43] we got the chip [01:19:44] and... [01:19:45] Yeah, [01:19:50] we got the chip [01:19:51] and, you know, [01:19:52] very quickly, [01:19:53] like, [01:19:53] a single engineer [01:19:54] in a few days [01:19:55] was able to confirm [01:19:56] that, you know, [01:19:57] everything worked. [01:19:58] Obviously, like, [01:19:58] we have to optimize [01:19:59] performance more. [01:20:01] But we're very confident [01:20:02] we can get in a very, [01:20:03] very good place. [01:20:04] Awesome. [01:20:05] That's awesome to hear. [01:20:06] So, look, [01:20:07] one of the more amazing things [01:20:08] that we've done recently, [01:20:09] right, [01:20:10] is our teams [01:20:10] have collaborated [01:20:11] even more closely [01:20:12] in terms of, [01:20:13] you know, [01:20:14] using AI [01:20:15] to program AMD GPUs. [01:20:17] So, you know, [01:20:17] would love to hear [01:20:18] your thoughts [01:20:19] and anything you can share [01:20:20] with the audience. [01:20:21] Yeah, no, [01:20:22] I mean, [01:20:23] everyone in this room [01:20:24] knows how much [01:20:25] agents are taking over [01:20:26] our jobs, [01:20:29] in a sense, [01:20:29] as kernel engineers. [01:20:31] And, [01:20:33] yeah, [01:20:34] and it's been amazing [01:20:35] just to see, like, [01:20:36] how much better [01:20:37] they've got over [01:20:38] just the past [01:20:39] six months. [01:20:40] And just now [01:20:40] they're extremely good [01:20:41] at generating [01:20:42] just high-quality [01:20:43] GPU kernels [01:20:44] in a way [01:20:44] that were not before. [01:20:46] And it's been [01:20:47] very great collaborating [01:20:48] with AMD [01:20:50] on just, like, [01:20:52] making sure [01:20:52] that agents [01:20:53] not only are capable [01:20:54] enough, [01:20:55] you know, [01:20:56] but also have [01:20:57] all the right context [01:20:58] they need [01:20:58] to be able [01:20:59] to make good decisions [01:21:00] and good optimizations. [01:21:02] That's great to hear. [01:21:03] Yeah, I think, [01:21:03] you know, [01:21:04] the quality of code [01:21:05] that's getting produced [01:21:06] by AI-assisted kernel generation [01:21:09] has been truly remarkable [01:21:10] through our collaboration. [01:21:12] Very excited. [01:21:13] Now, [01:21:13] as you look ahead, [01:21:15] right, [01:21:15] you know, [01:21:16] what crystal ball [01:21:17] do you have for us [01:21:18] for the future [01:21:19] of AI software ecosystem [01:21:20] and collaborations [01:21:21] like ours? [01:21:22] Yeah, [01:21:23] so there's two things [01:21:24] that really come to mind. [01:21:26] One, [01:21:26] you know, [01:21:27] as we keep talking, [01:21:28] is just how good [01:21:29] agents are getting. [01:21:31] And the other one is, [01:21:32] I think, [01:21:33] how much the open approach [01:21:35] to software of AMD [01:21:36] is enabling, [01:21:37] not only for OpenAI, [01:21:39] but I think [01:21:39] for the industry [01:21:40] as a whole. [01:21:41] You know, [01:21:42] I talked a little bit earlier [01:21:43] how we've been collaborating [01:21:44] on LVM code generation. [01:21:46] LVM is an open source project. [01:21:48] The entire compiler stack [01:21:49] from AMD is open source. [01:21:51] And that has allowed us [01:21:52] to make our agents [01:21:54] really good at very, [01:21:55] very low level code generation [01:21:56] down to, you know, [01:21:58] scheduling instructions [01:21:59] in assembly. [01:22:00] And this has led to, like, [01:22:02] very, very significant [01:22:04] performance gain [01:22:05] that I don't think [01:22:05] we have been able [01:22:06] to achieve [01:22:08] in a fully closed source stack. [01:22:10] Yeah, [01:22:10] that's one of the real [01:22:12] sort of big value propositions [01:22:14] that we have to offer, right? [01:22:15] Because everything that we do [01:22:17] is in the open [01:22:17] and the fact that AI [01:22:19] can actually pick that up [01:22:20] is a big one. [01:22:21] Thank you, Philippe. [01:22:22] We truly appreciate [01:22:23] the partnership [01:22:24] and we're really excited [01:22:25] about what we're going [01:22:26] to build together. [01:22:27] Yeah, my pleasure. [01:22:27] Thank you so much, Raheem. [01:22:34] What you just heard [01:22:35] is extremely important. [01:22:37] When we combine [01:22:38] better abstractions [01:22:39] with AI-assisted development [01:22:41] and closed hardware-software [01:22:43] co-design, [01:22:43] we can create [01:22:44] some incredibly [01:22:45] powerful capabilities. [01:22:46] And that's exactly [01:22:47] aligned with our strategy. [01:22:50] Now, everything [01:22:51] I've shown you so far [01:22:52] has been about AI. [01:22:53] But Raheem is also [01:22:55] a great stack [01:22:56] for developing [01:22:56] high-performance [01:22:57] and scientific computing [01:22:58] applications. [01:22:59] From frontier models [01:23:00] to climate simulation, [01:23:02] we use the same Raheem. [01:23:04] The same library, [01:23:04] same tools, [01:23:05] same AI-assisted development. [01:23:07] One open stack [01:23:08] for every single workload. [01:23:11] And in fact, [01:23:11] when you look at our roadmap, [01:23:13] we have retained [01:23:13] that very strong focus [01:23:15] on HPC [01:23:15] and scientific computing. [01:23:17] Just like MI-455 [01:23:19] leads in AI compute [01:23:20] with FP4 [01:23:21] and memory performance, [01:23:22] its MI-430X [01:23:24] leads HPC compute [01:23:25] with FP64 [01:23:27] and memory performance, [01:23:29] two platforms [01:23:29] purpose-built [01:23:30] for two very different [01:23:31] classes of workload. [01:23:34] The MI-430X [01:23:35] is built [01:23:36] by leveraging [01:23:37] our modular [01:23:37] chiplet architecture [01:23:38] to deliver [01:23:39] a purpose-built variant [01:23:40] with native [01:23:42] FP64 hardware, [01:23:44] giving customers [01:23:45] leadership performance [01:23:46] across AI and HPC [01:23:48] in a single platform. [01:23:50] And this is [01:23:51] hardware double precision. [01:23:52] not emulation. [01:23:54] For the scientific community, [01:23:56] that distinction [01:23:56] is everything. [01:23:58] This product delivers [01:23:59] 288 teraflops [01:24:01] of FP64 compute, [01:24:03] nearly nine times [01:24:04] better than competition. [01:24:06] And the same leadership, [01:24:07] memory capacity, [01:24:08] and bandwidth [01:24:09] of MI-455. [01:24:11] The MI-430X [01:24:12] ships [01:24:13] in the first half [01:24:15] of 2027. [01:24:16] moment. [01:24:16] This is why [01:24:19] leaders in sovereign AI [01:24:21] love this product. [01:24:23] The first sovereign [01:24:23] exascale AI factories [01:24:25] in both US and Europe [01:24:26] are being built [01:24:27] on the MI-430X. [01:24:29] Discovery [01:24:30] at Oak Ridge National Labs [01:24:31] and Alistair Koch [01:24:33] with CNES [01:24:33] and Gen-C [01:24:34] in Europe. [01:24:36] So look, [01:24:37] as we look ahead, [01:24:39] let me leave you [01:24:39] with this thought. [01:24:41] Every so often, [01:24:42] our industry goes [01:24:43] through inflection points. [01:24:45] We've all felt it [01:24:46] when neural networks [01:24:47] first started recognizing [01:24:49] cats and dogs, [01:24:50] when transformers changed [01:24:51] what was possible [01:24:52] with AI. [01:24:53] I believe we are experiencing [01:24:55] another one [01:24:55] right now, [01:24:57] a real inflection [01:24:58] in the ability [01:24:59] of AI [01:25:00] to program [01:25:00] complex hardware systems. [01:25:02] Building on the surface area [01:25:04] that's been exposed [01:25:05] from open source software [01:25:07] and leveraging the power [01:25:08] of abstracted interfaces, [01:25:10] AI is making it [01:25:11] dramatically easier [01:25:13] to program AI systems. [01:25:15] We cannot wait [01:25:16] what you will build with it. [01:25:19] Thank you. [01:25:28] Now, AI is getting [01:25:29] pervasively infused [01:25:31] into all forms of computing. [01:25:32] Enterprises are experiencing [01:25:34] the same shift [01:25:35] with their own [01:25:35] unique requirements. [01:25:37] So to tell us more [01:25:38] about that transformation, [01:25:40] please welcome to stage [01:25:41] Senior Vice President [01:25:42] and General Manager [01:25:43] of Compute [01:25:44] and Enterprise AI, [01:25:46] Dan McNamara. [01:25:55] Thank you, Vansi, [01:25:56] and good morning, everyone. [01:25:58] It is a great pleasure [01:26:00] for me to bring to you [01:26:02] the third pillar [01:26:02] of our strategy, [01:26:03] which is powering [01:26:04] AI everywhere. [01:26:06] And, you know, [01:26:06] the vision behind this [01:26:08] has remained the same [01:26:09] for several years [01:26:10] for us, [01:26:11] you know, [01:26:12] to deliver the right [01:26:13] compute engine [01:26:14] to the right workload [01:26:15] and then continuously [01:26:17] drive optimization points [01:26:19] across our entire portfolio [01:26:22] from CPUs, [01:26:24] GPUs, [01:26:25] networking, [01:26:25] and FPGAs. [01:26:27] So building on [01:26:29] what we've already shared, [01:26:30] I'll cover how our strategy [01:26:32] comes together [01:26:33] across our enterprise portfolio [01:26:34] and how the foundation [01:26:36] we built [01:26:37] with the CPU franchise [01:26:38] has truly positioned us [01:26:40] for the next era of AI [01:26:42] across multiple deployment models. [01:26:46] So Lisa talked about [01:26:47] the frontier model [01:26:48] and hyperscale landscape [01:26:49] and just the rapid [01:26:51] and massive advances [01:26:52] we've seen [01:26:53] over the last [01:26:54] six to 12 months. [01:26:56] And those advances [01:26:57] are driving [01:26:57] the adoption of AI [01:26:59] well beyond the cloud. [01:27:01] And as AI continues to grow, [01:27:03] it will demand compute [01:27:04] beyond the massive [01:27:05] purpose-built data centers [01:27:07] that we all know and love, [01:27:09] extending into enterprise, [01:27:10] personal, [01:27:11] and physical AI. [01:27:12] And each of these [01:27:13] brings a new [01:27:14] or unique set [01:27:16] of requirements. [01:27:17] Enterprises are usually [01:27:19] constrained by power and cooling. [01:27:20] Personal AI usually [01:27:22] must operate [01:27:23] within a fairly limited [01:27:25] compute and memory footprint. [01:27:26] And physical AI [01:27:28] must perform reliably [01:27:29] in real-world demanding environments. [01:27:33] And AMD [01:27:34] is the only company [01:27:35] delivering a complete portfolio [01:27:37] that spans [01:27:37] all of these [01:27:38] development models. [01:27:41] So I just mentioned [01:27:42] that our journey began [01:27:44] in the data center [01:27:46] with our CPU franchise [01:27:47] and I want to touch [01:27:48] on it a bit. [01:27:48] With each generation [01:27:50] of Epic, [01:27:50] our strategy has been [01:27:51] to listen to our customers, [01:27:53] address their evolving needs, [01:27:54] and deliver [01:27:55] the highest performance, [01:27:57] lowest total cost [01:27:58] of ownership, [01:27:59] and the fastest time [01:28:00] to value. [01:28:02] As Lisa mentioned earlier, [01:28:03] we've introduced [01:28:04] many industry firsts [01:28:06] across our generations [01:28:07] that our enterprise [01:28:09] customers [01:28:09] has actually asked for. [01:28:12] And that customer focus [01:28:13] has brought us [01:28:14] to where we are today [01:28:15] with Epic [01:28:16] as the leader [01:28:18] in enterprise computing. [01:28:20] And over the last three years, [01:28:21] there's been [01:28:21] very, very strong adoption [01:28:23] across the enterprise [01:28:24] with the world's largest companies [01:28:25] moving more and more [01:28:26] workloads to AMD. [01:28:28] And it's reflected [01:28:29] in a number of places. [01:28:31] First, public cloud, [01:28:32] where our VM consumption [01:28:34] has grown at 76% CAG [01:28:36] over the last three years. [01:28:38] And on-prem deployments [01:28:40] with our OEM [01:28:41] and ODM partners [01:28:42] is growing quite aggressively [01:28:44] across key industries. [01:28:46] But just as important, [01:28:49] this position allows us [01:28:51] to understand truly [01:28:52] the diversity [01:28:52] of the workloads [01:28:53] that our enterprise [01:28:54] customers run every day. [01:28:57] And each workload [01:28:59] places new demands [01:29:02] on the CPU. [01:29:03] Some require [01:29:03] maximum thread count, [01:29:05] some depend on [01:29:06] per-core performance, [01:29:07] while others depend [01:29:08] on memory band [01:29:09] with cache efficiency [01:29:10] or I.O. performance. [01:29:12] The really important point [01:29:14] is that there is [01:29:15] no single SKU [01:29:17] that solves [01:29:18] every enterprise application. [01:29:20] And that's why [01:29:21] our EPIC portfolio [01:29:22] with VENIS [01:29:23] spans 8 to 256 cores [01:29:26] with a range of power, [01:29:27] frequencies, [01:29:28] and I.O. capabilities, [01:29:29] giving customers [01:29:31] the flexibility [01:29:31] to choose [01:29:32] the right solution [01:29:33] for their workload [01:29:34] and most importantly, [01:29:36] at the most optimized cost. [01:29:39] So, Lisa mentioned [01:29:41] this earlier [01:29:42] and now we all know [01:29:43] we covered VENIS [01:29:45] and she showed [01:29:46] how we have [01:29:48] strong leadership [01:29:49] across AI host nodes [01:29:51] and the agentic workflows. [01:29:53] But I wanted to show you [01:29:54] how VENIS [01:29:55] actually leads [01:29:56] across the general purpose [01:29:58] workloads [01:29:58] of the enterprise [01:29:59] and these workloads [01:30:01] that are running, [01:30:01] our customers [01:30:03] rely on every day. [01:30:04] As you can see, [01:30:05] we have a minimum [01:30:06] of 2.5x [01:30:08] to performance [01:30:09] across our competition, [01:30:10] across these very, [01:30:11] very important [01:30:12] enterprise workloads. [01:30:15] So, now, [01:30:16] as we all know, [01:30:17] agentic AI [01:30:18] is really [01:30:20] the next major transition [01:30:21] for the enterprise [01:30:22] and it truly [01:30:23] is a force multiplier [01:30:25] for enterprises [01:30:26] and adoption [01:30:27] is accelerating. [01:30:28] In fact, [01:30:29] the latest data here [01:30:29] shows [01:30:30] is that nearly [01:30:31] every enterprise [01:30:31] next year [01:30:32] will have some form [01:30:34] of agentic deployments. [01:30:37] And as enterprises [01:30:38] move from, [01:30:39] you know, [01:30:40] pilots to production [01:30:41] deployments, [01:30:42] we've talked to CIOs [01:30:44] all the time [01:30:44] and they are looking [01:30:46] at three major things. [01:30:47] First, [01:30:48] predictable infrastructure [01:30:50] cost and token costs. [01:30:52] Secondly, [01:30:53] making sure [01:30:54] the data is secure [01:30:55] and then lastly, [01:30:57] and sometimes [01:30:58] this is overlooked, [01:30:59] right, [01:31:00] integrating AI [01:31:01] into their existing [01:31:02] infrastructure. [01:31:03] It's not all [01:31:04] greenfield out there [01:31:05] for the enterprise [01:31:06] customer. [01:31:08] So, [01:31:08] to solve [01:31:09] these problems, [01:31:10] we truly believe [01:31:11] that this is going [01:31:11] to be a distributed [01:31:12] deployment model [01:31:13] that leverages [01:31:14] frontier models, [01:31:16] cloud services, [01:31:17] on-prem infrastructure, [01:31:17] and of course, [01:31:18] AI clients. [01:31:20] And as AI [01:31:22] has moved [01:31:22] from chatbots [01:31:23] responding to prompts [01:31:24] to, you know, [01:31:25] agents doing real work, [01:31:26] writing code, [01:31:28] summarizing documents, [01:31:29] and running complex flows, [01:31:32] the requirements [01:31:33] have evolved also. [01:31:35] And those are exactly [01:31:36] the requirements [01:31:37] we've been designing [01:31:38] for for multiple years. [01:31:40] So, [01:31:41] while AI will be deployed [01:31:43] across all these [01:31:44] environments, [01:31:45] the on-prem deployments [01:31:46] in our enterprise [01:31:46] customers have [01:31:47] fairly unique constraints. [01:31:49] First and foremost, [01:31:50] power cooling. [01:31:52] Really big challenge [01:31:53] in space. [01:31:54] And of course, [01:31:56] costs remain [01:31:57] top of the list [01:31:58] for CIOs [01:31:59] that are trying to, [01:32:00] you know, [01:32:01] drive things forward [01:32:02] in the AI world. [01:32:03] And then lastly, [01:32:05] the full solutions [01:32:06] they need [01:32:07] to be easy [01:32:09] to deploy. [01:32:10] That's been common [01:32:12] for many, many years [01:32:13] for the enterprise. [01:32:14] But it has to, [01:32:15] that's something [01:32:16] we're very, very focused on. [01:32:17] So, [01:32:18] to address these constraints, [01:32:20] thank you, Laura. [01:32:22] I'm very happy [01:32:24] to announce [01:32:26] the launch [01:32:26] of the Instinct [01:32:27] MI350P. [01:32:34] So, today, [01:32:35] every enterprise [01:32:36] wants AI [01:32:37] in the data center. [01:32:38] But the challenge is, [01:32:39] how do you add [01:32:40] those capabilities [01:32:41] without starting over? [01:32:43] That's exactly [01:32:44] what the 350P [01:32:45] was designed to solve. [01:32:46] It's an air-cooled GPU [01:32:48] designed to fit [01:32:49] within the power [01:32:50] and cooling [01:32:50] of enterprise servers [01:32:52] today [01:32:52] without a facilities upgrade, [01:32:55] making it easy [01:32:56] to bring LLM scale [01:32:57] inference [01:32:58] into today's [01:32:59] enterprise data center. [01:33:00] And at the same time, [01:33:01] there's no compromise [01:33:02] in capability. [01:33:04] A single MI350P [01:33:06] can support [01:33:06] up to 260 billion parameters, [01:33:09] allowing customers [01:33:09] to run [01:33:10] the majority [01:33:11] of enterprise AI workloads [01:33:12] on a single GPU. [01:33:15] And the economics [01:33:16] are actually even better [01:33:18] and more compelling. [01:33:19] The 350P delivers [01:33:21] more than four times [01:33:22] the tokens per second [01:33:23] per dollar [01:33:24] than the competition. [01:33:26] It's pretty cool. [01:33:28] And it actually [01:33:32] turns the customer's [01:33:35] existing data center [01:33:36] into an AI data center. [01:33:38] Now, [01:33:39] we all know [01:33:41] that benchmarks [01:33:41] are one part [01:33:42] of the equation, right? [01:33:44] The other part [01:33:44] is really evaluating [01:33:46] does the advantage [01:33:47] hold up [01:33:48] across the actual workload [01:33:49] you run [01:33:49] and how your business runs. [01:33:51] So we tested [01:33:52] a set of workloads, [01:33:54] enterprise use cases [01:33:55] using production models. [01:33:57] And the 350P [01:33:58] delivers much higher [01:33:59] productivity [01:34:00] than the competition, [01:34:02] delivering up to [01:34:02] two to five times [01:34:04] tokens per, [01:34:05] faster tokens per second [01:34:07] than the competition. [01:34:08] So if you think [01:34:09] about that, [01:34:10] that's more tokens [01:34:11] per second, [01:34:12] more work completed, [01:34:14] more users supported, [01:34:16] all at a lower cost [01:34:18] for the business. [01:34:20] So another area [01:34:22] that I wanted [01:34:23] to talk about [01:34:23] is our earliest [01:34:24] customers, [01:34:25] our own AMD IT team. [01:34:28] And they're [01:34:30] a tough group [01:34:31] to work with [01:34:31] even for us. [01:34:32] So, [01:34:33] but they put us [01:34:35] through our paces [01:34:36] and, you know, [01:34:37] really this has been [01:34:38] our model [01:34:39] from day one [01:34:39] is we want [01:34:40] to put our technology [01:34:41] into our own [01:34:42] data center first. [01:34:43] So like many companies, [01:34:45] many enterprises, [01:34:46] we're looking at [01:34:46] how to deliver AI service [01:34:47] to our employee base [01:34:48] but also manage [01:34:50] costs [01:34:50] and protect our data [01:34:52] and manage [01:34:53] and control the data. [01:34:54] So we focused [01:34:56] on two use cases [01:34:57] that we wanted [01:34:58] to try out [01:34:59] and first was [01:35:00] autonomous threat detection [01:35:02] and, you know, [01:35:02] that is probably [01:35:04] one of the faster [01:35:04] growing agentic [01:35:05] applications [01:35:06] that we're seeing today. [01:35:07] And the second was, [01:35:09] you know, [01:35:09] a personalized AI assistant [01:35:10] running on OpenClaw. [01:35:12] So both of these [01:35:13] run route requests [01:35:15] through the router [01:35:16] and the gateway [01:35:18] sort of tests [01:35:20] the complexity, [01:35:21] the sensitivity [01:35:22] of the information [01:35:22] and the service level [01:35:24] requirements. [01:35:25] So some requests [01:35:26] are sent [01:35:27] to the Frontier model [01:35:28] while there's a router [01:35:29] to open weights [01:35:30] models running [01:35:31] on Epic [01:35:31] and an MI350P. [01:35:34] So with intelligent routing, [01:35:36] we reduced our token [01:35:37] cost by 43% [01:35:39] while delivering [01:35:41] up to 3x [01:35:42] faster response times [01:35:43] for the workloads [01:35:44] running locally. [01:35:46] So we believe [01:35:47] this is how [01:35:48] enterprise AI [01:35:49] will be deployed. [01:35:52] But we also know [01:35:54] that from our history [01:35:56] here with Epic [01:35:56] that enabling enterprise [01:35:58] takes a whole lot more [01:36:00] than great silicon. [01:36:02] It requires a complete [01:36:03] ecosystem of software [01:36:04] and solutions. [01:36:05] And over the last [01:36:06] several years, [01:36:06] we've truly grown [01:36:08] significantly [01:36:09] and expanded [01:36:09] our enterprise [01:36:10] AI ecosystem. [01:36:11] Today, [01:36:12] we support [01:36:12] numerous AI ISVs, [01:36:15] provide zero-day support [01:36:16] for all of the industry's [01:36:18] leading open models, [01:36:19] and deliver robust [01:36:20] frameworks [01:36:21] that help customers [01:36:22] bring AI [01:36:23] into production [01:36:24] as fast as possible. [01:36:26] And it's all built [01:36:27] on trusted platform [01:36:29] software and OEM [01:36:30] partnerships [01:36:30] that enterprises [01:36:32] rely on [01:36:32] every single day. [01:36:34] Now, [01:36:35] everything I just shared [01:36:38] comes down [01:36:39] comes down [01:36:39] to one goal [01:36:39] for us, [01:36:40] and that's [01:36:41] helping customers [01:36:42] deploy AI [01:36:42] where it creates [01:36:44] the most value [01:36:45] for them, [01:36:46] whether it's [01:36:47] in the cloud, [01:36:48] on-premise, [01:36:48] or at the edge. [01:36:50] So to bring this [01:36:51] to life, [01:36:52] I'd like to welcome [01:36:53] Jeremy Legg, [01:36:55] Chief Technology Officer [01:36:56] at AT&T. [01:37:04] Hey, crew, [01:37:05] good to see you. [01:37:06] Yeah, thanks for having me. [01:37:08] Yeah, this is great. [01:37:09] It's kind of a homecoming. [01:37:10] I grew up in the East Bay, [01:37:12] so it's good to be back. [01:37:13] Yeah, fantastic. [01:37:13] We really appreciate [01:37:14] you joining us today. [01:37:15] So, look, [01:37:16] I think everyone knows [01:37:18] AT&T is the largest [01:37:20] communications company, [01:37:21] one of the largest [01:37:22] in the world, [01:37:23] and I'm sure [01:37:23] all these people, [01:37:25] half of them out there [01:37:26] probably have AT&T [01:37:27] network right now, [01:37:29] but, you know, [01:37:29] hopefully more than half. [01:37:32] But, you know, [01:37:34] AI is touching [01:37:35] way more than the network, [01:37:36] and you and your team [01:37:37] are really doing [01:37:38] a fair amount of work [01:37:39] driving it. [01:37:41] So can you share [01:37:41] with us sort of [01:37:42] the opportunities [01:37:42] you approach, [01:37:44] you're working on, [01:37:44] and your experience to date? [01:37:46] Sure, happy to. [01:37:48] At AT&T, [01:37:49] we're now burning [01:37:50] about a trillion-plus [01:37:52] tokens per month, [01:37:54] and that number [01:37:55] has gone up [01:37:56] pretty dramatically. [01:37:57] It's moving up [01:37:58] double digits, [01:37:59] and so this is now [01:38:00] becoming very pervasive [01:38:02] inside of the enterprise. [01:38:04] We've got over [01:38:05] 100 Gen A models [01:38:06] in production [01:38:07] across our enterprise [01:38:10] as well, [01:38:11] and those are stretching [01:38:12] from, you know, [01:38:13] things that I think [01:38:14] a lot of folks [01:38:14] are doing in customer care, [01:38:16] things that we're doing [01:38:17] in fraud, [01:38:18] but also things [01:38:19] that you wouldn't [01:38:19] necessarily think about, [01:38:21] like where's the best place [01:38:22] to place a cell tower [01:38:24] or a RAN [01:38:25] that is the most effective [01:38:26] use of it [01:38:27] inside of a network. [01:38:29] So all of these things [01:38:30] are happening [01:38:30] across our enterprise, [01:38:32] but they're happening [01:38:32] at a scale [01:38:33] that is not necessarily normal. [01:38:36] We get over 300,000 calls a day. [01:38:39] The transcription [01:38:40] of those calls [01:38:41] is an enormous workload, [01:38:42] and then driving [01:38:43] the insights [01:38:44] out of those calls [01:38:45] that we then bring back [01:38:46] to the business [01:38:46] is another set of insights. [01:38:49] So it's really [01:38:50] an interesting thing, [01:38:51] and now we're beginning [01:38:52] to actually rebuild [01:38:54] entire workflows [01:38:55] inside of the company [01:38:56] beyond just the point [01:38:58] use of a use case. [01:39:00] So as we think about HR, [01:39:01] as we think about finance, [01:39:03] and we think about [01:39:03] different kinds of things [01:39:05] that each of those [01:39:05] organizations do, [01:39:07] cash forecasting, [01:39:09] things that we do [01:39:09] on onboarding employees [01:39:10] as examples, [01:39:11] we're now agentifying [01:39:13] those entire workloads. [01:39:14] Yeah, so it is, [01:39:17] you know, [01:39:17] we work closely [01:39:18] with your team, [01:39:19] and we've had [01:39:19] many conversations. [01:39:21] It's truly amazing [01:39:22] the number of actual, [01:39:23] you know, [01:39:25] opportunities [01:39:26] you're addressing [01:39:27] with AI. [01:39:28] So can, you know, [01:39:28] a super broad range of use. [01:39:30] So can you just explain [01:39:32] sort of how, [01:39:32] what are the key learnings [01:39:34] in moving these [01:39:36] to AI [01:39:37] at enterprise scale, [01:39:38] and then maybe a bit [01:39:39] on what we've done together? [01:39:41] Sure. [01:39:41] I mean, [01:39:42] I'll start at a place [01:39:43] that I don't think [01:39:44] everybody starts from here, [01:39:46] which is the human. [01:39:48] You need world-class teams [01:39:50] in order to do this, [01:39:51] and you need people [01:39:52] that think in workflows. [01:39:54] There's an enormous amount [01:39:55] of what I call [01:39:56] racing from stoplight [01:39:57] to stoplight. [01:39:59] Folks are going [01:40:00] from zero to 100, [01:40:01] and then they hit [01:40:01] the next stage [01:40:02] of the workflow, [01:40:03] and then there's [01:40:03] a red light, [01:40:04] and they haven't worked [01:40:04] their way all the way [01:40:05] through that. [01:40:07] So I would start [01:40:08] with people [01:40:09] and the quality [01:40:09] of the people [01:40:10] that you really have [01:40:13] across your enterprise [01:40:14] in order to do [01:40:14] these things. [01:40:15] The second point [01:40:16] is we are big believers [01:40:18] in data sovereignty, [01:40:20] and data sovereignty [01:40:21] for us means [01:40:22] not being tied [01:40:23] to a specific chipset, [01:40:25] not being tied [01:40:26] to a specific model [01:40:27] or a specific set [01:40:28] of development tools. [01:40:30] You need, [01:40:30] as an enterprise, [01:40:31] to manage your data. [01:40:32] Your data is your fuel [01:40:34] in terms of how [01:40:35] you implement AI [01:40:36] across that broader enterprise. [01:40:38] What we've done [01:40:39] with AMD, [01:40:40] and AMD's been [01:40:41] such a great partner [01:40:42] in this, [01:40:43] is leveraging [01:40:44] open source models, [01:40:45] leveraging different sets [01:40:46] of chipsets, [01:40:47] building our own models, [01:40:49] training other models. [01:40:50] We're one of the few [01:40:51] companies in the world [01:40:52] that's actually post-trained [01:40:53] a model on AMD [01:40:55] and then been able [01:40:56] to drive equivalency [01:40:57] in terms of performance [01:40:58] and accuracy [01:40:59] with other more expensive [01:41:00] chipsets [01:41:01] in the marketplace. [01:41:03] So as we have done that, [01:41:04] we've been able [01:41:05] to drive down token costs [01:41:07] essentially through moneyballing [01:41:08] across all of these [01:41:10] different models [01:41:10] and leveraging [01:41:11] that human capital [01:41:12] in a way [01:41:13] that's made a big difference [01:41:14] to our company. [01:41:16] And so increasingly, [01:41:17] even though token consumption [01:41:18] is going up, [01:41:19] we're able to manage [01:41:20] that underlying token cost [01:41:22] at enterprise scale, [01:41:23] which is not a small thing. [01:41:26] Yeah, it's just, [01:41:26] it's a bunch of great work. [01:41:28] Another key area [01:41:29] that's sort of groundbreaking, [01:41:31] which is exciting also, [01:41:32] is you're the first [01:41:33] telecom company [01:41:34] to actually train [01:41:35] a model on AMD [01:41:36] and also you're driving [01:41:38] fully open source models [01:41:40] in telecoms. [01:41:40] So tell us about [01:41:41] that experience [01:41:41] and why you felt [01:41:42] like that was needed. [01:41:44] We felt like the industry [01:41:45] was moving down a path [01:41:47] where folks were only [01:41:49] going to use [01:41:49] closed source models [01:41:50] to solve solutions. [01:41:52] And while closed source models [01:41:53] are an important ingredient [01:41:54] in that, [01:41:55] we didn't want to lose sight [01:41:56] of the fact that [01:41:56] open source models [01:41:57] were going to be [01:41:58] a very significant part [01:41:59] of the solution. [01:42:00] So we announced [01:42:01] what we call [01:42:02] Otel 1.0, [01:42:04] that's the open telco AI model, [01:42:06] a bit over a year ago [01:42:07] or so now, [01:42:08] and we had over 18 million [01:42:10] downloads of that model [01:42:12] since then. [01:42:13] We're excited [01:42:14] to announce today [01:42:15] the launch of Otel 2.0, [01:42:17] which is a more advanced [01:42:18] set of models [01:42:19] that we have trained [01:42:20] on AMD, [01:42:21] and we're now making [01:42:23] that available [01:42:23] via open source [01:42:24] across the broader industry. [01:42:27] And it, I think, [01:42:27] is going to give folks [01:42:28] an opportunity [01:42:29] to see [01:42:29] that not only [01:42:31] do you need [01:42:31] to think about [01:42:32] leveraging models, [01:42:33] but you also need [01:42:34] to think about [01:42:34] training those models [01:42:35] and how they work [01:42:36] across your enterprise. [01:42:38] So that's just great. [01:42:39] We certainly appreciate [01:42:40] the partnership, [01:42:41] and, you know, [01:42:41] we really feel like [01:42:43] this is germane [01:42:44] to the conversation. [01:42:44] It's exactly what [01:42:46] we've been talking [01:42:46] about today, [01:42:47] and I just want [01:42:48] to thank you again [01:42:48] for joining us [01:42:49] and the partnership. [01:42:50] Thank you. [01:42:51] You're a big part of it. [01:42:52] Thank you. [01:42:57] Okay, so let me close [01:42:59] with one last point. [01:43:01] We've established [01:43:02] ourselves as a leader [01:43:03] in enterprise computing [01:43:05] by listening to our customers [01:43:06] and delivering [01:43:07] the right solutions [01:43:08] for their workloads. [01:43:10] Today, we're extending [01:43:11] that leadership [01:43:12] to help customers [01:43:14] accelerate their AI journey. [01:43:17] This is our vision [01:43:18] for powering AI everywhere. [01:43:21] And to expand a little bit [01:43:23] beyond the data center [01:43:24] into physical [01:43:26] and client AI, [01:43:29] it's my pleasure [01:43:30] to introduce [01:43:30] Mr. Jack Yoon, [01:43:32] who leads our computing [01:43:33] and graphics group. [01:43:35] Thank you. [01:43:40] Thank you, Dan. [01:43:42] It's great to see [01:43:44] so many friends, [01:43:45] partners, [01:43:45] and developers [01:43:46] here in San Francisco. [01:43:47] Today, we have seen [01:43:49] how Andy is building [01:43:50] the compute foundation [01:43:51] for AI in a data center. [01:43:54] But the full AI transformation [01:43:56] requires two more ingredients. [01:43:58] Intelligence that becomes personal [01:43:59] and intelligence that can [01:44:01] understand and act [01:44:02] in the physical world. [01:44:04] Those are the two frontiers [01:44:05] I want to explore [01:44:06] with you here today. [01:44:07] Let's begin [01:44:08] with personal AI. [01:44:10] Enterprise AI [01:44:11] is already reshaping [01:44:12] the data center, [01:44:14] but agentic AI [01:44:15] will extend far beyond it. [01:44:17] Agents are the next [01:44:18] great productivity multiplier. [01:44:20] The jump from software [01:44:22] that responds [01:44:23] to software [01:44:24] that completes work. [01:44:25] A personal agent [01:44:26] can double [01:44:27] what one person can do. [01:44:29] A team of agents [01:44:29] can multiply that by 10. [01:44:31] And as agents [01:44:32] gain more capabilities, [01:44:34] often can grow exponentially. [01:44:36] But that promise [01:44:37] has a compute consequence. [01:44:39] Unlike a single AI query, [01:44:41] as agent reasons, [01:44:43] calls tools, [01:44:45] in accordance to other agents [01:44:46] and run nonstop, [01:44:50] agents become more capable [01:44:51] and pervasive. [01:44:52] So will token consumption. [01:44:55] Meaning that demand [01:44:56] will require every layer [01:44:57] of the computing ecosystem [01:44:58] to scale together. [01:45:00] Data centers remain [01:45:01] the foundation [01:45:02] for frontier AI [01:45:03] in the most demanding workloads. [01:45:05] But edge-inclined systems [01:45:07] can extend that capacity, [01:45:09] putting more intelligence [01:45:10] closer to where data [01:45:12] is created [01:45:12] and where the work gets done. [01:45:16] The PC is already [01:45:17] one of the most powerful [01:45:18] compute engines [01:45:19] ever placed [01:45:20] in the hands [01:45:21] of an individual. [01:45:22] Across billions of devices, [01:45:24] enormous computing capability [01:45:26] is already deployed [01:45:27] every single day. [01:45:28] Yet its potential [01:45:29] remains largely untapped. [01:45:32] Putting that capacity [01:45:33] to work [01:45:33] changed the economics [01:45:34] of computing. [01:45:36] More workloads [01:45:36] can run on systems [01:45:38] that are already in place, [01:45:39] reducing overall infrastructure costs [01:45:42] and allowing data centers [01:45:43] to focus on the largest, [01:45:45] most demanding task. [01:45:46] And local computing [01:45:47] brings inherent advantages. [01:45:50] Your data stays with you. [01:45:52] Your applications [01:45:52] remain responsive [01:45:53] and your system [01:45:55] keeps working [01:45:56] even when your network [01:45:58] does not. [01:45:59] And thanks to [01:45:59] a new generation [01:46:01] of smaller, [01:46:02] more capable AI models [01:46:04] will begin to unlock [01:46:05] that potential. [01:46:07] And the unlock [01:46:08] is happening faster [01:46:10] than anyone expected. [01:46:12] Last August, [01:46:13] GPT OSS [01:46:15] needed 120 billion parameters [01:46:17] to score 80 [01:46:18] on GPQA. [01:46:20] Just seven months later, [01:46:22] QN 3.5 scored even higher [01:46:25] with only 9 billion. [01:46:27] Better performance [01:46:28] with 13 times fewer parameters. [01:46:32] This is not [01:46:32] an incremental improvement. [01:46:34] It's a dramatic shift [01:46:35] in the compute required [01:46:37] to deliver [01:46:38] advanced intelligence. [01:46:40] And this is a trend [01:46:41] now an outlier. [01:46:43] Smaller open models [01:46:45] are closing [01:46:45] the frontier gap [01:46:46] at extraordinary speed. [01:46:49] QN 3.6 [01:46:50] 27 billion parameter miles [01:46:52] now outperforms [01:46:54] leading frontier miles [01:46:56] from the previous generation [01:46:57] on graduate level reasoning [01:46:59] at a fraction [01:47:00] of the size and cost. [01:47:02] That changes [01:47:03] is what is possible. [01:47:06] Workloads once reserved [01:47:08] for the cloud [01:47:09] can now run [01:47:10] on path compute engines [01:47:11] designed for personal AI compute. [01:47:15] Let's map [01:47:16] these model performance [01:47:18] breakthroughs [01:47:18] to any platforms. [01:47:20] Ryzen AI 400 [01:47:22] supports miles [01:47:24] up to 24 billion parameters. [01:47:27] More than enough [01:47:28] for QN 3.5 [01:47:29] 9 billion [01:47:30] with substantial headroom. [01:47:33] Ryzen AI Max [01:47:34] extends that capacity [01:47:36] to most as large [01:47:37] as 200 billion parameters [01:47:40] running natively [01:47:41] on a personal system. [01:47:43] This is not reduced AI [01:47:44] compressed to fit [01:47:46] on a PC. [01:47:47] It's a new class [01:47:48] of personal computing [01:47:49] built to run [01:47:51] powerful models [01:47:52] where the work [01:47:52] actually happens. [01:47:54] and this is exactly [01:47:56] what led us [01:47:57] to build [01:47:57] Ryzen AI Halo. [01:47:59] We wanted to give developers [01:48:01] a platform [01:48:02] powerful enough [01:48:03] to run these miles locally [01:48:04] yet simple enough [01:48:06] to use every day. [01:48:08] With 120 gigabytes [01:48:09] of unified memory [01:48:10] and support [01:48:12] for miles [01:48:12] up to 200 billion parameters [01:48:14] developers can build, [01:48:17] test, [01:48:18] and iterate locally [01:48:19] directly on their desk. [01:48:22] Ryzen AI Halo [01:48:24] embodies the future [01:48:25] of personal AI [01:48:26] for everyone. [01:48:27] A future [01:48:28] that shortens [01:48:28] the distance [01:48:29] between an idea [01:48:30] and a working model. [01:48:33] For developers, [01:48:34] that means [01:48:34] a dramatically [01:48:35] more efficient [01:48:36] and optimized path [01:48:37] from experimentation [01:48:38] to deployment. [01:48:41] And as Vanti [01:48:41] showed earlier, [01:48:43] rockam.ai [01:48:44] is making [01:48:44] hardware programming [01:48:45] radically more accessible. [01:48:48] Developers can understand, [01:48:51] modify, [01:48:52] and optimize code [01:48:53] for rockam [01:48:54] even without [01:48:55] deep rockam expertise. [01:48:58] And better yet, [01:48:59] the same open software [01:49:00] foundation extends [01:49:02] across all of AMD's hardware. [01:49:05] Developers can use [01:49:06] the frameworks [01:49:06] and tools [01:49:07] they already know, [01:49:08] start locally [01:49:09] on Ryzen AI Halo [01:49:10] and scale across [01:49:12] the full AMD compute platform [01:49:14] without rebuilding their work. [01:49:16] And we're not stopping there. [01:49:18] Today, [01:49:19] we're expanding our partnership [01:49:20] with Hug and Face [01:49:21] to bring the open AI ecosystem [01:49:24] directly on Ryzen AI Halo. [01:49:28] Power from compute [01:49:29] only creates value [01:49:30] when developers [01:49:31] can quickly access [01:49:33] the right models, [01:49:35] optimize them for the platform, [01:49:36] and turn them [01:49:37] into real applications. [01:49:39] Together, [01:49:41] AMD and Hug and Face [01:49:42] will deliver [01:49:43] Halo-optimized performance [01:49:45] for open models [01:49:46] and genetic workflows. [01:49:51] How cool is that? [01:49:54] And along with [01:49:55] co-engineered libraries [01:49:56] and toolkits [01:49:57] designed to help developers [01:49:58] spend less time [01:49:59] configuring infrastructure [01:50:00] and more time building. [01:50:03] And later this year, [01:50:05] every Ryzen AI Halo box [01:50:07] will include a full year [01:50:09] of Hug and Face Pro, [01:50:11] putting the power [01:50:12] of partnership [01:50:13] in developers' hands [01:50:14] from day one. [01:50:16] Freedom [01:50:16] and more power [01:50:18] to the user. [01:50:20] This is how [01:50:21] we accelerate innovation [01:50:22] in the AI era. [01:50:25] And we're already [01:50:25] pushing that vision [01:50:27] further [01:50:28] with Gorgon Halo. [01:50:30] And seeing it's believing, [01:50:32] look at how small [01:50:34] and beautiful [01:50:35] disk boxes. [01:50:41] We've increased [01:50:42] unified memory [01:50:43] from 128 [01:50:45] to 192 gigabytes, [01:50:48] the largest [01:50:49] unified memory pool [01:50:50] in this class, [01:50:51] and increased [01:50:52] mouse support [01:50:53] from 200 billion [01:50:54] to 300 billion [01:50:55] parameters. [01:50:57] Thank you, Laura. [01:51:01] And this is [01:51:02] no longer limited [01:51:03] to a single platform [01:51:04] form factor. [01:51:06] Together [01:51:06] with the world's [01:51:07] leading OEMs, [01:51:09] AMD now powers [01:51:10] the broadest portfolio [01:51:11] of personal AI systems [01:51:12] from laptops [01:51:13] and workstations [01:51:14] to many PCs [01:51:16] and developer platforms. [01:51:17] That means developers, [01:51:19] enterprises, [01:51:21] and creators [01:51:21] can choose [01:51:22] the right system [01:51:23] for the way [01:51:23] they want to build [01:51:24] and work. [01:51:26] Personal AI [01:51:27] is not a concept, [01:51:28] it is a category, [01:51:30] and is available [01:51:30] right now on AMD. [01:51:33] We now have [01:51:34] all the pieces. [01:51:36] Models are efficient enough. [01:51:38] Hardware is powerful enough. [01:51:40] Software stack [01:51:41] is open and ready. [01:51:43] It opened harnesses [01:51:44] like Lemonade Server [01:51:45] can route workloads [01:51:46] to the right model [01:51:47] and right compute. [01:51:48] The technology is here. [01:51:50] The next challenge [01:51:51] is taking it [01:51:52] from one developer's desk [01:51:53] to thousands of users [01:51:55] across an enterprise. [01:51:56] Cisco sees and believes [01:51:58] in that future [01:51:58] the same way we do. [01:52:00] To take us deeper [01:52:01] into how we turn [01:52:02] that vision [01:52:02] into enterprise-scale reality, [01:52:05] please join me [01:52:05] in welcoming to the stage [01:52:07] Jitu Patel, [01:52:08] President and Chief Product [01:52:09] Officer at Cisco. [01:52:14] Jitu, how are you doing? [01:52:15] How are you? [01:52:16] You look great. [01:52:17] You too. [01:52:17] Thank you for joining me here today. [01:52:18] Congratulations. [01:52:19] Thank you. [01:52:20] Thank you. [01:52:21] Now, Jitu, [01:52:21] you and I talked [01:52:22] how we're at inflection point. [01:52:24] And for the past few years, [01:52:26] Energy has spent time [01:52:27] just building intelligence. [01:52:29] The next phase [01:52:30] is how we deploy [01:52:31] this intelligence [01:52:32] at scale enterprises. [01:52:33] You spent so much [01:52:34] of your time [01:52:35] talking to CIOs, [01:52:36] technology leaders. [01:52:37] What are they telling you? [01:52:38] What's changing? [01:52:39] Well, firstly, [01:52:40] it's a really exciting time [01:52:41] to be alive in tech. [01:52:42] And if you just take [01:52:43] a step back right now [01:52:44] and see what's happening, [01:52:45] there's actually [01:52:46] a fundamental shift [01:52:47] that's happening [01:52:48] in the patterns of inferencing. [01:52:49] So if you think [01:52:50] about the patterns [01:52:51] of inferencing [01:52:51] that existed [01:52:52] during the chatbot era [01:52:53] that was human-led, [01:52:55] it was very, very spiky. [01:52:56] You ask a question, [01:52:57] you get an answer. [01:52:58] What you're starting [01:52:58] to see with agents [01:52:59] is they're working [01:53:00] seven by 24. [01:53:02] They're very consumptive [01:53:04] on bandwidth [01:53:04] and on infrastructure. [01:53:05] So you're starting [01:53:06] to see a much more [01:53:07] persistent pattern [01:53:08] of kind of demand [01:53:09] signal for infrastructure. [01:53:10] You know, [01:53:11] we usually joke [01:53:13] around internally, [01:53:14] humans click, [01:53:14] but agents swarm. [01:53:15] Yes. [01:53:16] You know? [01:53:17] And so as you start [01:53:18] to see this, [01:53:19] what you're going to have [01:53:20] is a tremendous amount [01:53:21] of, [01:53:23] and there's two big concerns [01:53:24] that are also there, [01:53:25] which is one is [01:53:26] every customer [01:53:26] is concerned about security [01:53:28] and every customer [01:53:30] is concerned about cost. [01:53:31] So what you're starting [01:53:32] to see happen [01:53:33] is inferencing [01:53:35] is not just going to be [01:53:36] limited to the data center. [01:53:38] Inferencing is going [01:53:38] to be distributed everywhere [01:53:39] and there's a whole new [01:53:40] class of computing [01:53:41] that's emerging [01:53:42] with this desk-side computing [01:53:44] that you just announced [01:53:46] where there's going [01:53:47] to be inferencing [01:53:47] that you're going to have [01:53:48] a laptop or a desktop [01:53:49] where humans work [01:53:51] and they might have [01:53:52] a desk-side computer [01:53:53] where agents are kind [01:53:54] of going out [01:53:54] and running their jobs [01:53:56] on behalf of humans. [01:53:58] So that's what [01:53:59] we're seeing essentially. [01:54:00] And we're fully aligned G2 [01:54:01] and that's exactly [01:54:02] the shift we're seeing. [01:54:03] You know, [01:54:03] if AI is moving closer [01:54:05] to employees, [01:54:06] compute has to move [01:54:06] closer as well, right? [01:54:08] The PC is not only [01:54:09] just a productivity device [01:54:10] because the intelligence node [01:54:12] in the enterprise, [01:54:14] but deploying intelligence [01:54:15] everywhere creates [01:54:16] a new channel for CIOs [01:54:17] that you and I [01:54:18] talk about so much. [01:54:19] We do, yeah. [01:54:19] How do you operate [01:54:20] thousands of these systems [01:54:22] with security, governance, [01:54:24] and control [01:54:24] enterprises require? [01:54:26] G2, how is Amy and Cisco [01:54:28] solving this [01:54:29] and what are we building together? [01:54:31] Yeah, so this is [01:54:32] an exciting time [01:54:33] because what you're [01:54:34] starting to see is [01:54:35] there's a few things [01:54:36] that need to really be done. [01:54:38] You can't just go out [01:54:39] and deploy your [01:54:40] desk-side computers [01:54:41] and hope that everything [01:54:43] works out in the enterprise. [01:54:44] You need to make sure [01:54:45] that it's governed, [01:54:46] it's managed. [01:54:47] And so there's a few things [01:54:48] that are really important. [01:54:49] Number one, [01:54:50] you have to have [01:54:50] an appropriate amount [01:54:51] of network bandwidth [01:54:52] to satiate the needs [01:54:53] of the agents [01:54:54] that you're going to run locally. [01:54:56] So this notion [01:54:57] of network infrastructure [01:54:58] that can get modernized [01:55:00] to keep up with the needs [01:55:02] that you're going to go out [01:55:03] and generate [01:55:03] is going to be very important. [01:55:04] That's number one. [01:55:05] Number two, [01:55:06] like I said, [01:55:07] every customer is really [01:55:08] worried about token costs. [01:55:10] And so we need to make sure [01:55:11] that we can actually contain [01:55:12] the costs of an agent's [01:55:14] consumption of tokens [01:55:15] when it actually starts [01:55:16] to go awry. [01:55:16] Up and down when needed. [01:55:17] Exactly. [01:55:18] Number three then is [01:55:20] how do we monitor agent behavior [01:55:22] for safety and security [01:55:24] so that if it does start [01:55:25] to do things [01:55:26] that we don't want it to do, [01:55:27] we can in runtime [01:55:28] provide enforcement guardrails. [01:55:30] And then fourthly, [01:55:32] it's just the overall safety [01:55:33] and security [01:55:34] that you need to have [01:55:35] from an apparatus perspective [01:55:36] to get this entire thing humming [01:55:38] in the way that you want it to hum. [01:55:39] Yes, G2, [01:55:40] let's hear how you and I [01:55:40] think very alike. [01:55:41] So, I mean, [01:55:42] the opportunity is not just [01:55:44] being at every device [01:55:45] as you and I talk. [01:55:46] How do we enable [01:55:46] enterprise to deploy, [01:55:48] manage and scale out [01:55:49] with confidence? [01:55:50] Why don't we take a look [01:55:50] at what we've been building [01:55:51] in the past few weeks and months? [01:55:53] Yeah, it's exciting. [01:55:55] So basically what we've done [01:55:56] is there's a full stack [01:55:58] that you can start to think about. [01:55:59] So what you folks have done [01:56:00] with Halo is you've got [01:56:02] an isolated secure agent sandbox, [01:56:06] you've got intelligent routing, [01:56:07] you've got an MCP, [01:56:09] core set of integrations. [01:56:13] And then what we've done [01:56:14] above that is we've made sure [01:56:16] that you have the right level [01:56:17] of security policy enforcement. [01:56:20] We've also got the right level [01:56:21] of observability [01:56:22] so that it's an end-to-end [01:56:24] resilient infrastructure stack. [01:56:26] So it's observability [01:56:27] of how is your apparatus [01:56:30] and the infrastructure working? [01:56:32] Is the agent behavior [01:56:33] being done in the right way? [01:56:35] Can you observe the agent behavior? [01:56:36] And are the tokenomics [01:56:38] within the guidelines [01:56:39] that you want them to be? [01:56:41] And then what we have [01:56:42] is a single unified management plane, [01:56:46] a control plane [01:56:47] that can look at every single [01:56:49] dimension that you have [01:56:51] and be managed [01:56:52] within one environment. [01:56:54] And so that's kind of [01:56:55] the full stack. [01:56:56] And we are really, really excited [01:56:57] about the partnership. [01:56:58] You know, us too. [01:56:59] Our teams love working [01:56:59] with your team, G2. [01:57:00] Yeah. [01:57:00] How is this all managed? [01:57:02] How does the CAO [01:57:03] deploy this at scale? [01:57:04] Yeah. [01:57:04] So let's take a look [01:57:05] at what this looks like, right? [01:57:06] Because what we would have [01:57:08] is this is Cisco Cloud Control, [01:57:10] which is the overall management plane. [01:57:12] So if you happen to have Halo devices [01:57:15] on every desktop, [01:57:17] which it might actually get to, [01:57:19] what you want to make sure [01:57:21] that you do is you are able [01:57:22] to go out and manage [01:57:23] your entire estate [01:57:26] of not just your Halo fleet, [01:57:28] but also what you have running [01:57:29] on the data centers [01:57:30] as well as what you have running [01:57:32] in the cloud. [01:57:32] So you can provide full visibility [01:57:35] for that entirety of the estate. [01:57:38] And then once you've got that, [01:57:39] what we also have is this notion [01:57:41] of how do we go out [01:57:42] and monitor tokenomics? [01:57:45] And so is the agent behaving [01:57:47] the way that it needs to behave? [01:57:49] Or is the agent actually going out [01:57:51] and being overly consumptive? [01:57:53] At which point you should be able [01:57:55] to isolate any individual Halo device [01:57:58] so that you can make sure [01:57:59] that you can quarantine it. [01:58:00] No, I love this. [01:58:00] Then we actually see [01:58:01] the return of invested capital. [01:58:02] It could pay for itself [01:58:03] in three or six months [01:58:04] even quicker. [01:58:05] Exactly. [01:58:05] So what you see over here basically [01:58:07] is you can tell [01:58:08] how every single agent [01:58:10] is performing. [01:58:11] and whether or not [01:58:12] they're consuming [01:58:13] way too much [01:58:13] kind of token costs. [01:58:15] And if they are, [01:58:16] how do you actually contain [01:58:17] the cost as you're going through it? [01:58:18] No, I love this G2. [01:58:19] Now, everyone else is wondering [01:58:21] when can they buy this? [01:58:23] Yeah. [01:58:23] When is this open for everyone? [01:58:25] Well, the good news [01:58:26] is you can buy the Halo [01:58:27] already today. [01:58:29] It does. [01:58:29] And then what we're doing [01:58:31] is we're making sure [01:58:31] that we provide [01:58:32] this entire management apparatus [01:58:34] that's currently [01:58:36] in early availability [01:58:37] for a select set of customers. [01:58:39] but we'll have it [01:58:40] in general availability [01:58:41] in the U.S. [01:58:43] in early fall. [01:58:45] So we're excited [01:58:46] to make sure [01:58:46] that we can provide [01:58:47] not just the compute [01:58:48] but all of the apparatus [01:58:49] that's needed [01:58:50] to go out and manage this [01:58:51] so that you can have inference [01:58:52] running in the cloud, [01:58:54] you can have inference [01:58:55] running in your data center [01:58:56] privately [01:58:56] or you can have inference [01:58:58] running on your desk side. [01:58:59] G2, we're so grateful [01:59:00] for a partnership [01:59:01] with you and Cisco [01:59:02] and we're so excited. [01:59:03] Thank you so much. [01:59:03] What you just saw [01:59:11] brings the vision [01:59:12] and scaling [01:59:12] AI seriously and securely [01:59:14] from personal devices [01:59:15] all the way [01:59:16] to the cloud [01:59:17] but in the age of AI [01:59:19] intelligence [01:59:19] will be pervasive [01:59:20] and it will be everywhere. [01:59:22] The next frontier [01:59:23] is where AI [01:59:24] leaves the screen [01:59:25] enters the world [01:59:26] and acts. [01:59:28] That frontier [01:59:29] here is physical AI [01:59:30] the ultimate application [01:59:32] of agentic frameworks [01:59:33] in the digital world [01:59:35] an agent plans [01:59:37] and completes a task [01:59:38] in the physical world [01:59:40] it must also sense motion [01:59:41] navigate uncertainty [01:59:43] and protect people [01:59:45] here intelligence [01:59:47] does not just produce [01:59:48] an answer [01:59:49] it produces an action [01:59:50] when answers [01:59:52] become actions [01:59:53] latency becomes safety [01:59:54] and reliability [01:59:55] becomes trust [01:59:56] that is why [01:59:58] the promise [01:59:59] of physical AI [01:59:59] is so incredibly powerful [02:00:01] it is not about [02:00:02] replacing human potential [02:00:04] it is about amplifying it [02:00:06] giving surgeons [02:00:07] greater precision [02:00:08] keeping people [02:00:10] out of harm's way [02:00:11] making farms [02:00:12] more productive [02:00:13] supply chains [02:00:15] more resilient [02:00:15] and factories [02:00:17] more efficient [02:00:18] the best technology [02:00:20] does not make people [02:00:20] smaller [02:00:21] it gives them [02:00:22] much greater reach [02:00:23] for more than 20 years [02:00:25] AMD has helped [02:00:27] power the core capabilities [02:00:28] that robotics depend on [02:00:30] from sensing and functional safety [02:00:32] to deterministic control [02:00:34] and precision motion [02:00:36] that foundation [02:00:37] is already trusted [02:00:38] by many of the companies [02:00:39] to find their future robotics [02:00:41] and now [02:00:42] that foundation [02:00:44] is coming alive [02:00:45] in extraordinary ways [02:00:46] these are not machines [02:00:48] following a script [02:00:49] they are beginning [02:00:50] to perceive the world [02:00:51] reason through uncertainty [02:00:53] and act with purpose [02:00:55] the promise of physical AI [02:00:58] already in motion [02:00:59] the next era of robotics [02:01:02] is not programmed [02:01:03] it is autonomous [02:01:05] while traditional robots [02:01:07] execute instructions [02:01:08] autonomous robots [02:01:09] are given an objective [02:01:11] and from there [02:01:12] they determine how to achieve it [02:01:14] and they do this [02:01:15] all at once [02:01:16] and all in real time [02:01:18] the robot must orchestrate [02:01:20] intelligence [02:01:21] motion [02:01:22] and safety [02:01:23] continuously [02:01:24] and every part [02:01:26] must respond [02:01:27] at the speed [02:01:28] of the world around it [02:01:29] that demands [02:01:31] an entirely new [02:01:32] kind of robotics brain [02:01:33] and to solve that [02:01:35] we rethought [02:01:36] its architecture [02:01:37] from the ground up [02:01:38] today [02:01:39] I'm so excited [02:01:40] to introduce [02:01:41] the AMD Korea [02:01:42] AI system [02:01:43] on module [02:01:44] powered [02:01:45] by Ryzen AI [02:01:46] embedded [02:01:47] X100 [02:01:48] it brings CPU [02:01:55] GPU [02:01:56] NPU [02:01:58] and unified memory [02:01:59] together [02:01:59] in one compact [02:02:01] open standard [02:02:02] architecture [02:02:03] this allows [02:02:04] robots [02:02:05] to process [02:02:05] perception [02:02:06] AI reasoning [02:02:08] and do it [02:02:09] all simultaneously [02:02:10] and in real time [02:02:12] the architecture [02:02:13] is different [02:02:14] but the results [02:02:15] are decisive [02:02:16] in third party [02:02:18] testing [02:02:18] against NVIDIA [02:02:19] Jensen Thor [02:02:20] Korea delivers [02:02:22] 3.4 times [02:02:23] better real time [02:02:24] results [02:02:25] 2.3 times [02:02:27] more concurrent [02:02:27] agents [02:02:28] and 1.6 times [02:02:30] more CPU [02:02:31] capacity [02:02:31] that means [02:02:33] faster reactions [02:02:34] and greater [02:02:36] headroom [02:02:36] for the robot [02:02:37] to take on [02:02:37] more complex work [02:02:38] and all this [02:02:40] without compromising [02:02:41] control [02:02:41] or safety [02:02:42] but [02:02:44] to find [02:02:45] the next era [02:02:46] of robotics [02:02:46] takes more [02:02:47] than a powerful [02:02:48] brain [02:02:48] developers [02:02:50] need a complete [02:02:51] path [02:02:51] from idea [02:02:52] to a machine [02:02:53] operating [02:02:54] in the real [02:02:55] world [02:02:56] today [02:02:57] we are introducing [02:02:59] AMD's [02:03:00] Korea AI [02:03:01] robotics [02:03:02] developer platform [02:03:03] the industry's [02:03:10] first turnkey [02:03:11] open [02:03:12] fully [02:03:13] integrative [02:03:14] robotic system [02:03:15] thank you [02:03:17] Laura so much [02:03:18] at the center [02:03:24] is the [02:03:24] Korea AI [02:03:25] system [02:03:25] on module [02:03:26] paired with [02:03:27] a robotics [02:03:28] carrier card [02:03:29] designed [02:03:29] for sensing [02:03:31] and connectivity [02:03:31] built [02:03:32] on Rockham [02:03:33] and Ross [02:03:34] 2 [02:03:34] this platform [02:03:35] gives developers [02:03:36] everything they need [02:03:37] to move [02:03:37] from concept [02:03:38] to prototype [02:03:39] in days [02:03:40] and the same [02:03:41] module [02:03:42] carries forward [02:03:43] into production [02:03:43] with no migration [02:03:45] or redesign [02:03:46] needed [02:03:47] at the finish line [02:03:47] you can build [02:03:49] faster [02:03:49] and move [02:03:50] from imagination [02:03:51] to deployment [02:03:52] with absolute [02:03:53] confidence [02:03:53] this is the [02:03:55] breath [02:03:56] only AMD [02:03:56] can bring [02:03:57] to Tom's [02:03:57] robotics [02:03:58] from the [02:03:59] first signal [02:04:00] to the [02:04:00] final motion [02:04:01] AMD [02:04:02] Korea [02:04:03] for the [02:04:03] brain [02:04:04] Verso [02:04:05] for the [02:04:05] spine [02:04:06] Zinc [02:04:07] for the [02:04:07] joints [02:04:08] and Spartan [02:04:09] for the [02:04:10] sensors [02:04:10] each layer [02:04:11] purpose [02:04:12] built [02:04:12] for its [02:04:13] role [02:04:13] working [02:04:14] together [02:04:15] in one [02:04:15] intelligent [02:04:16] system [02:04:16] connecting [02:04:17] perception [02:04:18] and high [02:04:19] level [02:04:19] reasoning [02:04:20] to real [02:04:21] time [02:04:21] control [02:04:22] the future [02:04:24] will not [02:04:24] be defined [02:04:25] by one [02:04:25] model [02:04:25] one [02:04:27] machine [02:04:27] or one [02:04:28] company [02:04:29] it will [02:04:29] be built [02:04:30] in the [02:04:30] open [02:04:30] across [02:04:32] silicon [02:04:32] software [02:04:34] systems [02:04:35] and an [02:04:36] ecosystem [02:04:36] moving forward [02:04:37] together [02:04:38] from the [02:04:39] cloud [02:04:39] to the [02:04:39] PC [02:04:40] and now [02:04:41] into the [02:04:41] physical [02:04:42] world [02:04:42] AMD [02:04:43] continues [02:04:43] to build [02:04:44] the [02:04:44] computer [02:04:44] foundation [02:04:45] for [02:04:45] intelligence [02:04:46] everywhere [02:04:47] with [02:04:48] laser [02:04:48] focus [02:04:49] on [02:04:49] the [02:04:49] future [02:04:49] so [02:04:50] our [02:04:50] solution [02:04:50] can [02:04:50] augment [02:04:51] and [02:04:51] expand [02:04:52] what [02:04:52] people [02:04:53] are [02:04:53] capable [02:04:53] of [02:04:53] achieving [02:04:53] the [02:04:55] next [02:04:55] era [02:04:55] of [02:04:56] AI [02:04:56] will [02:04:56] not [02:04:56] just [02:04:57] unfold [02:04:57] behind [02:04:57] a [02:04:57] screen [02:04:58] it [02:04:58] will [02:04:59] unfold [02:04:59] in [02:04:59] the [02:04:59] world [02:04:59] around [02:05:00] us [02:05:00] and [02:05:01] now [02:05:01] to bring [02:05:02] it [02:05:02] all [02:05:02] together [02:05:02] and [02:05:03] take [02:05:03] us [02:05:03] to [02:05:03] what [02:05:03] comes [02:05:04] next [02:05:04] please [02:05:05] welcome [02:05:05] back [02:05:05] Lisa [02:05:06] to [02:05:06] the [02:05:06] stage [02:05:07] hi [02:05:13] Lisa [02:05:13] all right [02:05:16] thank you [02:05:16] Jack [02:05:17] and [02:05:17] it's [02:05:18] been [02:05:18] a [02:05:18] big [02:05:19] day [02:05:19] we've [02:05:19] covered [02:05:20] a lot [02:05:20] from [02:05:21] our [02:05:21] data [02:05:21] center [02:05:22] products [02:05:22] to [02:05:23] enterprise [02:05:23] AI [02:05:23] to [02:05:24] personal [02:05:24] and [02:05:24] physical [02:05:25] AI [02:05:25] and [02:05:25] the [02:05:25] software [02:05:26] that [02:05:26] ties [02:05:26] it [02:05:26] all [02:05:26] together [02:05:27] this [02:05:28] is [02:05:28] the [02:05:28] strongest [02:05:29] product [02:05:29] portfolio [02:05:30] in our [02:05:30] history [02:05:31] and [02:05:31] we [02:05:32] believe [02:05:32] it's [02:05:32] the [02:05:32] strongest [02:05:33] portfolio [02:05:33] in the [02:05:34] industry [02:05:34] but [02:05:35] it's [02:05:35] really [02:05:35] just [02:05:36] the [02:05:36] beginning [02:05:36] because [02:05:36] what [02:05:37] our [02:05:37] customers [02:05:38] count [02:05:38] on [02:05:38] is [02:05:39] not [02:05:39] just [02:05:39] today [02:05:40] they [02:05:40] really [02:05:41] count [02:05:41] on [02:05:41] us [02:05:41] looking [02:05:42] ahead [02:05:42] and [02:05:43] pushing [02:05:43] the [02:05:43] bleeding [02:05:44] edge [02:05:44] of [02:05:44] technology [02:05:44] so [02:05:45] I [02:05:45] want [02:05:45] to [02:05:50] use [02:05:51] in [02:05:51] 2028 [02:05:52] we're [02:05:53] going [02:05:53] to [02:05:53] introduce [02:05:54] Florence [02:05:54] Florence [02:05:55] brings [02:05:56] the [02:05:56] next [02:05:56] gen [02:05:57] 7 [02:05:57] 7 [02:05:57] cores [02:05:58] it's [02:05:59] leading [02:05:59] edge [02:05:59] process [02:05:59] technology [02:06:00] it's [02:06:01] a new [02:06:01] set [02:06:02] of [02:06:02] AI [02:06:02] compute [02:06:02] extensions [02:06:03] to [02:06:03] really [02:06:04] ensure [02:06:05] that [02:06:05] we [02:06:05] have [02:06:05] all [02:06:05] of [02:06:05] the [02:06:06] AI [02:06:06] capability [02:06:06] and [02:06:07] it [02:06:07] supports [02:06:07] the [02:06:07] latest [02:06:08] memory [02:06:08] technologies [02:06:09] and [02:06:09] we're [02:06:10] not [02:06:10] stopping [02:06:10] there [02:06:10] we're [02:06:11] already [02:06:11] deep [02:06:12] in [02:06:12] development [02:06:12] of [02:06:13] Ravenna [02:06:13] our [02:06:14] 8th [02:06:14] generation [02:06:15] Epic [02:06:15] family [02:06:16] built [02:06:16] on [02:06:16] Zen 8 [02:06:17] and [02:06:17] that [02:06:17] family [02:06:18] is [02:06:18] already [02:06:18] well [02:06:19] under [02:06:19] development [02:06:19] for [02:06:20] 2030 [02:06:20] now [02:06:22] to [02:06:22] give [02:06:22] you [02:06:22] a [02:06:22] little [02:06:23] bit [02:06:23] of [02:06:23] a [02:06:23] flavor [02:06:24] Florence [02:06:24] is [02:06:25] again [02:06:25] going [02:06:26] to [02:06:26] extend [02:06:26] our [02:06:27] biggest [02:06:27] advantage [02:06:27] with [02:06:28] Epic [02:06:28] it's [02:06:28] all [02:06:29] about [02:06:29] product [02:06:30] breadth [02:06:30] and [02:06:31] having [02:06:31] the [02:06:31] right [02:06:32] foundation [02:06:33] and [02:06:33] also [02:06:34] the [02:06:34] right [02:06:34] family [02:06:35] for [02:06:35] each [02:06:35] workload [02:06:36] so [02:06:36] you [02:06:36] have [02:06:36] Florence [02:06:37] you [02:06:37] have [02:06:38] Ferrara [02:06:38] you [02:06:38] have [02:06:39] Fidenza [02:06:39] and [02:06:40] it's [02:06:40] built [02:06:41] on [02:06:43] that [02:06:43] is [02:06:43] optimized [02:06:44] for [02:06:44] these [02:06:44] different [02:06:45] workloads [02:06:45] that's [02:06:46] giving [02:06:46] the [02:06:46] customers [02:06:47] the [02:06:47] right [02:06:48] CPU [02:06:48] for [02:06:49] every [02:06:49] workload [02:06:50] and [02:06:50] that [02:06:50] is [02:06:51] our [02:06:51] commitment [02:06:51] with [02:06:52] Epic [02:06:52] now [02:06:53] moving [02:06:54] to [02:06:54] instinct [02:06:54] we [02:06:55] are [02:06:55] continuing [02:06:56] our [02:06:56] cadence [02:06:56] of [02:06:57] delivering [02:06:57] a [02:06:58] new [02:06:58] generation [02:06:59] every [02:06:59] single [02:07:00] year [02:07:00] MI [02:07:01] 500 [02:07:02] brings [02:07:02] next [02:07:03] generation [02:07:03] HBM [02:07:04] it [02:07:05] brings [02:07:05] a [02:07:05] larger [02:07:06] scale [02:07:06] up [02:07:06] domain [02:07:07] because [02:07:07] scaling [02:07:08] is [02:07:08] everything [02:07:08] and [02:07:09] it [02:07:09] introduced [02:07:10] new [02:07:10] copper [02:07:11] and [02:07:11] optical [02:07:11] interconnects [02:07:12] and [02:07:13] with [02:07:13] MI [02:07:14] 600 [02:07:14] it's [02:07:15] powered [02:07:15] by [02:07:15] our [02:07:15] CDNA [02:07:16] next [02:07:16] architecture [02:07:17] and [02:07:17] it's [02:07:17] already [02:07:18] deep [02:07:18] in [02:07:18] development [02:07:18] for [02:07:19] 2028 [02:07:20] now [02:07:21] to give [02:07:21] you [02:07:22] an [02:07:22] idea [02:07:22] of [02:07:23] what [02:07:23] we [02:07:23] see [02:07:23] from [02:07:24] a [02:07:24] performance [02:07:24] standpoint [02:07:25] every [02:07:26] generation [02:07:27] of [02:07:27] instinct [02:07:27] has [02:07:28] delivered [02:07:28] a [02:07:28] major [02:07:29] big [02:07:29] step [02:07:29] up [02:07:30] in [02:07:30] performance [02:07:30] and [02:07:31] with [02:07:31] MI [02:07:31] 455 [02:07:32] we've [02:07:32] talked [02:07:33] a lot [02:07:33] about [02:07:33] it [02:07:33] delivering [02:07:34] 35 [02:07:34] times [02:07:35] more [02:07:35] inference [02:07:36] throughput [02:07:36] compared [02:07:37] to [02:07:37] MI [02:07:38] 355 [02:07:38] but [02:07:39] we're [02:07:39] really [02:07:40] really [02:07:40] excited [02:07:41] about [02:07:42] what we [02:07:42] see [02:07:43] is [02:07:43] another [02:07:44] opportunity [02:07:44] to take [02:07:45] a major [02:07:46] step [02:07:46] up [02:07:47] and [02:07:47] actually [02:07:47] bend [02:07:47] the [02:07:48] curve [02:07:48] again [02:07:48] so [02:07:49] MI [02:07:49] 500 [02:07:50] will [02:07:50] deliver [02:07:51] the [02:07:51] largest [02:07:51] generational [02:07:52] leap [02:07:52] in the [02:07:53] history [02:07:53] of [02:07:53] instinct [02:07:53] putting [02:07:54] us [02:07:54] on track [02:07:55] to [02:07:55] deliver [02:07:55] more [02:07:56] than [02:07:56] 2,000 [02:07:57] times [02:07:57] higher [02:07:57] inference [02:07:58] throughput [02:07:58] in [02:07:59] just [02:07:59] four [02:08:00] years [02:08:00] now [02:08:07] what [02:08:08] I [02:08:08] can [02:08:08] tell [02:08:08] you [02:08:08] is [02:08:09] we [02:08:09] are [02:08:09] working [02:08:09] with [02:08:10] a [02:08:10] number [02:08:10] of [02:08:10] our [02:08:10] customers [02:08:11] already [02:08:11] on [02:08:11] this [02:08:12] design [02:08:12] point [02:08:12] MI [02:08:13] 500 [02:08:13] is [02:08:14] super [02:08:14] exciting [02:08:15] and [02:08:15] the [02:08:15] feedback [02:08:16] we're [02:08:16] getting [02:08:16] from [02:08:17] customers [02:08:17] is [02:08:17] just [02:08:18] fantastic [02:08:18] now [02:08:19] put [02:08:20] the [02:08:20] CPU [02:08:21] GPU [02:08:21] and [02:08:22] the [02:08:22] networking [02:08:22] roadmaps [02:08:23] together [02:08:23] and [02:08:24] you [02:08:24] can [02:08:24] now [02:08:24] see [02:08:25] a [02:08:25] complete [02:08:25] cadence [02:08:26] of [02:08:26] our [02:08:26] rack [02:08:27] scale [02:08:27] systems [02:08:27] so [02:08:28] you [02:08:28] should [02:08:28] expect [02:08:29] from [02:08:29] AMD [02:08:29] a [02:08:30] new [02:08:30] Helios [02:08:31] system [02:08:31] every [02:08:32] year [02:08:32] delivering [02:08:33] more [02:08:33] performance [02:08:34] more [02:08:35] capacity [02:08:35] so [02:08:36] that [02:08:36] we [02:08:36] can [02:08:36] run [02:08:36] larger [02:08:37] models [02:08:37] more [02:08:38] agents [02:08:39] at [02:08:39] better [02:08:39] total [02:08:40] cost [02:08:40] of [02:08:40] ownership [02:08:40] and [02:08:41] that [02:08:41] gives [02:08:42] our [02:08:42] customers [02:08:42] the [02:08:43] complete [02:08:43] predictable [02:08:44] roadmap [02:08:44] to [02:08:45] plan [02:08:46] and [02:08:46] scale [02:08:46] their [02:08:46] infrastructure [02:08:47] for [02:08:47] years [02:08:48] to [02:08:48] come [02:08:48] so [02:08:49] lots [02:08:50] of [02:08:50] exciting [02:08:50] things [02:08:51] in [02:08:51] store [02:08:51] over [02:08:51] the [02:08:53] with [02:08:53] that [02:08:54] let [02:08:54] me [02:08:54] wrap [02:08:54] this [02:08:55] up [02:08:55] for [02:08:55] this [02:08:55] morning [02:08:56] today [02:08:57] I [02:08:57] hope [02:08:57] you [02:08:57] saw [02:08:58] a [02:08:58] little [02:08:58] bit [02:08:59] about [02:08:59] how [02:08:59] we're [02:09:00] really [02:09:00] looking [02:09:01] at [02:09:01] that [02:09:02] compute [02:09:02] from [02:09:03] every [02:09:03] angle [02:09:03] so [02:09:04] every [02:09:04] aspect [02:09:04] of [02:09:05] AI [02:09:05] compute [02:09:05] whether [02:09:06] you're [02:09:06] talking [02:09:06] about [02:09:07] Helios [02:09:08] or [02:09:08] MI455 [02:09:09] or [02:09:09] Venice [02:09:09] for [02:09:10] the [02:09:10] world's [02:09:11] largest [02:09:11] AI [02:09:11] systems [02:09:12] or [02:09:12] you're [02:09:12] talking [02:09:13] about [02:09:13] MI430 [02:09:14] and [02:09:14] MI350 [02:09:15] for [02:09:15] sovereign [02:09:16] and [02:09:16] enterprise [02:09:17] AI [02:09:17] you [02:09:18] saw [02:09:18] the [02:09:18] passion [02:09:19] in [02:09:19] BOMC [02:09:19] with [02:09:19] Rockham [02:09:20] AI [02:09:20] it's [02:09:21] so [02:09:21] good [02:09:21] to [02:09:22] see [02:09:22] how [02:09:22] much [02:09:22] has [02:09:23] come [02:09:23] through [02:09:23] that [02:09:23] platform [02:09:24] in [02:09:24] terms [02:09:25] of [02:09:25] making [02:09:25] developers [02:09:26] make [02:09:26] it [02:09:26] much [02:09:27] easier [02:09:27] for [02:09:27] developers [02:09:28] to [02:09:28] access [02:09:29] the [02:09:29] AMD [02:09:30] platform [02:09:30] and [02:09:31] then [02:09:31] Jack [02:09:32] talked [02:09:32] about [02:09:32] Gorgon [02:09:33] Halo [02:09:33] and [02:09:33] our [02:09:34] new [02:09:34] KREA [02:09:34] AI [02:09:34] platforms [02:09:35] really [02:09:36] bringing [02:09:36] together [02:09:36] the [02:09:37] leadership [02:09:37] of [02:09:38] the [02:09:38] end [02:09:38] to [02:09:38] end [02:09:38] AI [02:09:39] story [02:09:39] including [02:09:40] PCs [02:09:40] and [02:09:41] physical [02:09:41] AI [02:09:41] so [02:09:42] lots [02:09:43] of [02:09:43] new [02:09:43] information [02:09:43] but [02:09:44] I [02:09:44] want [02:09:44] to [02:09:44] say [02:09:44] a [02:09:45] very [02:09:45] very [02:09:46] special [02:09:46] thank [02:09:46] you [02:09:47] to [02:09:47] all [02:09:47] of [02:09:47] our [02:09:47] partners [02:09:48] who [02:09:48] joined [02:09:48] us [02:09:48] today [02:09:49] because [02:09:49] it [02:09:49] is [02:09:50] really [02:09:50] through [02:09:50] those [02:09:51] partnerships [02:09:51] that [02:09:52] we're [02:09:52] able [02:09:52] to [02:09:52] do [02:09:53] the [02:09:53] most [02:09:53] amazing [02:09:54] things [02:09:54] together [02:09:55] so [02:09:56] if [02:09:56] I [02:09:56] just [02:09:56] leave [02:09:57] you [02:09:57] with [02:09:57] a [02:09:57] final [02:09:57] thought [02:09:58] when [02:09:59] I [02:09:59] think [02:09:59] about [02:10:00] where [02:10:00] AI [02:10:00] is [02:10:01] today [02:10:01] the [02:10:02] biggest [02:10:02] change [02:10:03] that [02:10:03] we [02:10:15] are [02:10:16] personal [02:10:16] lives [02:10:17] and [02:10:18] I [02:10:18] have [02:10:18] to [02:10:18] say [02:10:19] that [02:10:19] I [02:10:19] spent [02:10:19] my [02:10:20] entire [02:10:20] career [02:10:20] in [02:10:21] tech [02:10:21] believing [02:10:22] that [02:10:22] high [02:10:23] performance [02:10:23] computing [02:10:23] can [02:10:24] make [02:10:24] the [02:10:24] world [02:10:25] an [02:10:25] incredibly [02:10:26] better [02:10:26] place [02:10:26] and [02:10:27] I [02:10:27] have [02:10:27] never [02:10:28] ever [02:10:28] believed [02:10:29] that [02:10:29] more [02:10:29] than [02:10:29] I [02:10:29] do [02:10:29] today [02:10:30] at [02:10:31] AMD [02:10:31] what [02:10:32] we're [02:10:32] focused [02:10:32] on [02:10:33] is [02:10:33] building [02:10:33] the [02:10:33] technology [02:10:34] the [02:10:35] roadmaps [02:10:35] and [02:10:36] the [02:10:36] partnerships [02:10:36] and [02:10:37] there's [02:10:38] never [02:10:38] been [02:10:38] a [02:10:38] more [02:10:39] exciting [02:10:39] moment [02:10:40] than [02:10:40] today [02:10:40] for [02:10:41] our [02:10:41] 30,000 [02:10:42] plus [02:10:42] engineers [02:10:43] this [02:10:44] is [02:10:44] really [02:10:45] the [02:10:45] next [02:10:45] phase [02:10:45] of [02:10:45] AI [02:10:46] this [02:10:46] is [02:10:46] where [02:10:47] we [02:10:47] make [02:10:48] AI [02:10:48] give [02:10:49] the [02:10:49] AI [02:10:49] the [02:10:50] opportunity [02:10:50] with [02:10:51] all [02:10:51] the [02:10:51] tools [02:10:51] and [02:10:52] all [02:10:52] the [02:10:52] capabilities [02:10:52] to [02:10:53] really [02:10:53] bring [02:10:54] meaningful [02:10:55] real [02:10:55] world [02:10:55] impact [02:10:56] and [02:10:56] I [02:10:56] can [02:10:56] tell [02:10:57] you [02:10:57] I [02:10:57] could [02:10:57] not [02:10:58] be [02:10:58] more [02:10:58] excited [02:10:58] to [02:10:59] build [02:10:59] all [02:10:59] of [02:11:00] that [02:11:00] together [02:11:00] with [02:11:01] you [02:11:01] our [02:11:02] ecosystem [02:11:02] thank [02:11:03] you [02:11:03] so [02:11:03] much [02:11:04] for [02:11:04] joining [02:11:04] us [02:11:04] today

Transcribe Any Video or Podcast — Free

Paste a URL and get a full AI-powered transcript in minutes. Try ScribeHawk →