About this transcript: This is a full AI-generated transcript of This was a data center a year ago… Now it's on my desk from Alex Ziskind, published August 6, 2026. The transcript contains 3,022 words with timestamps and was generated using Whisper AI.
"It was just a year ago that this was a data center, and today it's next to my laptop on my desk. And everything you're about to see, dozens of AIs, all thinking at once, is running on this one machine, not a data center. We've been waiting a long time for the DJX station, and now it's finally here."
[00:00:00] Speaker 1: It was just a year ago that this was a data center, and today it's next to my laptop on my desk. And everything you're about to see, dozens of AIs, all thinking at once, is running on this one machine, not a data center. We've been waiting a long time for the DJX station, and now it's finally here. But the interesting part is not how fast it is, and it is plenty fast. And I'm going to test that out for you. But first, we're going to see what is this thing and what is it made out of, besides a lot of copper. This is the Asus Expert Center Pro ET900N G3. That's a long name, but inside, it's known as the DGX station. That's what powers it. That's the NVIDIA bit, and the GB300 is the chip. We first heard about it alongside the DGX Spark on stage when Jensen Juan announced it. It'll be a beast for a lot of industries that need heavy compute. But today, I'm looking at it as an AI developer. So it's all LLMs and agents. And I'm going big. GB300 is that chip inside, and GB300s have been in servers like the HGX for months now, since last year. I tested one earlier this year on the channel, actually. But this is the desktop variant. The desktop super chip, as it's known. Build for your desk. Ta-da! Now, yes, it has that crazy chip inside, but it also has 748 gigabytes of memory to run your AI. Its memory kind of works like Apple Silicon's unified memory. But instead of one pool, it fuses two. An ultra-fast 256 gig GPU pool and a big 500 plus gig grace pool into one 748 gig pool. I'm saying pool a lot. I want to go swimming. It's kind of hot outside. I'm also hungry. Think of it as unified memory with a fast lane and a not-so-fast lane. A slightly slower lane. I'll get back to that. There's also one neat thing that you don't normally find on machines that are on your desk. Our 400 gigabit NICs, network cards. These are called the Super NICs, and it's using ConnectX 8. These are basically things that you'd find in a data center connecting multiple GPUs together. And it's right here. If you see my videos connecting and clustering DGX Sparks, those have ConnectX 7, but it's basically the same idea. You'd be able to connect multiple boxes with ultra-low latency and high bandwidth. There's also three PCIe slots on there, so you can expand it with more GPUs if you want to. This loaner unit actually shipped with an RTX Pro 4000 inside, but you can also put an RTX Pro 6000 in there. Now, the super chip alone is about 360 watts idle and about 1,400 flat out when it's running. That's just a chip. The whole machine at the wall is more. And while you can plug it into a regular outlet in the United States, which gives you 115 volts and 15 amps, I opted for a 240 volt with 20 amp outlet. That way, I'm not running into any issues. All right, let's put it to work and see what it can do. We're running some big models today, and here's a quick look at exactly what we're running. Quad, Deep, Deep, Neemotron, GLM. Now, these models come in different quantizations, and Blackwell Architecture can do NVFP4 quantization, which means floating point 4, but it's the NVIDIA format, and it runs really well and efficiently on this machine. Here's an example of the NVFP4, and this is Neemotron 3 Super 120 billion parameter model. It's pretty big. I'm also going to be doing DeepSeq V4 Flash, Quen 3 235 billion parameter model, also NVFP4, and we're going to do this one, GLM 5.2. I know you're all excited about that one. That's the big one. As a side note, most models come in 16-bit, and quantization reduces the number of bits, so if it's 4-bit, then you basically shrunk your model 4 times. So GLM 5.2 becomes 465 gigabytes on disk instead of being like 2 terabytes. Not only are they smaller, but they also are faster. And NVFP4 is NVIDIA's own 4-bit format that Blackwell chips like this one and the ones on the Pro GPUs, they can run it natively. That's the whole trick that puts a data setter model on your desk. Now, if you notice, GLM 5.2, even the NVFP4 format, is 465 gigabytes. The DGX station high bandwidth memory, HBM, it's not going to fit the entire model because HBM is only limited to 256 gigabytes. That part is going to be really fast. But if the whole model is not in HBM, it's going to have to spill over to the rest of the memory. That's the system memory, the LPDDR5X. So together, they're considered coherent memory. But what the heck is coherent memory? Is it like memory that's aware of what's going on? Is it self-aware memory? All right, that was pretty dumb. Here's what's really happening. Look at how fast the HBM is. 7 to 8 terabytes per second. Well, if the model spills onto the slower memory, which is 500 gigabytes per second, what happens? Does the whole thing crawl? Well, no. Only the bytes that come from grace. But GLM is a mixture of experts model. And the offloaded experts are exactly what each token reads. So you feel most of the slow lane. That's 24 tokens per second. It's kind of like the same trick as running a giant model on a maxed out Mac Studio. Which are, oh, I moved. They're not here anymore. Bigger than the GPU's memory, but runs anyway because it's shared. Plus, this has an HBM fast tier that Macs don't have. So what's up with all these people buying Mac Minis for hosting their agents? Is that even doable? Or do I need this box? Well, a Mac Mini will run an agent. It'll just be on a much smaller scale. What the big box buys you is hard things and big models. And it does it fast. But we're going to talk about agents. Just hold that thought. I don't expect one AI prop to build an entire project for me. In reality, I'm constantly moving between models depending on the job. GPT for research, Claude for coding, Gemini for massive context, Nano Banana, Mid Journey, Flux for images. And then I've got Sea Dance and Kling for video. That's why Chat LLM by Abacus AI makes sense. It brings day one support for the latest GPT, Claude, Gemini, Grok, DeepSeq, and more in one place the moment they drop. Pick any model from the interface or let Route LLM automatically choose the best model for each prompt. Create professional presentations with graphs and charts and deep research detailed content. Need human sound and copy? Humanize rewrites text to defeat AI detectors. Need visuals? Pick frontier or open source models. And when you need more than chat, Abacus AI agent can help build complex apps and websites, connect payments, or run 24-7 agents that keep working through longer tasks. The best part is app hosting. Back-end database and auth support comes with the subscription. All that starts at just $10 a month. Way cheaper than paying for all those subscriptions separately. Check out chatllm.abacus.ai or click the link below. Now, if you're using this machine running models and you're plugging in your code editor, for example, this is going to be something else. Boom. Okay, there it is. Nemotron 120B going at 185 tokens per second right there in decode. Whoa, we're up to 230. 250? Oh, wow. It just keeps going up. What's going on here? Oh, it's doing a concurrency sweep. We didn't get to that part yet. Run single DGX. All right. Sorry about that. There we go. About 184 tokens per second decode. And we're getting some nice pre-fill there. Pre-fill is important for agents when you're sending in your prompt to process. Because your prompt is not just going to be, hello, how are you? Or write a story or hi, like I sometimes am guilty of showing off on this channel. But that's just chat. Your prompt is going to include all that code and all the files that the agent is automatically going to send in, along with all the context of your conversation thus far. And all that needs to be processed. For that, you need fast prompt processing, also known as pre-fill throughput. Here is the pre-fill for Nemotron. Single stream, so like chatting, basically. 2200 on Nemotron, 2500 on Quen, 1800 on DeepSeq, and GLM is a big chonker, so 243. Also, that's offloaded to that slower memory that I mentioned. That's why we're getting slower numbers here. But what would happen if we were not to use a single stream? If we were to use multiple streams? Everything so far was just me. One user, one chat. But a machine like this isn't built for just chatting with one person. It's built to serve a crowd. All right, they may be pushing it. Maybe like a few people in an office environment. Maybe a few agents. In fact, maybe a lot of agents. That's coming up soon. So let's talk about concurrency. And that's something I showed before on the channel many times. And people were like, why do I need 4,000 tokens a second? Ha! Hold that thought. Concurrency just means how many requests hit the model at the same time. A whole team using it at once. An app with 100 users. Or, as we'll see, a swarm of agents. And here's the interesting question. When I pile all those requests on at the same time, what happens? Does each person slow to a crawl? Does the machine choke? Or does something else happen? Well, let's see. I'm going to kick this off. And right now, I'm going to start with just me. About 180 tokens a second. We already saw this. Now, watch as I add more at the same time. This is 2 now. 250 tokens per second. 4. We're getting up to 300 tokens per second now. 350. 4. Concurrency of 4. Now we're at 400 tokens per second. At concurrency of 8. And we can keep going. This machine will handle it. Here's concurrency of 16. It just keeps going. 700 tokens per second. Now 750. We're over 800 tokens per second with concurrency of 16. And now, concurrency 32. Ladies and gentlemen, we just hit 1,000 tokens per second. And yeah, we have more. We have more. 64 concurrency. That's 64 users or 64 agents. We're at 1,500 tokens per second. All that being generated right here on the fly. 1,700. 1,800. We're at 128 concurrency. That's 128 requests. And it just keeps going. This is just insane. We're on Nemo Tron 120B. One of the really fast models that runs on this. Really tuned well with NVF before. And we've hit 2,600 tokens per second. Now it's not perfect. Each individual request takes a little bit longer. By the way, during that little experiment, we've generated 254,000 tokens in just about a minute or so. And the best peak was 4,096 tokens per second. That's continuous batching. The GPU packs all those requests together and runs them all in one pass. I did the other models too. There's Nemo Tron 120B that you just saw live. Here's Qen. We went to about the same, actually, 5,000. Just over 5,000 tokens per second for 128 users. And this is how concurrency scales. Notice there's a little bit of a dip on both Nemo Tron and Qen at 32, which is kind of weird. Because I would think that would be a little blip, an error. But both models did it several times in a row. So that must be real. DeepSeq does pretty well with that too. 64 seems to be the sweet spot for DeepSeq on this machine. And GLM, remember GLM is offloading to the slower memory. But still, it can all fit 35 tokens per second at a concurrency of 4 and 56 at a concurrency of 16. I didn't do 8. And reading the prompts does the exact same thing. Climbing to about 41,000 tokens per second for Qen 235. About 35,000 tokens per second on Nemo Tron. And we do see a dive in DeepSeq for prompt processing. 64 is really nice for DeepSeq. So this tells you that it depends on what kind of model you're using too. That matters. Even GLM is happy. 1,800 tokens per second at a concurrency of 16. So it's a trade-off. A few users keeps everyone snappy. Pack it full and you get maximum total throughput. And the longer the answer is, the higher the total climbs. But who actually needs thousands of tokens a second across dozens of parallel requests? Hmm, that's not a person. That's an agent. Running a lot of agents hides two different ideas. One is architecture. How do you wire agents together? Any laptop can do that, right? The other is throughput. How many actually run at the same time? That is hardware. The DJX doesn't change your design. It just raises your ceiling a little bit. So the real question is, how many agents can this one box actually run at once? Let's get there. First, just one agent. Here's Visual Studio Code with its agent configured to use Nemo Tron. I'm going to ask, what is this project about? Boom. It's evaluating all the files, processing, and some of this is the thinking stage. This is an agent, so it's not just doing one thing at a time. It's constantly sending back and forth. And look how fast that goes. Wow. We found out what it does. So it has tool use. Everything an agent needs. This project is a command line task tracker called TRK. That's right. I told you I tested four models. This is how fast each of the models is, including GLM 5.2. Quan, Deep Seek, Nemo Tron, GLM. Notice GLM 5.2 is just a little bit slower. That's because it's spilled over. It can't all fit inside high bandwidth memory like Nemo Tron can. Here I hooked it up to Codex. That's OpenAI's agent CLI. Model is NVIDIA, Nemo Tron, Super, 120B. And I'm just going to do a softball here, write a story, boom. And there is our proof that it's actually working. I'm running NVTOP on the actual machine, watching it. And the power is going up to about 715 watts, about 550 on average. 227 gigabytes out of the 250 on high bandwidth memory. It's all fitting inside on the GB300. But one agent is easy. A laptop can do it. Even your Mac mini can do it. Maybe not that fast. But here's the real test. Can this one box run a whole crowd of agents at the same time? Let's find out. So here I got a little agent view. This is the power bar. Right now it's at 205 watts. Temperature is 33. GPU is at 0%. And I can pick how many agents I want to launch. Let's start with 16, shall we? And launch. Here we go. GPU is up to 100%. Power is hitting about 750 watts. Temperature's gone up a little bit. Everything seems to be working. And we're cranking those numbers out. Look at that. Live 16 agents. 1,400 tokens per second. And this is how many total tokens we've done so far. A lot of work is being done right now. A lot of work we're not going to use or see or whatever. But, you know, it's a demo. We've got different kinds of agents too. A coder, a researcher, an analyst, a writer, a planner, a tester, an architect. You get the idea. Let's kick it up a notch. Let's go to 32. You know what? Let's go to 64. Boom. And 900 watts. Wow. 3,500 tokens per second. And we're generating 35, 37, 40,000 tokens. I can't even keep up speaking it. Wow. Look at them all go. These are all generating all at the same time. All doing their own thing. But we're not done yet. Let's try 128. And, of course, you might have guessed this is going to work. And it does work. But look at the speed at which this is happening. We're not hitting the 1,300 watts that this processor is capable of. But we're getting pretty close. We're about 1,000 watts right now, is what I've seen as maximum. Temperature is about 55 degrees. And so far, since I've been talking, we've already generated almost 150,000 tokens. This is a local agent swarm on the GB300. Now, the sweet spot is around 32 to 64 for this model. This, again, is NemoTron right here. It's good enough to peg the GPU while every agent stays responsive. And everyone gets the same big model. That's what a small machine can't do. 24-7, 100 agents, nothing leaves your desk. Zero per token cost. Of course, you have to buy the box. And these boxes are not cheap. Luckily, I didn't have to buy it. I'm borrowing it. Versus thousands of dollars a month in API bills, this is not bad. That's the pitch. Fast for one, scales for many, and 100 at once for free. Just pay your electricity bill, okay? Or they'll turn it off on you. Now, you can also put a GPU in one of those PCIe slots. And an RTX Pro 6000 is one of the ones you can actually fit in there. It's a nice GPU that also scales. If you can't buy a big machine like this, you can buy just the GPU. And you'll see my video right here on how I do that. Thanks for watching, and I'll see you next time.