Moving Data, Not Computing, Is AI's Real Bottleneck
An NVIDIA chip architect explains why moving data costs around 200x the energy of computing it, and what that means for teams building AI outside the data center.
Moving data, not doing math, is where AI systems spend most of their energy. That is the core claim from Mochamad Asri, a lead architect at NVIDIA, in a long interview with Gita Wirjawan on the Endgame podcast (episode 277, September 2026). Asri spent six years doing a PhD on exactly this problem at UT Austin, and he now designs chips for a living. When he says the expensive part of AI is not the part most people argue about, that is worth sitting with.
His number: moving data costs roughly 200 times more energy than computing on it. Not 2x. Not 20x. Around 200x, and that ratio shapes everything about how AI infrastructure gets built.
The full interview is on YouTube. This post pulls out the technical argument and what it implies for teams building outside the data center mainstream.
The bottleneck is the road, not the factory
Asri's analogy from the interview: imagine a government office that can serve 1,000 people a day, connected by roads that can only carry 100. The building is not the constraint. The road is.
AI systems have the same mismatch. Arithmetic on a modern chip is cheap. Getting the data to the arithmetic is not. Modern frontier models make this worse by design: a model with trillions of parameters does not fit on one GPU, so it gets sharded across clusters, and every forward pass means gathering data from far away and scattering results back.
Asri breaks the cost into three parts:
- Capacity. The model has to live somewhere, and one GPU cannot hold it.
- Bandwidth. Inside a chip, transfer rates are high. Asri's phrase: step off the chip and you leave the highway for an alley.
- Latency. Users expect answers in milliseconds. A model that responds in a week has no users.
This is not a fringe view. Mark Horowitz's ISSCC 2014 plenary made the same point a decade ago: on modern systems, data movement, not arithmetic, dominates energy and latency. The industry has known for years. The AI boom just made it the binding constraint, which is why interconnects and photonics get so much investment.
Electricity is layer zero
Gita pushed the conversation toward energy supply, and the numbers there are brutal for most of the world. Indonesia generates about 372 TWh of electricity a year (Ember data via World Power Monitor), which works out to roughly 1,300 kWh per person. Gita framed about 10,000 kWh per person as the range where the economies shaping AI already sit. Indonesia's installed capacity would need to grow by hundreds of gigawatts to close that gap. At the current build rate, he estimated, that takes on the order of a century.
Neither speaker treats that as a reason to give up. It is a constraint, and constraints pick the winning designs. Asri's point about DeepSeek applies here: the team worked around hardware export limits with techniques others never needed to invent. Necessity is the mother of invention. An energy-constrained country will not out-build a gigawatt data center, but it can out-invent around one.
The edge way in
The escape route Asri describes is edge AI and small language models. The logic:
- Large models are generalists. They need data-center scale because they promise to do anything.
- Edge models are specialists. A model for one factory's quality control, or one crop disease, does not need general-purpose intelligence, so it does not need data-center power budgets.
- The edge approach also started from privacy: keep sensitive data on the device, and you must shrink the model until the device can run it.
Two supporting facts make this path real today. MIT's Song Han publishes his entire EfficientML.ai course (6.5940) free on YouTube, covering quantization, pruning, and small-model design. And the open-weights ecosystem gives small teams a starting point instead of a blank page. Even Jensen Huang entered X for the first time in July 2026 to co-sign an open letter backing frontier open-weight models, alongside Microsoft and Meta (The New Stack). Open weights plus efficient-model training are the two inputs an energy-poor country needs to participate, and both are available now.
The education angle matters as much as the hardware one. Asri credits a single guest lecture at Tokyo Tech for sending him to Silicon Valley. Stanford's first-year CS153 course now puts Huang, Altman, and Nadella in front of freshmen (WIRED called it AI Coachella), and the lectures are public. Exposure is free. The talent pipeline starts with people seeing what is possible.
What this means for builders
We build Uteke, a local-first memory engine, so read this section knowing that. But the interview is one of the clearest external arguments for why we built it this way.
Uteke gives AI agents persistent memory that runs entirely on the local machine: 98.4% recall@5 on LongMemEval-S, around 45 ms warm query time, zero LLM tokens per query, CPU only, fully offline. Our benchmark methodology post has the full benchmark methodology. Install is one shell command:
curl -sSL codecora.dev/uteke/install | sh
No API keys, no cloud round trips, no data center meter running. That is the same bet Asri describes at infrastructure scale, applied to one layer of the agent stack: move less data, spend less energy, keep the work local. A memory engine is not a model, and we make no claims about that level. It is the boring, local layer that lets the interesting parts run closer to home.
The interview closes with Asri refusing to pick a side in the hardware-versus-software question: "I'm a person. What do you need?" That is the systems view the whole hour argues for. Optimize the real constraint, not the fashionable one. For most of the world, the fashionable one is compute, and the real one is everything that moves data to it.
Watch the interview. Then check what your own stack spends on data movement. The answer is usually uncomfortable, and that is the useful kind of answer.