Energy · Breakthrough
Jason Eshraghian and Rui-Jie Zhu built an AI that runs on the power of a lightbulb.
Jason Eshraghian, Rui-Jie Zhu · Santa Cruz, CA, USA
Large language models are notorious for their energy appetite. A UC Santa Cruz professor and his graduate student found a way to strip the most expensive step out of the math, and in 2024 ran a model with a billion parameters on 13 watts, about the draw of a lightbulb.
The person and the place
Jason Eshraghian, assistant professor of electrical and computer engineering at UC Santa Cruz's Baskin School of Engineering, and Rui-Jie Zhu, his graduate student and the paper's first author.
The problem Jason and Rui-Jie cared about
Running large language models costs real money and real energy, and Eshraghian wanted to know whether that cost was actually necessary. "The idea started by acknowledging the fact that language models like ChatGPT are incredibly expensive in terms of the amount of resources that you need to run them," he said.
The agency moment
In 2024, instead of accepting that bigger, pricier hardware was the only path forward, Eshraghian and Zhu set out to eliminate matrix multiplication, the single most computationally expensive step in running a neural network, and then built custom hardware to prove it could be done. "We got the same performance at way less cost — all we had to do was fundamentally change how neural networks work. Then we took it a step further and built custom hardware," Eshraghian said. They built the low-power prototype system in three weeks.
What changed
Their custom board ran a model with a billion parameters while drawing 13 watts. The same job on GPUs takes roughly 700 watts, which makes their board more than fifty times as efficient. On standard GPUs, the same approach cut memory use about tenfold and ran about 25% faster, matching the performance of language models its own size. That result went out in 2024 as a preprint (arXiv 2406.02528) rather than a peer-reviewed paper, and the comparison is to models at the same scale, not to systems like GPT-4, which are estimated to carry more than a trillion parameters.
The move worth copying
They didn't wait for cheaper chips to make AI less expensive to run. They took the priciest part of the math out of the machine. The copyable part is the question, not the chip design: which expensive step is everyone treating as fixed? Know someone asking that about their own field? Feed the scout at goodinprogress.org.
"We are just a small academic lab that started less than two years ago and we are capable of competing with the giants." — Jason Eshraghian, Santa Cruz Sentinel via GovTech Insider
Sources
Here is where this story came from. Go read the originals.
Curated by Good in Progress from the public record.
Sources last verified .
