750 Tokens Per Second: Why Cerebras Wins the Inference Speed Race
Cerebras and OpenAI launched Ultrafast, running GPT-5.6 Sol at 750 tokens per second — 14x the standard rate. The secret is on-chip memory, not raw compute.
Read ArticleArticles about AI Engineering
Cerebras and OpenAI launched Ultrafast, running GPT-5.6 Sol at 750 tokens per second — 14x the standard rate. The secret is on-chip memory, not raw compute.
Read Article
IBM and Together AI just committed $240M to proving open-source inference can undercut the hyperscalers. Here's why this changes the math for every team running AI in production.
Read Article
Meta shipped a 30B agent model that runs on a single consumer GPU under Apache 2.0. The benchmarks are competitive, the hardware bar is lower than you think, and the build-vs-buy math just shifted.
Read Article
Tesla and SpaceX just committed $16.8 billion to build a chip factory in Texas. The real story isn't the money—it's what vertical integration at this scale reveals about who actually controls AI compute.
Read Article
AMD acquired Taalas, which etches AI model weights directly into silicon. At 16,960 tokens/sec on Llama 3.1 8B with no HBM--this isn't optimization, it's a paradigm shift for inference.
Read Article
A one-year-old startup built CUDA-equivalent chip software in 10 hours using AI coding agents. The implications for Nvidia's moat—and for every team running AI infrastructure—are significant.
Read Article
Three rigorous 2026 studies reveal a striking paradox: engineers using AI feel 20% faster while measuring 19% slower, and aggregate PR throughput is up just 10% despite 93% adoption.
Read Article
On July 28, 1,178 engineers at OpenAI, Anthropic, Google, and Meta asked the US government to build an AI slowdown mechanism. When the builders raise the alarm, that's the signal practitioners can't ignore.
Read Article
Qualcomm closed its $3.9B Modular acquisition on July 29. For the first time, there's a serious, well-funded alternative to CUDA's software moat—here's the operational read.
Read Article