750 Tokens Per Second: Why Cerebras Wins the Inference Speed Race
Cerebras and OpenAI launched Ultrafast, running GPT-5.6 Sol at 750 tokens per second — 14x the standard rate. The secret is on-chip memory, not raw compute.
Read ArticleArticles about Infrastructure
Cerebras and OpenAI launched Ultrafast, running GPT-5.6 Sol at 750 tokens per second — 14x the standard rate. The secret is on-chip memory, not raw compute.
Read Article
IBM and Together AI just committed $240M to proving open-source inference can undercut the hyperscalers. Here's why this changes the math for every team running AI in production.
Read Article
Mojo 1.0 shipped today with a stability contract, an Apache 2.0 standard library, and MLIR-backed benchmarks that match CUDA on memory-bound workloads. Here's what it means for AI infrastructure.
Read Article
ByteDance has started the largest AI training run ever attempted — on domestic Chinese chips. Here's what the infrastructure reality looks like, and why the hardware story matters more than the parameter count.
Read Article
A use-after-free lurking in Linux's SCTP stack since 2008 just became a full container-escape chain. Here's the operational breakdown and what you need to do right now.
Read Article
Tesla and SpaceX just committed $16.8 billion to build a chip factory in Texas. The real story isn't the money—it's what vertical integration at this scale reveals about who actually controls AI compute.
Read Article
AMD acquired Taalas, which etches AI model weights directly into silicon. At 16,960 tokens/sec on Llama 3.1 8B with no HBM--this isn't optimization, it's a paradigm shift for inference.
Read Article
University of Toronto researchers showed at Black Hat 2026 that GDDR6 Rowhammer attacks can reach a host CPU root shell by bypassing IOMMU — the control the industry assumed was sufficient.
Read Article
A one-year-old startup built CUDA-equivalent chip software in 10 hours using AI coding agents. The implications for Nvidia's moat—and for every team running AI infrastructure—are significant.
Read Article