750 Tokens Per Second: Why Cerebras Wins the Inference Speed Race
Cerebras and OpenAI launched Ultrafast, running GPT-5.6 Sol at 750 tokens per second — 14x the standard rate. The secret is on-chip memory, not raw compute.
Read ArticleArticles about Technology
Cerebras and OpenAI launched Ultrafast, running GPT-5.6 Sol at 750 tokens per second — 14x the standard rate. The secret is on-chip memory, not raw compute.
Read Article
Mojo 1.0 shipped today with a stability contract, an Apache 2.0 standard library, and MLIR-backed benchmarks that match CUDA on memory-bound workloads. Here's what it means for AI infrastructure.
Read Article
ByteDance has started the largest AI training run ever attempted — on domestic Chinese chips. Here's what the infrastructure reality looks like, and why the hardware story matters more than the parameter count.
Read Article
A use-after-free lurking in Linux's SCTP stack since 2008 just became a full container-escape chain. Here's the operational breakdown and what you need to do right now.
Read Article
Tesla and SpaceX just committed $16.8 billion to build a chip factory in Texas. The real story isn't the money—it's what vertical integration at this scale reveals about who actually controls AI compute.
Read Article
AMD acquired Taalas, which etches AI model weights directly into silicon. At 16,960 tokens/sec on Llama 3.1 8B with no HBM--this isn't optimization, it's a paradigm shift for inference.
Read Article
University of Toronto researchers showed at Black Hat 2026 that GDDR6 Rowhammer attacks can reach a host CPU root shell by bypassing IOMMU — the control the industry assumed was sufficient.
Read Article
Andy Pavlo spent a decade teaching CMU students to read ClickHouse internals. Now he's joined the company — and the timing tells you everything about where real-time analytics is headed.
Read Article
Go 1.27 lands three production-ready changes: generic methods that close the generics gap, post-quantum ML-DSA in the standard library, and a json/v2 rewrite that finally enforces UTF-8 and rejects duplicate keys.
Read Article