MP Marc Popeinfra · open source
sys: online · EST Let's talk
the log · 132 entries

#ai-engineering

Every entry filed under AI Engineering.

RSS
2026.08.167 min read AI API Prices Are Falling Fast. Model Routing Is Becoming Core Infrastructure. Rapid AI price cuts make a fixed model choice expensive to maintain. Task-aware routing offers a way to reduce costs while keeping harder work on more capable models. technologyai-engineeringinfrastructure 2026.08.156 min read Qwen 3.8-27B Gives Self-Hosted Agents a Stronger Case Alibaba's Qwen 3.8-27B combines Apache 2.0 licensing, single-GPU deployment and stronger reported agent benchmarks. The next test is how those gains hold up in local workloads. programmingai-engineeringinfrastructure 2026.08.147 min read Cerebras Reports 750 Tokens Per Second for GPT-5.6 Sol Ultrafast Cerebras and OpenAI's Ultrafast tier promises up to 750 output tokens per second. On-chip memory helps explain the speed, though production performance and capacity remain open questions. technologyai-engineeringinfrastructure 2026.08.137 min read IBM's $240 Million Together AI Deal Puts Open-Model Inference Costs in Focus IBM and Together AI plan a 2,000-chip Blackwell inference cluster for Q1 2027. The deal gives enterprises another option to compare on price, model quality, and operating requirements. ai-engineeringinfrastructuretechnology-leadership 2026.08.117 min read Meta's Muse Glimmer Brings a 30B Agent Model to a Single Consumer GPU Meta says Muse Glimmer runs on a single 24 GB GPU, with competitive agent benchmarks and Apache 2.0 licensing. The release makes local inference more practical, though performance and… programmingai-engineeringsoftware-development 2026.08.088 min read Terafab: Tesla and SpaceX's $16.8 Billion Plan to Control Chip Production Tesla and SpaceX plan to bring chip manufacturing, memory, packaging, and testing into one Texas complex. The case rests on faster development and reliable supply, but building a working… technologyai-engineeringinfrastructure 2026.08.077 min read AMD Plans to Acquire Taalas, Betting on AI Models Built Into Silicon AMD's planned Taalas acquisition targets a costly part of AI inference: moving model weights from memory. Taalas builds those weights into silicon, trading flexibility for claimed gains in… technologyai-engineeringinfrastructure 2026.08.058 min read Infinity's 10-Hour Chip Software Build Tests Nvidia's CUDA Advantage Infinity reportedly built CUDA-like software for an inference chip in about 10 hours. That could lower switching costs, but production testing and Nvidia's broader software stack remain… programmingai-engineeringinfrastructure 2026.08.017 min read AI Coding in 2026: Why Developers Can Feel Faster While Output Barely Changes A randomized trial found developers took 19% longer with AI while believing they were 20% faster. Broader data shows modest throughput gains, with results depending on the task and workflow. ai-engineeringtechnology-leadershipsoftware-development
open channel

Have an infrastructure problem that needs an owner?

Backup, recovery, remote access, fleet automation, or a platform that has to stay up. Tell me what is breaking.

Start a conversation