Infinity, a one-year-old startup founded by Jeremy Nixon, reportedly used AI coding agents to build CUDA-like software for D-Matrix's inference hardware in roughly 10 hours. The result suggests that some of the software work needed to support alternative AI chips could take much less time than it has in the past.
According to reporting on Infinity's work, the company raised $15 million in July 2026, with backing from Touring Capital and researchers from OpenAI and Anthropic. Its target was Corsair, an inference accelerator from Microsoft-backed D-Matrix.
The reported build is a proof of concept, not evidence that Infinity has reproduced the whole CUDA platform or delivered a replacement ready for production at scale. Even with that qualification, it raises a practical issue for Nvidia: software integration costs have helped keep customers on its hardware, and AI coding agents may reduce some of those costs.
CUDA's advantage extends well beyond an API
CUDA, short for Compute Unified Device Architecture, is roughly two decades of software development around Nvidia hardware. It includes low-level operations, optimized kernel libraries, debugging tools and profilers. It also has an enormous installed base of production code. A kernel is a routine that runs on an accelerator, often performing a specific mathematical operation.
Ian Buck, CUDA's creator, now runs Hyperscale and High-Performance Computing at Nvidia. The platform he helped build is deeply embedded in machine learning workflows. A research team can write a training loop without directly choosing CUDA, yet still depend on it through PyTorch internals, cuDNN and other parts of the software stack. Dependencies can extend into checkpoint handling, where a training run's saved state is stored and restored.
That makes switching hardware more complicated than buying a different chip. Existing code must work correctly on the new platform, and it must run efficiently enough to justify the move. Libraries, tools and operational practices all contribute to the switching cost.
This software investment is a major part of Nvidia's competitive protection. A rival accelerator can perform well on a particular task and still be difficult to adopt.
Inference gives competitors a narrower target
CUDA's hold is especially strong in training. Training systems can depend on custom CUDA kernels, fused operators that combine several operations, gradient checkpointing logic and distributed frameworks carefully tuned for Nvidia hardware. Moving that work to AMD ROCm or Google TPUs can require months of engineering for a model family.
Inference, the process of running a trained model to produce an answer, presents a different integration problem. Production serving teams focus on measures such as time to first token and decode throughput, or how quickly a model continues generating text after it starts responding. These workloads are narrower than training and increasingly sit behind runtimes such as vLLM, TensorRT-LLM and SGLang.
Marshall Choy at Rebellions argues that CUDA adoption is no longer a factor for inference-focused deployments. That is a strong assessment. The underlying opportunity is that inference runtimes can be retargeted, allowing an alternative chip to support a defined serving workload without reproducing every capability used in training.
Retargeting still takes work. Infinity's proposition is that coding agents can handle enough of that work to make alternative hardware practical sooner.
What the 10-hour build covered
Infinity calls its core technology the Omega algorithm. It is a machine learning system that generates and evaluates algorithms through automated feedback loops. In this application, it writes low-level hardware code, tests the output, scores the results and iterates.
Nixon's team applied that process to D-Matrix's Corsair accelerator and reportedly produced working software described as CUDA-equivalent in about 10 hours. That description concerns the demonstrated implementation. It should not be read as equivalence with CUDA's full collection of libraries, tools and supported workloads.
Corsair is a purpose-built inference accelerator rather than a conventional GPU. It uses TSMC's N6 process, SRAM-based in-memory compute and LP-DDR5 memory. The design targets the decode phase of autoregressive inference, when a model generates successive output tokens.
Reported independent tests pairing Corsair cards with GPUs reduced a 24-second baseline response to under two seconds, described as roughly a 10x speedup. That result concerns the tested setup rather than every inference workload. D-Matrix announced that Corsair entered full production in June 2026, with shipments to priority hyperscalers through the summer.
Like other alternative accelerators, Corsair needs software that makes existing machine learning workloads practical to run. Hardware performance alone cannot cover an expensive or unreliable integration. Infinity's demonstration suggests that agents can sharply compress one part of that engineering effort, though the 10-hour result does not establish the full cost of deploying and maintaining it.
Infinity's broader plan is a universal inference library that works across AI chips. It is designed to generate and optimize low-level code dynamically for the target hardware. If that approach scales as Nixon expects, the software barrier to adopting non-Nvidia inference chips could fall substantially.
Generating code leaves a verification problem
Chris Lattner, co-founder of Modular and creator of Swift, LLVM and MLIR, has called the excitement around AI-generated chip software “very overblown.” His objection is that code generation accounts for only part of production software development.
Teams still need to verify correctness, debug failures, characterize performance and test integration with the rest of the system. A feedback loop that tests generated code is useful, but it does not by itself establish that the code behaves correctly across the conditions a production service will encounter.
Bing Xu describes verification as the biggest bottleneck. Under that view, competitive advantage may shift toward whoever can reliably automate correctness checks for generated low-level code. That remains an unsolved problem.
These objections limit what can reasonably be concluded from Infinity's demonstration. A working implementation in 10 hours is promising; a supported production system requires more evidence. The optimistic forecast is that the gap could close nearer to 18 months than five years, given improvements in AI coding capabilities since 2024. The demonstration alone cannot establish that timeline.
Nvidia also controls infrastructure above the chip
Nvidia's software position extends beyond CUDA. In December 2025, it acquired SchedMD, the company that develops Slurm, a leading job scheduler for high-performance computing. Slurm runs on more than half of the systems in both the top 10 and top 100 of the TOP500 supercomputer list.
A scheduler decides where and when workloads run on a cluster. That places it above individual accelerators in the infrastructure stack and makes hardware support important to operators managing shared computing resources.
Nvidia described Slurm as part of the critical infrastructure needed for generative AI. Daniel Newman at Futurum Group interpreted the acquisition as deepening Nvidia's CUDA advantage. Owning SchedMD gives Nvidia influence over the development of a widely used part of cluster infrastructure, including the pace and quality of hardware integration.
Nvidia said it would continue developing Slurm as vendor-neutral open-source software. There is no finding here that Slurm preferentially routes jobs to Nvidia hardware. The concern is conditional: if Nvidia-native hardware receives better integration or earlier support, competing chips could face an adoption obstacle even after their own software improves.
That makes the integration roadmap worth following. The vendor-neutral commitment matters, as does how it is carried out over time.
Decisions for infrastructure teams
As of August 5, 2026, the practical response is to distinguish new inference deployments from established training systems. Their switching costs are different, and a fast software demonstration does not make them interchangeable.
- Evaluate alternatives for new inference workloads. Corsair is in production. Reporting on the AMD and Anthropic deal announced in July describes MI450-series GPUs shipping into AMD Helios at a 2-gigawatt scale. These developments strengthen the case for a real workload evaluation rather than automatically selecting H100s or B200s. A deployment announcement or an isolated benchmark still does not substitute for testing the intended serving stack.
- Treat existing training dependencies as a real cost. A codebase built around CUDA-specific kernels and custom operators tuned for Nvidia hardware is unlikely to move next quarter. Staying on Nvidia can be a rational engineering choice. New code can still avoid adding unnecessary hardware-specific dependencies.
- Track Slurm's development. Operators of multi-tenant clusters should watch hardware support, integration priorities and Nvidia's handling of its vendor-neutral commitment. Those details can affect the practicality of running mixed hardware.
- Separate faster integration from finished modernization. Infinity's approach may also apply to platform engineering and legacy-system integration. Work previously too expensive to attempt could become affordable. Verification, debugging and ongoing support remain part of the cost, even when agents write the initial code quickly.
CUDA benefits from a reinforcing cycle: adoption produces more CUDA-dependent code, which raises switching costs and encourages further adoption. AI-generated hardware software could weaken that cycle in inference by making parts of the porting process cheaper. The next useful evidence will be whether those ports remain correct, fast and maintainable across real production workloads.