Samsung Electronics forecast roughly 89.4 trillion won in Q2 2026 operating profit on July 7, an increase of 1,810% from a year earlier. If realized, that would mark its third consecutive record quarter. Demand for AI memory, particularly HBM4, is central to the forecast.
The infrastructure problem behind those earnings extends beyond Samsung. Advanced GPUs need high-bandwidth memory, and only a few manufacturers can supply it at scale. Their production schedules affect when AI servers can ship. The shift toward AI memory also puts pressure on the supply and prices of memory used in conventional servers, phones, and storage.
Why memory bandwidth limits AI performance
High Bandwidth Memory, or HBM, is DRAM built from vertically stacked memory dies. Connections called through-silicon vias run through the dies, and the memory package sits beside the GPU on an interposer, a structure that connects the components.
This arrangement supports a much wider memory bus than conventional GDDR memory. The bandwidth comparison is substantial: 5.1 terabytes per second for Nvidia's H100, versus roughly 600 gigabytes per second for a high-end consumer GPU using GDDR6X. Moving data to the processor quickly is a major part of making large-model inference practical.
A 70-billion-parameter model stored in bfloat16 takes about 140 GB for its weights alone. Those weights must be read during a forward pass. When inference is limited by how quickly hardware can read memory, adding compute capacity measured in FLOPS won't provide the same benefit as increasing memory bandwidth.
That helps explain the value of accelerators such as Nvidia's H100 and H200. The next generation adds another supply dependency. Samsung began mass-producing HBM4 in February 2026, taking the first-to-market position for the sixth-generation standard. HBM4 offers roughly twice the bandwidth per stack of HBM3e, with improved power efficiency.
Nvidia's Vera Rubin platform, shipping in 2026, requires HBM4. An older memory generation isn't a drop-in replacement. Package dimensions, interposer interfaces, and the memory protocol are designed together, so a shortage of the required memory can hold up a complete GPU package.
Memory prices show the pressure on supply
Samsung's memory business reportedly achieved a profit margin exceeding 70% in Q1 2026. The report places that above Nvidia's GPU margins and TSMC's foundry margins. Such high margins suggest considerable pricing power for a component that buyers need and manufacturers can't quickly produce in greater quantities.
Average DRAM selling prices rose 44% quarter-on-quarter in Q2 2026. NAND prices increased 53% over the same period. Direct AI demand accounts for part of the pressure, with hyperscalers and cloud providers buying large quantities of HBM from Samsung and SK Hynix.
Conventional DRAM also faces an indirect effect. Manufacturers are reallocating wafer starts toward HBM in factories that also produce memory for phones and servers. A wafer start is the beginning of production for a silicon wafer. Devoting more production to HBM leaves less capacity available for standard DRAM unless total factory output expands.
That makes memory procurement an issue even for infrastructure teams that don't buy GPUs. Server refreshes, database deployments, and storage purchases are exposed to the broader price increases. Higher Q2 prices are already feeding into vendor quotes, and Samsung has warned that supply conditions may tighten further in 2027.
Most HBM supply depends on two South Korean companies
Three companies have meaningful HBM production capacity: Samsung, SK Hynix, and Micron. Samsung and SK Hynix account for the overwhelming majority of supply. SK Hynix was first to ship HBM3e in volume and currently supplies roughly 50% to 60% of Nvidia's HBM needs. Samsung is gaining ground with its early HBM4 production, while Micron trails the South Korean pair in capacity and product maturity.
As a result, delivery schedules for Nvidia's most advanced platforms depend substantially on production decisions in Hwaseong and Icheon, South Korea. GPU buyers may place orders with a server manufacturer or reserve capacity through a cloud provider, but those suppliers still depend on the same limited pool of memory manufacturers.
This geographic concentration creates a strategic risk for the United States and Europe, comparable in kind to their dependence on TSMC for leading-edge logic chips. Access to AI compute depends on access to advanced memory as well as GPU designs and fabrication capacity.
South Korea's expansion plan faces power and water limits
On June 29, South Korean President Lee Jae-myung announced a ₩1,350 trillion, approximately $880 billion, public-private investment plan covering semiconductors, AI data centers, and robotics over the next decade. The plan brings forward major fab expansions previously expected to extend into the 2040s, targeting the mid-2030s instead.
Samsung alone is committing approximately ₩1,000 trillion, about $648 billion, domestically across chips, AI infrastructure, next-generation batteries, and displays. Those commitments extend well beyond HBM, but advanced memory is an important part of the strategic rationale. South Korea has a lead in the technology needed to supply large-scale AI deployments and intends to expand it.
The national plan also targets 8.4 gigawatts of AI data center capacity by 2029, putting its ambitions on a scale comparable to major US hyperscaler investments.
Capital isn't the only limit on construction. A single megacluster in the planned zones would require power equivalent to roughly a quarter of Seoul's total demand. Semiconductor fabrication also needs substantial water supplies. Electricity and water infrastructure will affect how quickly new capacity can enter service, even when funding is available. Expansions planned for the next decade offer little immediate relief for procurement through 2027.
Planning for the next 12 to 18 months
HBM is in a familiar supply-cycle position: demand for an enabling component has grown faster than production capacity, giving its manufacturers unusually strong margins. New factories can eventually ease that pressure, but infrastructure buyers need plans that account for the capacity available before then.
- GPU delivery dates depend partly on memory output. Nvidia's ability to ship Vera Rubin at scale is tied to HBM4 production. Samsung's ramp is therefore relevant to GPU availability, alongside GPU fabrication and packaging. H200 reservations also face the broader constraints on HBM supply, though H200 and Vera Rubin use different memory generations.
- Conventional memory budgets need room for sustained high prices. The Q2 increases reflect a shift in production priorities, which supports the expectation that DRAM and NAND prices will remain elevated. Server memory refreshes and storage procurement should account for that risk through at least 2027 rather than assume a quick reversal.
- Inference efficiency can reduce pressure on scarce hardware. Caching can avoid repeated work, while quantization reduces the memory needed to hold model weights. Speculative decoding is another optimization option, though its benefit depends on the workload and implementation. These techniques affect hardware capacity requirements as well as operating costs.
- The 2027 shortage warning belongs in capacity forecasts. Samsung has flagged the possibility that HBM4 demand will grow faster than wafer starts can adjust. As of July 9, 2026, the next two to four quarters fall within the planning period for major deployments, while reservation windows for large AI compute contracts are already extending beyond it.
A GPU order or cloud reservation doesn't remove the underlying memory dependency. For deployments with fixed deadlines, capacity plans need to account for the required HBM generation, the suppliers' production ramps, and how much useful inference each reserved system can deliver.