Share with your CTO
SK Telecom is betting that AI inference, the compute-intensive process of running a trained model in production to generate answers for users, is where domestic chip alternatives can finally compete with Nvidia. The Korean telco has deployed Rebellions neural processing units to run its 519 billion-parameter A.X K1 model in its own data centers, generating 136.2 billion won ($92.1 million) in AI data center revenue. Longer term, SK Telecom is planning 5 gigawatts of domestic AI data center capacity starting in 2029, with ambitions to reach 15 gigawatts total.
What this means for your business
The 92% Nvidia training dominance figure is not really the story here. Training share is sticky because CUDA, Nvidia’s software ecosystem that developers use to write and optimize AI workloads, creates switching costs that hardware alone can’t dissolve. Inference is a different game. When you’re running a model millions of times a day for paying users, cost per query and watts per query matter more than peak benchmark performance, and that’s exactly where an alternative chip architecture can compete on a level surface.
The strategic posture SK Telecom is building, mixing Nvidia GPUs for training with domestic accelerators for inference, is likely to become the dominant architecture pattern for any organization running large models at commercial scale. The logic is sound regardless of geography. If you’re a CTO whose AI infrastructure roadmap still assumes a single-vendor GPU stack end-to-end, the SK Telecom model is a direct challenge to that assumption. The organizations insulated from this pressure are those still in evaluation mode, running small pilots; the ones exposed are those already locking multi-year, single-architecture commitments for inference workloads.
The vendor whose position quietly weakens here isn’t Nvidia. It’s whichever cloud provider or colocation partner sold you the idea that a homogeneous Nvidia stack is the simplest path to production AI. Heterogeneous inference infrastructure is operationally harder to manage, but the cost and efficiency gap at scale is wide enough that “simpler” stops being a defensible justification. The leading indicator to watch is whether your current inference vendor is investing in workload-routing capabilities that can direct queries to the cheapest-per-token accelerator, not just the one they already sell.
Concept deep-dive: Inference
Inference is what happens after a model is trained: a user asks a question, and the model generates a response. Think of training as writing a textbook and inference as answering a student’s question using that textbook, millions of times per day. Training is a one-time (or periodic) capital expense; inference is an ongoing operational cost. At commercial scale, the energy and compute bill for inference typically dwarfs the training bill, which is why chip economics matter most there.
Based on reporting from SK Telecom, Rebellions expand Korean AI chip infrastructure, originally published 2026-08-07 18:44:00.

