Share with your CTO
xAI is positioning Grok 4.5 as the cost-efficient workhorse for engineering teams, trained alongside Cursor on datasets spanning coding, math, science, and engineering. The model runs at 80 tokens per second and is priced at $2 per million input tokens and $6 per million output tokens. xAI claims it resolves SWE Marathon benchmarks at 29%, ahead of Anthropic’s Opus 4.8 at 26%, while consuming roughly half the output tokens per task compared to leading competitors, making it the cheapest path to comparable agentic coding performance.
What this means for your business
Token efficiency is the number that should stop your infrastructure team. If your agentic coding pipelines are running on Opus 4.8 or GPT-5.5 at 67,000 output tokens per SWE Bench task, and Grok 4.5 resolves the same task at roughly 16,000 tokens, the cost differential is not marginal. At scale across hundreds of concurrent agentic workflows, that gap becomes a capex decision, not a model preference.
The Cursor partnership is the more strategically interesting signal. xAI trained Grok 4.5 with Cursor’s active involvement, not just on code datasets, but on real engineering task distributions. That is meaningfully different from training on GitHub scrapes. When a model learns from the actual prompt patterns of professional software engineers, it converges faster on what “correct” means in production contexts, not just in benchmark harnesses. Your engineering teams already using Cursor are effectively already inside xAI’s training feedback loop.
The benchmark picture is genuinely mixed, and that honesty matters. On DeepSWE 1.1, Grok 4.5 trails both Fable and Opus 4.8 by a meaningful margin. On SWE Marathon it leads. No model wins every evaluation, and any vendor claiming otherwise is selecting benchmarks. The question worth holding: which benchmark distribution actually maps to your team’s specific task mix, long-horizon agentic runs versus single-shot debugging versus greenfield scaffolding?
Concept deep-dive: Agentic rollouts
Agentic rollouts are multi-step AI task executions where the model plans, acts, checks results, and iterates without human intervention at each step. They exist because most real engineering work cannot be completed in a single prompt-response cycle. Think of it like the difference between asking a contractor a question versus hiring them to build a room. xAI’s RL training stack lets these rollouts run for hours across tens of thousands of GPUs simultaneously, which is why token efficiency compounds so dramatically into cost savings at enterprise scale.
Based on reporting from Introducing Grok 4.5 | SpaceXAI, originally published 2026-07-16 03:00:00.

