Grok 3 Review: What Happened to xAI’s AI Assistant in 2026
- August 24, 2026
- 0
Search for “Grok 3 review” today and you’ll find dozens of pages still treating it like the current thing to try. It isn’t. xAI retired the grok-3 model
Search for “Grok 3 review” today and you’ll find dozens of pages still treating it like the current thing to try. It isn’t. xAI retired the grok-3 model
Search for “Grok 3 review” today and you’ll find dozens of pages still treating it like the current thing to try. It isn’t. xAI retired the grok-3 model from its API on May 15, 2026, and every request that still points to it gets quietly redirected to a newer model. The consumer apps moved on even earlier. If you’re trying to decide whether Grok 3 is worth your time in 2026, the honest answer is that you can’t use it anymore, and the more useful question is what actually happened to it and what you should look at instead.
That’s the review this article sets out to give. Not a rehash of 2025 launch-day marketing, but a clear-eyed look at what Grok 3 got right, why it didn’t stick around, and where the Grok product line stands now that it’s part of SpaceX.
xAI unveiled Grok 3 on February 17, 2025, during a live-streamed demo where Elon Musk called it “the smartest AI on Earth.” Underneath the showmanship was a real model, detailed in xAI’s own Grok 3 launch announcement. It was trained on the Colossus supercomputer in Memphis using roughly 200,000 Nvidia H100 GPUs, about ten times the compute xAI had used for Grok 2. The launch introduced two features that defined the model going forward: Think mode, a visible chain-of-reasoning process for harder questions, and DeepSearch, an agent that pulled from the web and X to compile research summaries.
Grok 3 positioned itself directly against GPT-4o, early Gemini 2.0, Claude 3.7 Sonnet, and DeepSeek-R1. xAI’s own benchmarks claimed leading scores in math, science, and coding, though it’s worth noting that several of the headline numbers relied on a consensus-of-64-samples scoring method, which isn’t directly comparable to the single-answer scores most competitors reported. That distinction mattered then and matters now if you’re weighing any AI vendor’s self-published benchmarks.
On paper, Grok 3’s numbers were strong. xAI reported a 93.3% score on the 2025 AIME math competition and 84.6% on GPQA Diamond, a graduate-level science benchmark. Independent testers generally confirmed the model was competitive with the top reasoning models of early 2025, particularly for multi-step math and structured logic problems where Big Brain mode allocated extra compute to work through the answer.
Coding was a genuine strength. Grok 3 scored around 79.4% on LiveCodeBench, and developers who tried it for debugging and code generation tasks generally found it comparable to GPT-4o and early Claude 3.7 releases. It wasn’t purpose-built the way today’s dedicated coding models are. If you want a sense of how far coding-specific tools have come since then, our Cursor AI vs Windsurf vs Claude Code comparison covers the current generation of purpose-built coding assistants.
DeepSearch was arguably Grok 3’s most distinctive feature. Instead of a single web search, it broke a query into steps, pulled from multiple sources including X, and returned a synthesized report with its reasoning shown. It was a genuinely useful research tool for anyone who wanted a fast first pass on a topic. If deep research tooling is what you’re actually after now, it’s worth comparing DeepSearch’s approach against dedicated research assistants like the one covered in our NotebookLM review.
Grok’s tightest advantage was always its access to X. Because the model could pull recent posts and trending discussion directly from the platform, it had a genuine edge on questions about breaking news, live events, and public sentiment that other assistants without real-time web access simply couldn’t match at the time.
That strength came with real caveats, though. X data skews toward whatever is loudest on the platform at a given moment, which isn’t the same as balanced or verified information. Grok’s responses on contested topics often reflected that skew, and xAI’s own moderation choices around the model drew scrutiny more than once. If you needed an assistant for professional research where source reliability matters, the X integration was a double-edged feature rather than a clean advantage.
At launch, Grok 3 was briefly free for all users before settling into a paid structure. Access required either an X Premium+ subscription, which cost $22 a month at the time, or a dedicated SuperGrok plan priced around $30 a month or $300 a year, which added DeepSearch and Big Brain mode on top of higher usage limits.
None of that pricing structure carried through unchanged. xAI’s plans have been restructured multiple times since, most recently in 2026 with a SuperGrok Lite tier added in March and a shift to a single shared weekly usage pool across all paid tiers starting in June. Whatever number you find quoted for “Grok 3 pricing” elsewhere is almost certainly out of date, because the model itself no longer has a standalone price. Check xAI’s official pricing page directly if you’re comparing current plans.
xAI retired the original grok-3 API model on May 15, 2026, alongside several other legacy identifiers including grok-4, grok-4.1, and grok-code-fast-1. Requests sent to the old model slug don’t error out; they get silently rerouted to grok-4.3. On the consumer side, the free and paid apps had already moved on to newer Grok generations well before that point.
The bigger context here is corporate, not technical. SpaceX acquired xAI in a deal that closed on February 2, 2026, valuing the combined company at roughly $1.25 trillion. Following SpaceX’s IPO in June, the AI division was formally folded into SpaceX and rebranded SpaceXAI in July. xAI no longer exists as an independent company, and Grok is now a SpaceX product line rather than a standalone AI lab’s flagship.
The model that inherited Grok 3’s spot is a very different animal. As of August 2026, the current flagship is Grok 4.6, released August 12, with a 500,000-token context window built for long agent runs, coding, and research work. Below it sits Grok 4.3 and the Grok 4.20 variants as mid-tier options, and Grok Build 0.1 as a dedicated coding model.
Subscription plans have expanded too. The free tier now caps out around 10 prompts every two hours on a limited version of Grok 4.3, and image or video generation has moved entirely behind a paywall. Paid tiers range from SuperGrok Lite at roughly $10 a month up through SuperGrok Heavy at around $300, which adds a multi-agent mode designed for the hardest, most involved tasks. Every paid tier now includes DeepSearch, Big Brain-style extended reasoning, and voice mode as standard features rather than upsells.
Grok 3 competed on math and reasoning benchmarks with the frontier models of its time, and by most independent accounts it held its own without clearly beating them. That competitive positioning hasn’t really changed with the current Grok lineup, just the names involved. Grok’s real differentiator remains the same one it launched with: live access to X and the web through DeepSearch, which gives it an edge for anything tied to current events or public discussion that a model without that integration would have to search for separately.
Where it still tends to lag is trust and consistency for professional or sensitive work. Enterprises weighing AI vendors for research, coding, or writing tend to prioritize predictable behavior and clear data handling over raw benchmark scores, and that’s an area where Grok has had a rockier public track record than its main rivals. If your work leans toward writing quality specifically, it’s worth comparing Grok’s output against dedicated tools like the ones in our Grammarly vs Hemingway vs ProWritingAid breakdown, since general-purpose assistants and specialized writing tools solve different problems.
Pros
Cons
Grok 3 was a legitimate frontier model when it launched, not a marketing stunt. It held its own against GPT-4o and Claude 3.7 Sonnet on reasoning and coding, and DeepSearch was a real innovation in how AI assistants handled current information. None of that changes the fact that it’s no longer something you can actually use.
If you’re evaluating Grok today, the question isn’t about Grok 3 at all. It’s whether the current Grok 4.6 and SuperGrok lineup fit your needs, and that comes down to how much you value live X integration and DeepSearch-style research against the higher, more complex pricing structure and the trust concerns that have followed the product since launch. For coding and writing specifically, dedicated tools often still edge out a general-purpose assistant. For fast-moving research on current events, Grok’s X access remains genuinely hard to replicate elsewhere.