In July 2025, xAI announced the release of Grok 4, a new AI model described as highly capable in reasoning tasks. Accompanying this is Grok 4 Heavy, an enhanced version that utilizes multi-agent collaboration. This post provides an explanation of Grok 4 and Grok 4 Heavy, details from the announcement, and upcoming developments in xAI’s offerings.
Grok 4 Heavy unlocks you instant access to all of this and there’s much more on the way!… pic.twitter.com/l4xhoowZkW
— X Freeze (@amXFreeze) August 2, 2025
Grok now automatically decides how much to think about your question!
You can override “Auto” mode and force heavy thinking at will with one tap. https://t.co/lCU0YI1ALq
— Elon Musk (@elonmusk) July 26, 2025
Most people think Grok Heavy is just another $300 AI plan. It’s not.
Grok 4 Heavy runs multiple smart agents at once, they brainstorm, cross-check, and correct each other in real time.
The result? More depth, better accuracy, zero echo chamber. It’s like having a team of AIs… pic.twitter.com/sIkD1pkGey
— Open Insights (@OpenInsightss) August 2, 2025
My thoughts on Grok 4 Heavy after 12hrs:
Crazy good!
“Create an animation of a crowd of people walking to form “Hello world, I am Grok” as camera changes to birds-eye.”
And it 1-shotted the *entire* thing.
No other model comes close.
Watch the full clip. pic.twitter.com/4j1GXNIF9O
— Mckay Wrigley (@mckaywrigley) July 10, 2025
The Announcement: Grok 4’s Capabilities
xAI introduced Grok 4 on July 9, 2025, highlighting its performance across various benchmarks. The model demonstrates strong results in reasoning-oriented evaluations, achieving high scores in tests that previous models found challenging.Key details include:
- Benchmark Performance: Grok 4 achieved 15.9% on the ARC-AGI-2 benchmark, outperforming other models in this area. It also reached 50% on Humanity’s Last Exam (HLE), a benchmark with over 2,500 problems in math, sciences, engineering, and humanities. Additional results: 88% on GPQA Diamond, 94% on AIME 2024, and 87% on MMLU-Pro.
- Technical Features: It includes a 256k token context window, support for text and image inputs, function calling, and structured outputs. It processes at approximately 75 output tokens per second.
- Reasoning Approach: As a reasoning model, Grok 4 processes queries with deliberation, and users can select a “heavy thinking” mode for more detailed analysis.
According to xAI’s communications, Grok 4 is intended to support advancements in scientific and technological fields over the next 1-2 years. This release positions xAI ahead in certain independent evaluations compared to models from other organizations.
What is Grok 4 Heavy? A Multi-Agent System
Grok 4 Heavy builds on Grok 4 by incorporating a multi-agent framework to address complex queries. This approach involves multiple agents working together.
- Multi-Agent Structure: Agents operate in parallel, generating ideas, reviewing outputs, and correcting errors collaboratively. This method aims to improve accuracy for tasks prone to errors in single-agent systems.
- Performance Metrics: On HLE, Grok 4 Heavy scores 50.7%, and it performs well in reasoning-intensive benchmarks such as coding (LiveCodeBench & SciCode) and math (AIME24 & MATH-500). It is noted for providing detailed responses in advanced applications.
Grok 4 Heavy requires significant computational resources and is designed for scenarios demanding high reliability.
Access and Subscription Options
Grok 4 is accessible to SuperGrok and X Premium+ subscribers on platforms including grok.com, x.com, and the Grok/X apps for iOS and Android. Grok 4 Heavy is available through the SuperGrok Heavy subscription, which offers access to the model, increased usage limits on Grok 4, and early access to new features. This tier is suited for users requiring advanced capabilities.For information on pricing and plans, visit https://x.ai/grok. Details on xAI’s API for model integration are available at https://x.ai/api.
As of August 2, 2025, the SuperGrok Heavy subscription from xAI, which grants access to Grok 4 Heavy along with higher usage limits, priority features, and early betas, is reported to cost $300 per month based on announcements from July 2025.
This positions it as one of the more expensive AI subscription tiers available, targeting users who require advanced reasoning capabilities for complex tasks.
Upcoming FeaturesxAI is introducing beta features alongside the models:
- Grok Imagine: A beta tool for AI-generated videos with sound, focused on creative content. It is in early stages and may have limitations. A more advanced video model is under development using 110,000 NVIDIA GB200 GPUs.
- Valentine Beta: An additional feature in beta, with access expanding to Grok Heavy subscribers.
- Other Enhancements: Grok 3 includes voice mode on iOS and Android apps, with potential future additions like text-to-video, simulation, and vision capabilities, possibly in Grok 5.
