In the early hours of February 20, 2025, xAI launched Grok 3 and made it available to all users of X (formerly Twitter). Elon Musk himself stated that “Grok 3 is the smartest AI in the world”, but since he owns the company, this statement may be biased.
So, to test this statement, we carried out a independent assessment comparing Grok 3 with the main large language models on the market: Claude 3.5 Sonnet, OpenAI o1, Gemini 2.0 Flash and Llama 3.2. We submitted these six models to three business challenges and asked the AIs themselves to evaluate the responses – first without knowing who the author was (blind test) and then with the names revealed.
The results? Grok 3 and its variant Grok DeepSearch outperformed other AIs by the assessment made by competing Artificial Intelligences themselves, who went so far as to describe it as “innovative”. But does that mean he’s really the best? Let’s explore the details.
Methodology: How do we evaluate AIs?
To guarantee a impartial test, we used two distinct approaches:
1 – Blind Test:
The LLMs analyzed the outputs without knowing who had produced them. The results of each AI were assigned to “Person 1”, “Person 2”, and so on up to “Person 6”. We then asked each Artificial Intelligence to evaluate the material produced by these “people” and state which solution was the best.
Responses from all language models were recorded in a public spreadsheet, allowing for independent fact check.
Below is a screenshot (in Brazilian Portuguese) from OpenAI o1 commenting that Person 1 (Grok) did better than the others:
2 – Revealing the Names:
At this stage, we revealed the model names and asked Claude, GPT-o1, Gemini and Llama to analyze the answers again, now knowing that the answers were given by their “peers”.
Claude 3.5 Sonnet highlighted Grok 3, in both versions, as the best in the three challenges. GPT-o1 and Gemini 2.0 Flash did the same. Llama 3.2 was unable to evaluate, due to its limitation in dealing with very large texts.
For everyone wanting to check these reviews, we left links to them at the end of this article.
What did the other AIs say?
Claudius 3.5 Sonnet: “If I had to choose the AI that performed best in all three challenges, I would recommend Grok 3 DeepSearch for its combination of detail, practicality, efficient resource allocation and realistic consideration of specific challenges.”
OpenAI o1 also ranked Grok at the top of its evaluation, although next to itself when we sent it the names of the “participants”.
Below, a screenshot of Gemini, who when evaluating the first challenge (he was unable to evaluate them all due to limited tokens), continued to point out Grok 3 DeepSearch as the best, placing itself in second place.
Bottom line: Which AI is best for you?
The tests revealed different characteristics in the models evaluated. Grok 3 and its variant Grok DeepSearch were frequently highlighted for combining innovation with practical solutions, according to the analyzes.
Meanwhile, Claude 3.5 Sonnet stood out for the clarity and organization of his answers. OpenAI o1 impressed us for its depth in more complex analyses, showing strength in scenarios that require detail.
Although Grok 3 has been praised by competitors, choosing the best AI depends on the context and specific needs of each user. The tests, focused on professional scenarios, show that performance varies by task or area, and is not universal for all cases.
Right now, Grok 3 differentiates itself by offering its most advanced model for free and by including features such as internet search and integration with the X database. While OpenAI, Claude and Gemini usually charge a fee of approximately 20 dollars per month for more intense use of their most advanced AI models.
Recent data shows that it is already the second most downloaded productivity application on the AppStore, ahead of competitors such as Gemini and Microsoft 365 Copilot.
Artificial Intelligence continues to evolve rapidly, with constant advances that promise to further transform its capabilities. Experts point out that current models are just the beginning, and many innovations are to come in the next years.

Learn more about the AIs that were tested
Grok 3 (xAI)
Developed by xAI, by Elon Musk, Grok 3 was launched in February 20, 2025. Its main innovation is the integration with X (Twitter), allowing access to data in real time. Additionally, it includes self-correction mechanisms to enhance responses and special modes, such as “Big Brain”, which improves your performance in mathematics and programming.
Grok DeepSearch (xAI)
Specialized version of Grok 3, aimed at advanced reasoning and creativity. Uses a special mode called “DeepSearch”, which allows you to explore out-of-the-box solutions, combining sophisticated heuristics and market insights.
Claude 3.5 Sonnet (Anthropic)
Created by Anthropic, a company founded by former OpenAI researchers, Claudius 3.5 Sonnet is known for its clarity and security in the answers. It avoids ambiguities and prioritizes organized and well-structured strategies, being a reliable AI for more conservative planning.
OpenAI o1 (OpenAI)
OpenAI o1 was launched in 2024 as a model focused on deep reasoning and complex problem solving. It breaks problems down into smaller steps and offers detailed analysis, including accurate timelines and metrics.
Gemini 2.0 Flash (Google)
Part of the family Gemini, the model 2.0 Flash is a lighter and faster version, optimized for operational efficiency. Your focus is to deliver direct and easy-to-execute answers, being useful for practical and quick tasks.
Llama 3.2 (Meta AI)
Developed by Meta AI, the Call 3.2 is a model open-source, accessible and efficient in basic tasks. However, is still in the maturation phase and presented difficulties in dealing with complex challenges.
Challenges submitted for each AI
We submit the models to three real business scenarios, evaluating their creativity, feasibility and clarity.
Challenge 1: Sustainable Fashion E-commerce
The first challenge was proposed through the following prompt:
A Brazilian affordable fashion e-commerce startup, focusing on sustainable clothing, wants to increase its sales by 30% in 3 months. She operates with a marketing budget of R$2,000 (around $400) and has a lean team of 5 people. Create a detailed plan for a growth campaign, considering the high competition from big players like Shein and Renner, and the young audience’s preference for social networks. Include specific strategies, budget allocation, and measurable success metrics.
Challenge 2: Coffee Cooperative
The second prompt question was as follows:
A cooperative of small coffee producers in the interior of Minas Gerais wants to expand its distribution to regional supermarkets within a 200 km radius in 6 weeks, with an operating budget of R$5,000 (about $1,000). They have 10 associated farmers, limited production and face challenges such as rural logistics and negotiation with retailers. Develop a practical plan to achieve this growth, detailing actions, costs and how to measure success (e.g. volume sold or new contracts).
Challenge 3: SaaS for SMEs
The third challenge involved a Software as a Service company:
A Brazilian software-as-a-service (SaaS) company that offers a financial management tool for small and medium-sized businesses wants to acquire 50 new customers in 2 months, with an acquisition budget of R$3,000 (around $600). The operation is 100% digital, but the sales team is made up of just 3 people. Create a detailed growth plan, including prospecting tactics, resource allocation, and KPIs to measure success (e.g., conversion rate or cost per customer). Consider SMB resistance to recurring subscriptions.
Links to videos and access to conversations
Comparative Analysis between AIs
Teste cego (Grok 3 DeepSearch): https://x.com/i/grok/share/0n7hgzTI34uYXTi5f9fizv2qr
Blind test (OpenAI o1): https://chatgpt.com/share/67b8cfb8-0f0c-8001-b1ad-ab8f320236a0
Blind test (Claude): Claude.mp4
Blind test (Gemini): Gemini.webm
Test with revealed names (Grok 3 DeepSearch):https://x.com/i/grok/share/zRyxfPF7jaU078nxOabXuUUgF
Test with revealed names (Open AI o1): https://chatgpt.com/share/67b8d591-2b54-8001-bd44-b36a9f93013d
Tests with names revealed (Claude): Claude (1).webm
Tests with names revealed (Gemini): Gemini (1).webm
Spreadsheet link with all results: Click here
Links for response to Grok 3: Prompt 1, Prompt 2 and Prompt 3
Links for responses to Grok 3 DeepSearch: Prompt 1, Prompt 2 and Prompt 3
Links to Claude 3.5 Sonnet’s answers: Prompt 1, Prompt 2 and Prompt 3
Links to OpenAI o1 answers: Prompt 1, Prompt 2 and Prompt 3
Links to Gemini 2.0 Flash Answers: Prompt 1, Prompt 2 and Prompt 3
