New Study Reveals How ChatGPT, Claude, Grok, and DeepSeek Use Web Search Differently
Summarized by AI from reporting by ArXiv cs.AI, published under our editorial policy.
A first-of-its-kind study analyzed how four major AI chatbots—ChatGPT, Claude, Grok, and DeepSeek—use web search. Researchers found significant differences in search invocation, query formulation, and domain preferences, impacting the reliability and accuracy of their answers.

Key takeaways
- Researchers analyzed web search strategies across four major AI chatbots: ChatGPT, Claude, Grok, and DeepSeek.
- The study found significant differences in how each platform invokes web search and formulates queries.
- Understanding these differences can help users choose the most reliable platform for specific types of queries.
Researchers from multiple institutions released a study titled 'Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses'. The paper, published on arXiv, examines how four major AI chatbots—ChatGPT, Claude, Grok, and DeepSeek—use web search to answer user queries.
Study Methodology: Real-World and Controlled Experiments
The study combined real-world user interactions (in vivo) with controlled experiments using the same platforms' models via their APIs (in vitro). Researchers investigated the quality of the agents' decisions to invoke web search, their strategies for formulating queries, and potential domain preferences in the search results. They found that each platform had distinct approaches to web search, with varying levels of effectiveness and accuracy.
Key Differences Across Platforms
The research highlighted several key differences:
1. Search Invocation: Some platforms were more likely to use web search for certain types of queries, while others relied more on their pre-trained knowledge. 2. Query Formulation: The study found that the way each platform formulated search queries varied significantly, affecting the relevance and quality of the results. 3. Domain Preferences: Certain platforms showed a preference for specific domains or types of websites, which could influence the accuracy and comprehensiveness of the answers provided.
Why It Matters for Everyday Users
This study sheds light on how AI chatbots gather and present information, which can impact the reliability and usefulness of their responses. For everyday users, understanding these differences can help in choosing the right platform for specific types of queries. For example, if you need up-to-date information, a platform that frequently uses web search might be more reliable than one that relies heavily on pre-trained data.
What You Can Do Today
If you use AI chatbots like ChatGPT, Claude, Grok, or DeepSeek, pay attention to how they source their information. For instance, if you ask ChatGPT a question that requires recent data, you can check if it mentions using web search. If not, you might want to verify the information through other sources. Additionally, you can experiment with different platforms to see which one provides the most accurate and relevant answers for your needs.
Frequently asked
- Which AI chatbots were included in the study?
- The study included ChatGPT, Claude, Grok, and DeepSeek.
- What were the main findings of the study?
- The study found differences in search invocation, query formulation, and domain preferences across the four platforms.
- How can this research help everyday users?
- It can help users understand which AI chatbot is best suited for different types of queries, especially those requiring up-to-date information.