The reliability of general-purpose large language models (LLMs) in specialized industrial environments remains critically low, according to a new report from Reshape Automation. The Industrial AI Accuracy Index 2026 reveals that even the most advanced models struggle to navigate the complex relationships inherent in industrial parts catalogs and manufacturer specifications. While these models can perform broad linguistic tasks, they frequently fail when required to provide the precise, technical data necessary for industrial procurement and engineering. This performance gap suggests that "bolted-on" search capabilities are insufficient for high-stakes enterprise applications where a single incorrect part number can lead to significant operational errors and incorrect shipments.
Reshape Automation Benchmarks GPT-6 Astra, Claude, and Gemini
The Industrial AI Accuracy Index 2026 evaluated three prominent models—GPT-6 Astra, Claude, and Gemini—using 100 questions derived from real-world inquiries directed at distributors and Original Equipment Manufacturers (OEMs). The test covered data from 14 industrial manufacturers, including Siemens, Festo, Rittal, and ATI Industrial Automation. The results indicate a massive disparity between model performance with and without web search capabilities. Without web access, all three models scored between 12% and 14% accuracy. When web search was enabled, GPT-6 Astra achieved a top score of 52%, while Claude reached 40.5%.
The report highlights a dangerous trend of "confidence without correctness." Approximately one-third of the incorrect answers provided by the models included specific part numbers or figures without any qualifying caveats. In one instance, Claude, utilizing web search, incorrectly confirmed three separate times that a Siemens communication module was compatible with a specific soft starter, despite the verified answer being "no." Furthermore, the models demonstrated significant inconsistency; in 66 out of 500 total test pairings, the models provided different answers to the exact same question. This lack of determinism poses a direct risk to industrial workflows where consistency is mandatory.
Technical Limitations in Configured Parts and Knowledge Retrieval
A critical failure point identified in the study involves configured parts—items where the part number is generated based on specific manufacturer ordering rules rather than being a static entry in a catalog. In these scenarios, GPT-6 Astra with web search scored only 22.7%, while Claude scored 0%. This suggests that current LLM architectures struggle to reconstruct the logical relationships between disparate data points, such as those found in a catalog, a datasheet, and a cross-reference table.
Reshape Automation CEO Juan Aparicio notes that industrial questions often require synthesizing relationships that may not be explicitly written down in a single document. To address this, the company is positioning its ReshapeX family of AI agents as a specialized alternative. Unlike general-purpose models, ReshapeX utilizes a knowledge graph built from manufacturer catalogs, datasheets, and engineering rules. This architecture is designed to ensure that every fact carries a verifiable source and that the agent provides the same answer to the same question every time. The company claims this approach can result in 60% to 90% time savings in technical support and customer service, as well as an 18.6% increase in average order value.
Key Takeaways
- Leading AI models scored a maximum of 52% accuracy on industrial parts questions when using web search, falling to 12%–14% without it.
- Approximately 33% of incorrect model responses provided specific part numbers or figures without any cautionary language or caveats.
- Performance dropped significantly on configured parts, with Claude scoring 0% and GPT-6 Astra scoring 22.7% in this category.
TechInsyte's Take
In our view, the Reshape Automation report serves as a necessary reality check for enterprise leaders rushing to integrate general-purpose LLMs into technical workflows. The data signals that "hallucinations" in an industrial context are not merely inconveniences; they are high-cost liabilities. When a model provides a specific, incorrect SKU with absolute certainty, it bypasses the natural skepticism a human operator might apply to a vague answer. This creates a "false sense of security" that can degrade the integrity of the entire supply chain. For CIOs and CTOs, the strategic takeaway is clear: general-purpose models with web-search grounding are suitable for broad information retrieval, but they are currently unfit for autonomous technical decision-making. Success in industrial AI will likely require moving away from "bolted-on" search toward structured, graph-based knowledge architectures that prioritize deterministic accuracy over linguistic fluency.
Questions & Answers
How do general-purpose AI models fail when handling industrial part inquiries?
General-purpose models often fail because they attempt to rebuild complex relationships between facts—such as those found in catalogs, datasheets, and cross-reference tables—on the fly. This leads to inconsistent answers, where a model might provide different SKUs for the same request, or "confident" errors where a specific but incorrect part number is provided without any disclaimer.
What is the specific risk of using LLMs for industrial procurement and distribution?
The primary risk is the shipment of incorrect, incompatible components due to inaccurate AI advice. As demonstrated in the Siemens test case, a model can incorrectly validate part compatibility, leading to orders for equipment that will not fit the intended application, thereby causing operational delays and increased costs.
How does a knowledge graph approach differ from standard web-search grounding?
Standard web-search grounding relies on the model retrieving and interpreting information from the live web, which can be inconsistent or unverified. A knowledge graph approach, as used by ReshapeX, assembles manufacturer-specific data—including catalogs and engineering rules—into a structured format before any questions are asked, ensuring that every answer is tied to a verified source and remains deterministic.
Why did the models perform significantly worse on configured parts?
Configured parts require the AI to follow specific manufacturer ordering logic to build a valid part number rather than simply looking up a pre-existing string. Current models struggle to execute these multi-step logical rules accurately, resulting in much lower accuracy scores compared to simple catalog lookups.
Source: Reshape Automation