Why can’t AI use the engineering data we already have?
Most industrial companies do not lack engineering data. They have decades of product structures, drawings, requirements, calculations, simulation results, test evidence, quality records and manufacturing information.
The problem is that this information was created to support individual tools, processes and departments. It is stored across PLM, PDM, ERP, requirements, simulation, test and manufacturing systems, as well as file shares, spreadsheets and local databases. Each environment may work for its original purpose while remaining difficult to connect to the wider engineering process.
AI needs more than access to documents or a large volume of records. It needs current, understandable and traceable context. It must be possible to establish which product, variant, material, requirement or revision a piece of information belongs to; whether it is authoritative; and how it relates to decisions elsewhere in the lifecycle.
Engineering data becomes a foundation for industrial AI only when authorized teams and applications can find it, trust it and reuse it across system and departmental boundaries.
What is engineering data?
Engineering data is the information created and used to define, develop, validate, manufacture and change a product.
It includes obvious records such as CAD models, drawings, bills of material and requirements. It also includes the information surrounding them: assumptions, calculations, simulation inputs and outputs, test results, material properties, software configurations, deviations, change decisions and links between product variants.
Much of the most valuable engineering data is not a neat table. It may be embedded in documents, model files, spreadsheet formulas, test reports or the working knowledge of experienced engineers. Its meaning depends on product context and lifecycle state.
This makes engineering data different from generic corporate information. A value without its unit, configuration, revision, source and relationship to the product can be worse than useless. It can produce a confident but wrong engineering conclusion.
Why does industrial AI depend on connected engineering data?
Useful industrial AI must operate in the context of a real product and a real engineering decision.
An assistant reviewing a requirement needs to know the applicable product variant and regulatory context. A system assessing an engineering change needs to trace affected drawings, components, simulations, tests and downstream manufacturing information. A tool preparing simulation inputs needs current geometry, material properties and boundary conditions.
Without those relationships, AI can search or summarize isolated content but cannot reliably support the workflow.
The dependency is practical:
- AI needs trustworthy context to produce useful output.
- Trustworthy context depends on traceable relationships between engineering records.
- Traceability depends on shared identifiers, clear ownership and authoritative sources.
- AI can only act at workflow speed when that information is available through repeatable, governed access.
This is why adding a model on top of a document repository is not the same as creating an engineering data foundation.
Where does engineering data become trapped?
Engineering information is commonly constrained in three ways.
It is locked inside the system that governs it
PLM, PDM and ERP systems are designed to protect authoritative records through access control, lifecycle states, audit trails and change processes. These controls are necessary. But many core systems were not originally designed to make their information easy to consume across a modern digital organization.
When another team needs the data, it submits a request, creates a custom report or exports a file. The authoritative source remains governed, but its value outside the system is limited.
It is copied into local working tools
Engineers need to get work done. If systems cannot exchange information, they download it into spreadsheets, scripts and local databases. Those copies make the immediate task possible, but they gradually diverge from the source.
Soon, the organization has several plausible versions of the same information. People must determine which file is current and whether a calculation used the right revision. AI trained or prompted with those copies inherits the ambiguity.
Its meaning depends on people
Experienced engineers know which fields matter, which test setup was unusual and which naming convention identifies the right product variant. That context may never have been made explicit in the systems.
AI cannot reliably infer organizational conventions that are inconsistent or undocumented. When meaning depends on knowing whom to ask, the data is not yet an organizational asset.
Why are spreadsheets and exports such a persistent problem?
Spreadsheets are not inherently bad. They are flexible, familiar and valuable for exploration, calculation and one-off analysis.
They become a structural problem when they serve as the permanent connection between core engineering systems. Every export separates information from its authoritative lifecycle. Formulas hide business logic inside cells. Email creates uncontrolled versions. Local naming conventions break traceability.
The apparent efficiency of an export often transfers cost downstream. Several engineers spend time reconciling data, checking revisions and repairing errors because a system-to-system connection was never created.
For AI, the consequence is severe. The model may receive a readable file, but it cannot know whether the file is complete, current or linked to the correct configuration unless that context is preserved.
The objective is not to ban spreadsheets. It is to stop using people and files as the standard integration layer for recurring engineering workflows.
What makes engineering data reusable without losing control?
Reusable engineering data has five characteristics.
1. There is an authoritative source
The organization knows which system owns each important type of information and who is accountable for its meaning and quality. Reuse does not create a new competing master.
2. Product context travels with the data
Information retains its relationship to the product, variant, revision, requirement, material, test or decision to which it belongs. Units, lifecycle state and provenance remain clear.
3. Common identifiers connect the lifecycle
Stable identifiers for products, parts, materials, documents and other key objects allow information to be matched across functions. Names and descriptions change; identifiers preserve the connection.
4. Access is governed and repeatable
Authorized teams and applications can obtain information through documented interfaces or managed data products. Access does not depend on personal contacts or a bespoke extraction project each time.
5. Ownership extends beyond the source system
Someone is responsible not only for storing the data but also for making it understandable and usable by legitimate downstream consumers. Quality is managed in relation to business use.
Together, these characteristics preserve control while allowing information to create value outside the system where it originated.
Why are common identifiers so important?
Industrial AI becomes valuable when it can connect evidence across the product lifecycle.
Suppose leadership wants to reduce the time required to assess an engineering change. The relevant context may include a part in PLM, material information in ERP, requirements in a specialist tool, simulation results on a data platform, test evidence in another system and manufacturing constraints at several sites.
If each environment describes the product differently, teams must match records manually. AI faces the same problem. Similar names are not enough when a wrong match can affect product quality or compliance.
Common identifiers provide the backbone for traceability. They allow the organization to ask which requirements, analyses, tests, suppliers and manufacturing processes relate to a specific product configuration.
This work may sound technical, but the decision is organizational. Functions must agree to preserve and use shared identity across their processes rather than optimizing only within their own system.
Do we need one central engineering data platform?
Not necessarily.
A central platform can make discovery, analytics and cross-system use easier. But moving every record into one place does not automatically resolve ownership, meaning or traceability. It can create another large repository containing disconnected copies.
The more useful goal is a connected engineering data environment:
- Authoritative systems continue governing their records.
- Important data is available through secure, documented interfaces.
- Shared identifiers preserve relationships across systems.
- Reusable data products present information for common engineering needs.
- Teams can discover what exists and understand how it may be used.
- Governance follows the sensitivity and purpose of the data.
The architecture may combine source-system interfaces, event flows, shared platforms and curated data products. For executives, the test is simpler: can an authorized team reuse reliable engineering information without months of custom integration and uncontrolled copying?
What does a good engineering data foundation enable?
A strong data foundation improves engineering before advanced AI is introduced.
It can reduce time spent finding and reconciling information, improve change impact analysis, connect requirements to verification evidence, automate reporting, reuse simulation and test results, and create more reliable feedback between development and manufacturing.
AI compounds those benefits. It can help engineers navigate connected evidence, identify gaps, prepare decisions and orchestrate work across a larger context. Because the underlying information is traceable, outputs can be checked against authoritative sources.
This is the strategic effect: each connected workflow produces better structured and more reusable information, which makes the next automation or AI application easier to build.
How should leaders build the data foundation?
Avoid launching an enterprise-wide “clean all engineering data” program. It is difficult to prioritize, slow to show value and easily disconnected from operational work.
Build the foundation through important engineering workflows.
1. Start with a material decision or delay
Choose a problem such as engineering change impact, requirements verification, simulation preparation, test-result reuse or transfer of product information into manufacturing.
2. Trace the information the workflow requires
Identify where each record originates, how it moves, where copies are created and which context is lost. Document the manual reconciliation performed by engineers.
3. Establish authority and ownership
Agree which system owns each record and which role is accountable for its meaning and quality. Resolve ambiguity before automating it.
4. Preserve identity across systems
Use common product and lifecycle identifiers so that information can be connected reliably. Do not depend on filenames or descriptions as permanent keys.
5. Create governed, reusable access
Expose the required information through documented interfaces or managed data products. Design for the next legitimate consumer, not only the first project.
6. Measure operational improvement and reuse
Track reduced manual effort, faster decisions, fewer errors and shorter integration time. Also track whether another workflow can reuse what was created.
This approach produces value now while incrementally building the wider data foundation.
Who owns engineering data?
No single function can own all aspects of engineering data.
Engineering domains own the meaning, use and quality of information in their processes. System owners protect authoritative records and lifecycle controls. IT provides integration, security, platforms and enterprise standards. Data teams may provide discovery, reusable products and analytical capability.
Leadership must make these responsibilities explicit. When everyone is responsible for “data” in general, nobody is accountable for whether a specific engineering record can be trusted and reused.
A practical ownership question is:
Who must act when another authorized engineering team cannot understand or reliably consume this information?
If the answer is unclear, the data is not yet managed as an organizational product.
What should an engineering executive do first?
Ask leaders from design, simulation, testing, manufacturing and IT to identify one high-value workflow that repeatedly crosses their boundaries.
Do not begin by asking where AI could be added. Ask where engineers currently search for information, copy it, reconcile conflicting versions or depend on a colleague’s knowledge. Quantify the delay and risk created by those steps.
Then fund the workflow improvement with two outcomes: solve the immediate operational problem and make the required data securely reusable for the next team.
The first visible success should not be a new data lake or another dashboard. It should be an engineering decision that becomes faster and more reliable because the right information can finally flow.
What are the implications for engineering leadership?
- Engineering data accessibility is a business performance issue, not only an IT concern.
- Governance and reuse must be designed together; control without accessibility creates workarounds.
- Shared product identifiers require cross-functional leadership alignment.
- Data quality should be improved in the context of valuable workflows.
- Every integration should reduce future integration cost, not solve only one project.
- AI outputs used in engineering decisions must remain traceable to authoritative sources.
Industrial AI does not begin with a model. It begins with an organization that can connect the information required to understand and improve how products are developed.
Frequently asked questions
Do we need perfect data before starting with AI?
No. Perfect enterprise data is not a realistic prerequisite. Start with a valuable workflow, define the information it needs and improve quality at the source. The important requirement is that limitations are understood and outputs remain traceable.
Can AI clean and connect our engineering data for us?
AI can help classify records, extract information and suggest relationships. It cannot decide the authoritative source, product identity, acceptable quality or ownership model on behalf of the organization. Those are engineering and leadership decisions.
Is PLM the source of truth for all engineering data?
PLM commonly governs important product structures and lifecycle records, but engineering decisions also depend on requirements, simulations, tests, software, materials and manufacturing information. The goal is not one source for everything; it is clear authority and reliable connection across sources.
Should data be copied out of core systems for AI?
Sometimes managed copies are necessary for performance, analytics or model use. They should be created through governed processes that preserve source, revision, access rules and update logic—not through ad hoc exports.
How do we justify investment in interfaces and identifiers?
Tie the investment to recurring operational cost: manual transfers, change delays, duplicated analysis, quality risk and repeated project integration. Reusable foundations become more valuable each time another workflow uses them.
What is the first sign that engineering data is becoming reusable?
A second team can discover and consume information created for the first workflow without repeating the original extraction, interpretation and governance effort.
References
- Kevin Pilch, The AI-Ready Product Engineering Organization: How Industrial Companies Need to Rethink Product Engineering for the AI Era, Version 1.3, June 2026.
- McKinsey & Company, Software-defined hardware in the age of AI.
- McKinsey Global Institute, The economic potential of generative AI: The next productivity frontier.
- DORA Research, 2025 State of AI-assisted Software Development.