← BlogEngineering knowledge

Why a digital twin starts with the P&ID drawing

Most plants keep their P&IDs as PDFs or scans, out of step with their instrument lists. Why that blocks digitization, and how AI helps.

Every process plant has thousands of P&ID sheets: piping and instrumentation diagrams that show where every pipe, valve, pump and instrument is and what it connects to. In most plants these drawings are PDFs or scanned paper, and the instrument index is kept in a separate Excel file. After years of maintenance and revamps the two drift apart: a tag on the drawing is missing from the list, or the other way round. Anyone who has worked on a revamp knows these mismatches.

The cost of information nobody can find

The problem is not new, and it is not cheap. A 2004 study by the US National Institute of Standards and Technology (NIST) estimated that poor information interoperability cost the US capital-facilities industry at least USD 15.8 billion a year; about two thirds fell on owners and operators, and 57% arose in operations and maintenance, not construction. A 2010 study that recorded the work of 78 design engineers hour by hour found they spent 55.75% of their time finding, reading and exchanging information.

No digital twin without readable drawings

A digital twin, a living software model of the plant that links real data to the structure of its equipment, needs structured data: which tag is which instrument, on which line, next to which equipment. The Industrial Digital Twin Association (IDTA) states plainly in its 2023 standard that the lack of machine-readable P&IDs can be a showstopper for digital use cases.

The industry has built standards for this. DEXPI 2.0, for exchanging P&ID data, was published in October 2025, and CFIHOS 2.0, for handing over equipment information, in November 2025. Both assume the data is already structured. For a plant that has run for decades, the first step is converting those PDFs.

What AI can do now

Machine vision has made real progress in reading engineering drawings. In a study published in 2022 in the Journal of Computational Design and Engineering, a deep-learning system recognized symbols on real P&IDs with 96.65% precision and text with 90.65% precision. Those numbers mean the machine does most of the work, but not all of it. The right approach is for the machine to take the volume and an engineer to check the doubtful cases, never to accept the output unreviewed.

A practical path

  1. Start with the tags. Instrument tags are read from the drawings and extracted: the most valuable and most checkable part of the data.
  2. Compare with the list. The extracted tags are matched against the instrument index in Excel.
  3. Report the mismatches. Every tag that is in one and not the other, or written differently, is flagged in a report.
  4. An engineer reviews. The engineering team checks only the mismatches, not thousands of tags.
  5. Then build the structure. The cleaned data becomes the basis for a structured model, standards such as DEXPI, and eventually a digital twin.

Vakav's view

Vakav's document-vision product starts this path at step one. It detects instrument tags in PDF drawings with OCR, matches them automatically against Excel lists, reports every mismatch in detail, and delivers an annotated PDF and an Excel report ready for the engineering team to review. Processing runs on graphics hardware inside the organization, and no drawing ever leaves it. The product's final name and hardware specifications will be announced soon.

If you have a revamp, documentation or digitization project ahead, see the product page or talk to us.