Xentara Knowledge Base

How IT/OT Convergence Gets Physical AI onto Real Machines

Industrial automation | By Bastian Kleimann, Faizan Ahmed | July 22, 2026

Physical AI covers robots and autonomous systems that perceive, decide, and act on the factory floor, and it is still early: Deloitte found that only 3% of firms have it extensively integrated into operations. Whether these projects make it to production usually depends less on the model or the robot than on how well IT and OT are connected underneath. Written as the engineering companion to "Physical AI Needs an Industrial Nervous System", this article works through the four gaps that keep IT and OT apart: protocols, data meaning, timing, and security. For each one, it covers what a convergence layer has to do and how Xentara does it. It ends with a brownfield path that starts with a single control loop.
How IT/OT Convergence Gets Physical AI onto Real Machines

Physical AI is still early, and the numbers say so. In Physical AI: The moment of acceleration, published in March 2026, Deloitte cites its own State of AI in the Enterprise survey: only 5% of firms say physical AI is transforming their organization today, while 41% expect it to within three years. Just 3% have it extensively integrated into operations, a share forecast to reach 18% within two years.

The hardware is already arriving. According to the International Federation of Robotics, factories installed more than 600,000 new industrial robots in 2025, an 11% jump, and the number of robots in operation worldwide passed 5 million. IFR expects 655,000 installations in 2026 and 806,000 in 2029. It also names easier programming and system integration as one of the things bringing deployment costs down.

Integration is where most physical AI projects will actually get stuck. The robot and the model are rarely the problem. The problem is the gap between the IT systems that hold the intelligence and the OT systems that touch the physical world. When Deloitte asked about barriers, cost came first at 41%, but talent and skills gaps (33%) and technology or data availability (31%) were close behind. In our experience, a large part of those last two comes down to IT and OT not being connected properly.

IT/OT convergence means merging information technology (the data, analytics, and enterprise software that run the business) with operational technology (the PLCs, sensors, drives, and control systems that run the plant) so the two can share data and coordinate in close to real time. It is less a single tool than the work of closing four gaps that have long kept the two apart: protocols, data meaning, timing, and security.

This article is a companion to Physical AI Needs an Industrial Nervous System, in which our Chief AI Officer Andreas Geiss makes the strategic case. The market has lined up behind the chips, the models, and the robots. The runtime that carries perception into real-time action on existing machines is the layer nobody owns yet. That paper explains why the layer matters. This one is the engineering view: what the layer has to do, gap by gap, and how to start building it in a brownfield plant. For the basics of how IT and OT differ, see IT/OT real-time and non-real-time technology.

What physical AI needs from the factory floor

Physical AI is the industry's term for machines that combine artificial intelligence with a body (robots, cobots, autonomous mobile units) and that perceive their surroundings, decide, and act without step-by-step scripting. That is a real break from decades of factory automation, in which a robot repeated one programmed motion and failed as soon as reality drifted from the script.

The difference matters because the consequences are physical. A language model that writes a bad sentence can regenerate it in milliseconds. A robot arm that misjudges a grasp drops the part, damages the tool, or hurts the person standing next to it. Physical AI has to cope with the continuous, noisy physics of real objects, and it has to do so on a clock.

That clock is the hard part. An autonomous system on the plant floor needs three things at the same time, and they have to stay in sync:

None of these sits neatly in IT or in OT. Perception is an OT concern. The model and its training pipeline belong to IT. The return path cuts straight across the boundary between them, so physical AI ends up living on the seam between the two.

Why the IT/OT divide holds physical AI back

For decades, IT and OT were built to stay apart. IT runs on elastic, best-effort infrastructure tuned for throughput and security, where a few seconds of delay rarely matter. OT controls physical processes on fieldbuses such as PROFINET, EtherCAT, and Modbus, where a PLC has to react within milliseconds and a missed deadline can mean a safety incident or a stopped line. The two grew up with different priorities, different toolchains, and different teams who rarely shared a vocabulary. Physical AI runs into every one of those differences.

Protocols come first. A real plant is a patchwork of fieldbuses; Modbus TCP and RTU, OPC UA, MQTT, and proprietary legacy interfaces, and the lines between them are blurrier than the IT/OT labels suggest. Modbus TCP runs over ordinary TCP/IP. Classic OPC UA client/server usually sits a level above the fieldbus, at the supervisory layer, and isn't a hard real-time protocol on its own. An AI system that reads a dozen of these sources and writes back to three controllers has to speak all of them, reliably and in order. Every middleware layer added to bridge one connection is another place for latency and failure to creep in.

Timing is the deepest problem, and our article on IT/OT real-time versus non-real-time covers it in detail. OT needs deterministic, low-jitter execution. Cloud-centric AI inherits IT's best-effort assumptions. Sending raw high-frequency data to a cloud server for inference costs too much in bandwidth and blows the latency budget of any closed control loop. Existing tools split along the same line. As the nervous system paper puts it, edge and IoT platforms can move and visualize data but can't close a control loop at two milliseconds, while PLC platforms can close the loop but can't host an AI model in the same runtime.

Then there's meaning. A raw Modbus register reads 40021 = 72.4. An analytics model on the IT side needs to know that this is the bearing temperature of pump 3, in Celsius, sampled at a known interval. Without a shared model of what the data means, every integration project rebuilds that context by hand, and physical AI can't generalize across machines because it never sees consistent structure.

Security is the fourth gap. OT systems were rarely designed for networked environments, and connecting them to IT and to AI services widens the attack surface across critical infrastructure. Convergence can't simply mean opening a firewall port.

The technical gaps have a human counterpart. OT engineers optimize for uptime and physical safety, IT teams for data and software. Deloitte's finding that a third of firms see talent and skills as a barrier fits that picture: physical AI needs people who understand both sides, and most organizations have kept those people in separate departments.

A better model won't get you past any of this. A generalist robot policy or a vision-language-action model is only as useful as the infrastructure that feeds it perception and carries its decisions to the machine, and that infrastructure is what IT/OT convergence provides.

What the convergence layer has to do

Convergence isn't a product you install. Whatever platform sits between your machines and your AI has to handle four jobs, and physical AI needs all four working together.

Connect everything without a middleware maze

Physical AI can't perceive what the layer can't read. The layer has to speak the interfaces a real plant runs on natively, so that old PLCs and new equipment feed one runtime. Every middleware shim bolted on to bridge a protocol is one more spot where latency and failure can creep in.

Give the data shared meaning

This is the semantics gap from above. The layer needs a semantic model that attaches context at the edge, so a register value of 72.4 arrives as the bearing temperature of pump 3, in Celsius, at a known sample rate. Without it, every integration rebuilds meaning by hand. With it, an AI model sees the same structure on every machine and can generalize across them.

Execute deterministically, on a shared clock

Perception, inference, and the control response have to run in a deterministic, low-jitter loop near the machine, not on a best-effort server somewhere across the network. The test is simple: does the AI's decision land inside the control cycle, or does it arrive too late to act on? For anything close to real-time control, a cloud round trip fails that test.

Two details are easy to miss. First, the fast control cycle and the slower inference step have to coexist in one schedule. A vision model that needs a few milliseconds per frame can't be allowed to stall a control task running every 250 microseconds. Second, every device in the loop needs the same notion of time, or the timestamps on perception data can't be trusted.

Stay secure, and leave safety where it is

Connecting OT to IT and to AI services widens the attack surface, so the layer needs fine-grained access control, encrypted connections, isolation between components, and a preference for keeping sensitive production data on-premise instead of shipping it off-site for inference. Security has to be part of the design from the start rather than a port opened later.

Safety is a separate question. Certified functional safety belongs in dedicated SIL-rated subsystems, and a convergence layer should sit alongside them rather than try to replace them. That keeps your safety case intact while the AI layer changes around it.

How Xentara implements the convergence layer

Xentara, embedded ocean's software-defined automation platform, was built to handle all four jobs in a single runtime. Here is how each one maps to a concrete, documented capability.

For connectivity, Xentara uses Skills, plugins that run inside the runtime rather than as separate software. The technical documentation covers drivers for OPC UA (client and server), Modbus TCP and RTU, EtherCAT, Siemens S7, Beckhoff ADS, Hilscher cifX, and MQTT, among others. A legacy PLC can join over Modbus while newer equipment connects over OPC UA, and both feed the same runtime without a separate piece of middleware for each connection. If a protocol is missing, the Plugin Framework lets you add your own.

For semantics, the semantic model adds structure and context to production data at the point of capture. Each value becomes an identified, typed property of a known asset instead of an anonymous register, which is the consistent structure a fleet-wide AI model needs.

For timing, the timing model supports cycle times as short as 10 microseconds (100 kHz). In a test on a single Intel Atom core controlling a motor over EtherCAT at 30% load, the maximum deviation on a 250-microsecond cycle was ±0.6 microseconds, and switching reactions were accurate to ±10 nanoseconds. Clocks across distributed systems are synchronized with the Precision Time Protocol (IEEE 1588 or TSN IEEE 802.1AS), which covers the timestamp problem. The scheduler also lets real-time and non-real-time processes run side by side in pipelines and tracks, which is how a millisecond-scale inference step and a microsecond-scale control task share one runtime. Real-Time Tasks vs. Non-Real-Time Tasks explains the mechanics.

For inference, models run locally on an engine built on PREEMPT_RT Linux, through the ONNX engine or the Torch engine, on CPU or GPU. Our article on machine learning at the industrial edge reports that vision CNNs on NVIDIA Jetson-class hardware often run under 20 milliseconds per frame. The same article describes "back channeling," where the result goes straight to a PLC or robot controller over native protocols with no cloud round trip. For simulation and virtual commissioning, the FMU driver brings FMI/FMU models into the same runtime.

For security, the security model lets you assign access rights down to a single data point, a data group, a device, or an entire bus, with inheritable user roles. Connections use TLS where the protocol allows it, and third-party protocols default to their highest available security level. Remote clients can authenticate with OAuth 2.0, certificates, or username and password. Xentara can also run in containers (see the Docker quick start guide), and data stays on-premise.

What this looks like in practice

The clearest physical AI example is an anonymized proof of concept with a European specialty-machinery OEM, described in the nervous system paper. The OEM built a vision-based quality-control station in which real-time computer vision and EtherCAT machine control run on one platform. That replaced two separately synchronized systems and cut the station's cycle time in half. It is the perceive-decide-act loop from the top of this article, running in a single runtime.

A second example shows the data side of convergence. In a brownfield retrofit, a manufacturer of precision turned components used Xentara to aggregate and preprocess shop-floor data from mixed legacy controls for its MES and ERP systems, with live dashboards and OEE analytics on top. This wasn't a physical AI project, but it built the kind of foundation one would need. Unplanned downtime fell by 30%, maintenance costs dropped by 10%, productivity rose 5% per line, and machine-specific energy use fell by 10 to 15%. The general manager described the goal this way: "we need to connect the machine level directly with business intelligence."

A practical path to converged infrastructure

The common mistake is treating convergence as a rip-and-replace megaproject. Most factories are brownfield, with PLCs, SCADA systems, and machines expected to last ten to fifteen years, so the realistic approach is additive. Leave the certified safety and motion systems where they are, and add a real-time software layer alongside them. Here is the sequence we recommend.

  1. Start with one loop, not the whole plant. Pick a single high-value use case where the latency argument is obvious, such as vision-based quality inspection or vibration-based predictive maintenance on a critical asset, and prove the local perceive-decide-act loop on one line before scaling. This is also how the IT and OT teams, who have to own the result together, start to trust each other.
  2. Model the data before you model the AI. Set up the semantic layer for the assets in that first loop. With consistent structure in place, the second and third use cases can reuse the groundwork instead of starting from zero. Skipping this step is one of the most common reasons pilots never scale.
  3. Keep inference and control in the same deterministic runtime. It's tempting to put perception on the edge and decision-making in the cloud "for now," but that split brings back the latency and reliability gap you're trying to close. The vision QC station above is a good illustration: merging two synchronized systems into one runtime is what halved the cycle time.
  4. Converge the teams as well as the systems. A unified configuration environment and open APIs let OT and IT engineers work on the same data pipeline without heavy custom development. Shared ownership is the harder part, because physical AI results depend on OT's process knowledge and IT's data discipline meeting at the same table.

Done this way, convergence stops holding physical AI back and starts making each new project cheaper, since the connectors, the semantics, and the security model are already in place.

Where this leaves the roadmap

Physical AI is usually told through robots and foundation models, because those are the parts you can see. What decides success is the converged layer underneath. It lets an autonomous system see the process at machine speed, make the decision where it needs to be made, and act on the machine within the cycle. Without IT/OT convergence, a physical AI project has intelligence but no way to put it to work on the floor.

That changes the order of questions. Before picking a robot or a model, ask whether your architecture can feed physical AI time-accurate perception and carry its decisions back to the floor deterministically and securely. If you do one thing this quarter, pick a single control loop where cloud latency clearly fails and prove convergence there. It's the smallest step that makes the rest of the roadmap real.


See what convergence is worth on your floor

Before you choose a robot or a model, put a number on the IT/OT gap for your own operation.