LLM-Generated PLC Code: Safety, Verification and Guardrails

LLM-Generated PLC Code: Safety, Verification and Guardrails

LLM PLC Code Generation: Safety, Verification and Guardrails

A language model can write a plausible Structured Text function block in four seconds. It can also write one that compiles cleanly, passes a casual read, and quietly lets a conveyor restart after an emergency stop is released. The first fact is why LLM PLC code generation is spreading through engineering tools; the second is why it cannot be treated like generating a web form.

Programmable logic controllers sit between software and physical consequence. A defect does not throw an exception that a retry fixes. It moves a motor, opens a valve, or fails to stop a press. Vendors have noticed the productivity opportunity: Siemens ships a copilot that writes SCL, Beckhoff announced an AI assistant inside TwinCAT in June 2026, and academic groups have published closed-loop systems that couple generators to model checkers. None of these claims that a model is a safety engineer.

This article gives you a reference pipeline: a sequence of verification gates that turns an unreliable generator into a source of reviewable candidates, plus a clear boundary for where the technique must not be used. You will leave with a gate design, a repair-loop pattern, a way to map the work onto IEC 61508 thinking, and a checklist for adopting it without borrowing risk you cannot see.

What this covers: the state of vendor and research tools, a four-gate reference architecture, the formal verification step with PLCverif and nuXmv, simulation and test evidence, the IEC 61508 and IEC 61131-3 constraints, failure modes, and a practical adoption checklist.

Context and Background

IEC 61131-3 is the international standard that defines the programming languages of programmable controllers: Ladder Diagram, Function Block Diagram, Structured Text, and Sequential Function Chart. Edition 4 was published in May 2025. According to the PLCopen status page it keeps the backward compatibility line with earlier editions, and summaries of the new edition from the community report that Instruction List is no longer in the main text, that UTF-8 string types were added, and that object-oriented features and a concurrency chapter were extended. Vendors may keep supporting what the standard drops, so your installed base will carry mixed dialects for years.

Structured Text (ST) is the natural target for language models. It is textual, Pascal-like, and well represented in public code compared with graphical languages. Siemens calls its flavor SCL (Structured Control Language), and CODESYS and TwinCAT implement ST with vendor extensions. Ladder Diagram, still dominant in many plants, is a poor fit for token-based generation because its meaning lives in graphical topology that must first be serialized.

The vendor landscape moved quickly between 2024 and 2026. Siemens announced a generative AI feature for its Industrial Copilot in 2024, built with Microsoft, that generates SCL and integrates it into the TIA Portal project, with general availability planned for summer 2024 according to trade coverage of the announcement. Beckhoff announced TwinCAT 3 CoAgent for Engineering (TE1700) on June 23, 2026, describing it as model-independent, including locally hosted models for air-gapped operation, with engineers reviewing and approving every AI-generated result before it is applied. Beckhoff also reports productivity gains of 20 to 30 percent in its own tests, particularly onboarding to existing code, documentation and maintenance. That figure is a vendor-reported result and the press release does not publish the method behind it.

Research has been working on the harder question of correctness. LLM4PLC, published at ICSE SEIP 2024, proposed an iterative pipeline combining a model, grammar checking, a compiler and the nuXmv model checker. Agents4PLC followed in late 2024 with five cooperating agents. Both are covered in detail below. For readers choosing between control languages first, our comparison of IEC 61499 and IEC 61131-3 for distributed control sets the language context this article assumes.

The central thesis is simple. The model is not the product; the verification chain around it is. Treat the generator as an untrusted proposer, as you would treat a contractor you have never audited, and invest your engineering effort in the checkers, the specifications, and the evidence trail.

A Reference Architecture for Verified Generation

The pattern that survives scrutiny is a proposer and a stack of independent checkers. The LLM drafts Structured Text from a requirement; each gate (compile, static rules, model checking, simulation) can reject the draft with machine-readable evidence; failures feed a capped repair loop; only a draft that passes every gate reaches a human reviewer, who signs the release. The model never gets to approve itself.

LLM PLC code generation reference pipeline with four verification gates

Figure 1: Reference pipeline for LLM PLC code generation. Dashed edges are repair feedback to the generator; the human review gate is not optional.

Figure 1 shows the flow. Requirements and the I/O list enter on the left of the chain, a formalization step turns them into properties and tests, the generator drafts code with vendor documentation retrieval, and four automated gates run in sequence. Dashed arrows return diagnostics, findings, counterexamples and failed traces to the generator. The last two boxes are human review and a release record in version control. Everything above the human step is automation that produces evidence; nothing above it produces authority.

Formalize before you generate

The single most valuable step in the pipeline happens before any code exists. If the requirement is a paragraph of prose, the model and the reviewer each interpret it differently, and verification has nothing to check against. Convert the requirement into three artifacts: a signal list with types, ranges and units; a set of properties in a form a model checker can read; and a set of scenario tests with expected outputs.

Properties for control code have recognizable shapes. Safety properties say something bad never happens: the motor output is never true while the guard-closed input is false. Reachability properties say something useful remains possible: the cycle can always return to the idle state. Timing properties bound responses: after a stop command the output drops within a stated number of scan cycles. CERN’s PLCverif, discussed later, offers requirement patterns that capture exactly these categories so that engineers need not write temporal logic by hand.

An important failure mode lives here. The LLM4PLC paper has the model translate constraints into formal specifications as part of the pipeline, and the Agents4PLC authors point out that in earlier work the correctness of specifications is verified only at the design level, while the correctness of the generated code remains questionable. If the same model writes both the specification and the code, a shared misunderstanding passes every gate. Properties should be authored or at least reviewed by someone who owns the hazard analysis.

The generator and its context

The generator needs more than a prompt. Vendor ST dialects differ in how they declare function blocks, handle strings, call timers, and name system functions. Retrieval of vendor documentation, library signatures and in-house code patterns sharply reduces hallucinated function names. AutoPLC, an arXiv framework aimed at vendor-aware ST generation, builds exactly this: an API library indexed by functional summaries plus a case database of proven implementations, with repair driven by real vendor compiler feedback. Its authors report compile pass rates above 90 percent on their own 914-task benchmark, though an independent summary notes the case database and benchmark come from the same source libraries and that removing case retrieval drops the pass rate by 35 points. Read this as evidence that retrieval matters a great deal, and as a reminder that compile rate is a low bar.

Context also includes your coding standard. A model told to use a house template, fixed naming, no pointer arithmetic, and explicit state enumerations produces code that is easier to verify. Constrain the output shape and the checkers have less to fight.

Why the gates must be independent

Each gate catches a different class of fault, and they must not share assumptions. A compiler proves the code is well-formed, not that it does the right thing. Static rules prove structural hygiene. A model checker explores the state space against properties you wrote. A simulator exercises timing and I/O behavior over time. If one tool failure can blind two gates, you have fewer gates than you think. Using a different tool family at each step, and a different model for review aids than for generation, is cheap insurance.

Verification gate sequence for IEC 61131-3 Structured Text AI output

Figure 2: Gate sequence for IEC 61131-3 Structured Text AI output, with a capped repair loop on any failure.

Figure 2 zooms into the gate chain. A failure at any gate sends a structured diagnostic to the repair loop and the loop re-enters at the top, because a fix for a counterexample may break syntax or style. The retry cap matters: an unbounded loop will eventually produce something that passes by gaming the checks, and it burns tokens while doing so.

Deeper Analysis: What the Verification Evidence Shows

The strongest argument for verified generation is also the most sobering: in the published research, the verification stage is where most generated programs fail. That is a feature. It means the gates are doing work that a human skim would miss, and it sets honest expectations for throughput.

What the research actually reports

LLM4PLC, by Fakih and colleagues, runs a pipeline of model-based design, ST generation with one-shot prompts and optional LoRA fine-tuning, syntax checking with the open-source MATIEC compiler, a feedback loop of compiler errors, SMV model generation, and verification with nuXmv, with an optional human in the loop. The authors evaluated GPT-3.5, GPT-4, and base and fine-tuned Code Llama models at 7B and 34B parameters. The abstract reports that the generation success rate rose from 47 percent to 72 percent, and that a survey-of-experts code quality score rose from 2.25 to 7.75 out of 10. The code dataset came from the OSCAT library, and a Fischertechnik manufacturing testbed was used for validation. The test set was small, forty samples, and the quality score depends on expert panels, so treat the figures as an indication of direction rather than a forecast.

Agents4PLC, by Liu and colleagues, replaces the single loop with five agents: retrieval, planning, coding, debugging and validation. It uses MATIEC and RuSTy for compilation and nuXmv and PLCverif for functional verification, where PLCverif translates ST into SMV or CBMC for model checking. The authors built a benchmark of 23 programming tasks with formal specifications: a 16-task easy set carrying 58 properties and a 7-task medium set carrying 43 properties. With GPT-4o, the paper’s table reports that Agents4PLC compiled 16 of 16 easy tasks and 7 of 7 medium tasks, but the verifiable rate, meaning the program also satisfies the formal properties, was 11 of 16 (68.8 percent) on easy tasks and 3 of 7 (42.9 percent) on medium tasks. The same table gives LLM4PLC with GPT-4o a verifiable rate of 2 of 16 and 0 of 7.

Read those numbers carefully. A compile rate of 100 percent coexists with a verifiable rate of 42.9 percent on the harder tier. Compiling tells you almost nothing about behavior. The benchmark is also tiny, the tasks are textbook-scale, and nothing in these papers speaks to the thousand-line, hardware-coupled programs that real plants run. They show that a closed loop helps, not that the problem is solved.

Model checking with PLCverif and nuXmv

PLCverif is CERN’s model checking platform for PLC programs. According to its ICALEPCS 2021 paper, it was first released internally in 2019 and has been available as open source since September 2020; it supports Siemens languages and a CBMC backend alongside symbolic model checkers. It reads the PLC program, builds a formal model of one scan cycle with nondeterministic inputs, and checks requirements expressed through patterns such as “always, if the condition holds then the assertion holds”.

The scan-cycle model is the key idea. A PLC executes a cycle of reading inputs, running logic and writing outputs. The verifier treats each cycle as a transition and lets the inputs take any value the type allows, so it explores situations a bench test would never reach: a guard input flickering, a sensor stuck, two commands arriving together. This is why model checking finds the stop-and-restart class of bug that scenario tests miss.

Be realistic about limits. State explosion is the standing problem; the Agents4PLC authors name it explicitly. Programs with wide integers, long timers, or many interacting state machines can exhaust a model checker, and abstraction is then your job. A 2026 arXiv preprint, ESBMC-PLC+, proposes a successor route: it accepts Ladder Diagram, graphical function block diagrams and ST/SCL through the MATIEC pipeline into the ESBMC backend using k-induction, and reports large speedups over nuXmv on integer-timer programs in its own 18-program test set. It is a preprint with a small benchmark and no connection to language models, so treat it as a promising tool to evaluate rather than a replacement for PLCverif.

A worked example: where a gate earns its keep

Consider a conveyor start/stop block with a guard and an emergency stop. A model asked for “start the conveyor when the start button is pressed, stop on stop or e-stop, do not restart after e-stop until reset” might produce the following.

FUNCTION_BLOCK FB_Conveyor
VAR_INPUT
  Start    : BOOL;
  Stop     : BOOL;
  EStopOk  : BOOL;  (* TRUE when the e-stop circuit is healthy *)
  Reset    : BOOL;
END_VAR
VAR_OUTPUT
  Motor    : BOOL;
END_VAR
VAR
  Latched  : BOOL;
END_VAR

IF NOT EStopOk THEN
  Latched := FALSE;
END_IF;
IF Start AND EStopOk AND NOT Stop THEN
  Latched := TRUE;
END_IF;
IF Stop THEN
  Latched := FALSE;
END_IF;
Motor := Latched AND EStopOk;

It compiles. It reads well. It also violates the stated requirement: when the e-stop is released, a still-held Start input latches the motor again immediately, with no Reset needed. The property “after an e-stop, Motor stays false until Reset is true” fails, and the model checker returns a counterexample trace of three cycles: e-stop open, e-stop restored with Start held, Motor true. The repair loop hands that trace to the generator, which adds an explicit Tripped state cleared only by Reset. A human reviewer skimming the first version might have caught it; the gate catches it every time, at the same cost, on the fortieth block of the day.

This example is deliberately small, and it also shows the boundary of the method. The e-stop here is a standard input in a standard controller. A real emergency stop function is a safety function with its own certified hardware path; the standard-controller logic above must never be the thing keeping people safe. That distinction is the subject of the next major section.

Sequence of agent, compiler, model checker and simulator interactions in LLM PLC code generation

Figure 3: Interaction sequence of an agentic LLM PLC code generation loop: compiler diagnostics, a model checker counterexample, then simulation, then an evidence pack for the engineer.

Figure 3 traces the loop as a conversation. The agent drafts, the compiler returns diagnostics, the agent revises, the model checker returns a counterexample, the agent fixes it, the simulator runs scenario tests, and finally the engineer receives the candidate together with an evidence pack: the properties checked, the tool versions, the traces seen and fixed, and the simulation logs. That pack is what makes review fast and what lets an auditor reconstruct what happened.

Simulation and test evidence

Model checking proves properties of an abstraction; simulation exercises the code against a plant behavior. Both are needed. Siemens PLCSIM, CODESYS and TwinCAT virtual runtimes, and open runtimes such as OpenPLC let you run the compiled logic without hardware. A digital twin of the cell, even a crude one with timed sensors and actuators, turns scenario tests into closed-loop tests: the conveyor model reports a jam, the logic responds, the check confirms the response time.

Three practices make simulation evidence trustworthy for generated code. First, write the scenario tests before generation and freeze them, so the model cannot tailor tests to its output. Second, include fault injection: stuck inputs, dropped signals, out-of-range analog values. Third, run the same suite against the existing hand-written version where one exists, and require identical behavior or documented differences. A differential run against known-good code is the cheapest oracle you have.

Test coverage should be measured, not assumed. Branch coverage of the generated blocks from the scenario suite tells you whether the tests reach the code the model wrote; uncovered branches in generated code are exactly where plausible-but-wrong logic hides.

Static rules as a cheap, powerful gate

Between the compiler and the model checker sits an inexpensive gate that is often overlooked. Static rules encode your coding standard in machine-checkable form: no recursion, no dynamic memory, every loop has a provable bound, no writes to an output from more than one place, every enumerated state has a handler, no implicit type narrowing, no use of vendor system calls outside an allowed list. The PLCopen coding guidelines, vendor linters and general static analysis tools cover much of this ground, and your own rules cover the rest.

These rules matter more for generated code than for human code. A person has habits and a model has distributions; it will occasionally reach for a construct that is legal but unwise, such as a WHILE loop polling an input inside one scan cycle, which can starve the cycle and trip the watchdog. A rule that forbids it costs nothing and removes the class of fault. Rules also shrink the verification problem: restricting the language subset the model may emit reduces the state space the checker must explore.

One useful extension is a diff gate for modifications. When the model edits existing code rather than writing new code, require a minimal diff, forbid changes outside the targeted function block, and re-run the full verification suite on everything the change can reach. Unrestricted rewrites of working code are where regressions are born, and the diff gate makes them visible.

Choosing the repair-loop policy

The repair loop is where cost and quality trade. Three policy decisions shape it. The first is the retry cap: a small number, such as three to five attempts per gate, is usually enough to separate easy repairs from conceptual misunderstandings. A draft that needs more attempts than the cap should be escalated to a human, not forced through. The second is what feedback the generator sees. Raw compiler messages are useful; a model checker counterexample translated into a plain statement of the failing cycle sequence is far more useful than a raw trace file. The third is whether the repaired draft restarts the gate chain from the top. It should, since a fix for one property can disturb another.

Log every iteration. The iteration history is diagnostic gold: a requirement that consistently takes six loops to satisfy is probably ambiguous, and a recurring counterexample pattern points at a missing rule in your template or prompt. Over time the loop teaches you where your specifications are weak, which is valuable independent of the model.

IEC 61508 and the Safety Boundary

Functional safety standards do not ban a tool; they demand evidence and control of the tool’s influence. IEC 61508 is the generic functional safety standard for electrical, electronic and programmable electronic safety-related systems, with Part 3 covering software and Part 4 the definitions. IEC 61511 applies similar thinking to the process industry. Neither document mentions language models, so everything below is interpretation, and any real project must confirm it with its functional safety assessor.

Where an LLM sits in the tool classification

IEC 61508 classifies software tools by their potential to affect the safety-related software. As summarized in a certification-body presentation, a T1 tool generates no output that contributes to executable code, such as a text editor; a T2 tool supports testing or verification and may fail to detect errors, such as static analyzers and coverage tools; and a T3 tool generates output that directly or indirectly contributes to executable code, for which evidence is needed that the tool conforms to its specification and that failures attributable to the tool are controlled.

An LLM that writes code which is then loaded into a controller is, by that definition, a T3-style tool in the worst reading. This is the heart of the problem. Qualification of a T3 tool usually leans on a documented specification, a validation suite, and a stable, versioned behavior. A hosted language model has none of these in the conventional sense: it has no specification of what it will output, it is non-deterministic by default, and the provider may change it without notice. Compare a certified compiler, which has a precise input language, deterministic output and a validation history.

The practical consequence is a design rule, not a legal conclusion: do not rely on the model’s output being correct, and do not claim any credit for the model in the safety argument. Instead, make every output pass through independent, deterministic checks and human review so that the tool’s failures are detected downstream. In tool-classification language, you are trying to make the risk the model introduces detectable by other means, which is the same argument that justifies using a non-qualified compiler with extensive output verification. Whether an assessor accepts that argument at a given integrity level is a judgment for the assessor.

What IEC 61508 software requirements imply

Without reciting clause numbers, the software lifecycle in IEC 61508-3 expects a defined development process, a software safety requirements specification, architecture and design stages, module testing, integration testing, verification at each phase, traceability from requirement to test, configuration management and controlled modification. The rigor of the recommended techniques increases with Safety Integrity Level, SIL 1 to SIL 4. These expectations are about the process and its evidence, which suggests the right place for a model: inside the process as a drafting aid whose output enters at the same gates as any human-written code.

Two consequences stand out. First, traceability: every generated block must link to the requirement it satisfies and to the verification evidence, which the evidence pack in the pipeline provides. Second, coding practice: standards and guidance in this area favor restricted, analyzable language subsets. IEC 61508 and IEC 61511 distinguish limited variability languages from full variability languages, and a common reading treats Ladder Diagram and Function Block Diagram as limited variability while general-purpose text languages are closer to full variability. That is an interpretation worth confirming for your toolchain, but the direction is clear: the narrower the language subset, the easier verification becomes, and the better the case for constraining the model’s output to a small template.

IEC 61131-3 itself defines the language semantics and does not claim functional safety. PLCopen publishes safety function blocks for use with certified safety controllers, and IEC 61131-6 addresses functional safety for programmable controllers. A certified safety PLC offers its own constrained programming environment, and a vendor certificate applies to that environment, not to code arriving from elsewhere.

BPCS versus safety instrumented function

The boundary that matters most is architectural. Industrial plants separate the Basic Process Control System (BPCS) from the Safety Instrumented System (SIS). The BPCS runs routine control; the SIS performs independent safety functions with its own sensors, logic solver and final elements, certified to a target SIL.

Decision boundary for LLM PLC code generation across safety functions and basic control

Figure 4: Where an LLM may sit relative to a safety function. Anything that can reach a safety instrumented function stays assistive only.

Figure 4 turns this into a decision. If an output could reach a safety function, restrict the model to drafting tests, documentation and review notes, with no code path into the safety logic solver. If the output reaches only the running process in the BPCS, use the full verification chain with human sign-off and change control. If it reaches neither, such as code explanation or documentation, usage is low risk. A model may help a person understand safety logic; it must not author it.

This position is conservative on purpose. The SIL concept exists to bound the probability of dangerous failure, and a stochastic generator has no quantified failure rate. You cannot compute an average probability of dangerous failure on demand for a model’s drafting errors, so you cannot place it inside the quantified part of a safety argument. You can place it in front of a process that does the quantification.

Cybersecurity is a second boundary

A language model that reads engineering documents and writes controller code is also an attack surface. Prompt injection through a spreadsheet of requirements, a library comment, or a retrieved manual page can steer output toward unsafe behavior or leak project data. Our analysis of agentic AI security and prompt injection lays out the attack classes; the relevant rule here is that retrieved content is data, never instructions, and that the generation agent should have no credentials to a live controller. The pipeline ends at a release record in version control, and a separate, human-controlled deployment path loads the controller.

Tool access deserves the same discipline. Agents increasingly reach controllers and data through protocol servers; the guidance in our piece on OPC UA companion specifications as tools for AI agents describes how to expose read-oriented, scoped capabilities. For code generation, the equivalent is read access to documentation and libraries, write access to a scratch repository, and nothing else.

Trade-offs, Gotchas, and What Goes Wrong

Verified generation has real costs and sharp edges. Most of the failures practitioners report are quiet.

Specification debt moves, it does not vanish. The pipeline demands precise properties and tests. For well-understood patterns (motor starters, interlocks, sequencers) that precision exists in the plant’s design documents. For novel process logic it does not, and writing the properties may take as long as writing the code. If you cannot state the property, the pipeline cannot protect you, and the model’s fluent output will hide the gap.

Verification inherits the abstraction. A model checker proves that the model of the program satisfies the property. Whether the model reflects real hardware behavior, such as scan-time jitter, I/O update ordering, retentive memory, and vendor-specific runtime quirks, depends on how the tool was built. Verified logic can still misbehave if the plant sensor is slower than the abstraction assumed.

Passing the gates can become the goal. A generator inside a repair loop optimizes for gate success. It may weaken the logic until a property is vacuously true, for example by making an output unreachable. Guard against this with coverage and reachability checks on the properties themselves, and with a diff-based review that flags removed behavior.

Non-determinism breaks reproducibility. The same prompt can yield different code tomorrow, and a hosted model may change under you. Pin the model version where the provider allows it, record prompts, seeds where available, and tool versions, and treat the committed code, not the prompt, as the artifact of record. Locally hosted models, which TwinCAT CoAgent explicitly allows for air-gapped use, give you version control over the generator at the price of operating it.

Vendor dialect drift. Code that compiles under one runtime may use a function another lacks. Retrieval from the actual target’s library helps, and the compile gate must use the real vendor compiler, not only an open-source stand-in. MATIEC is excellent for IEC 61131-3 conformance and says nothing about a vendor’s extensions.

Reviewer complacency. Humans over-trust fluent, tested-looking code. A reviewer who sees a green evidence pack may skim. Counter this by having reviewers verify requirements traceability and the properties first, the code second, and by periodically seeding known-bad candidates to measure whether review still catches them.

Data and IP exposure. Prompts carry plant structure, tag names and proprietary logic. Siemens states that its copilot offers a private Azure OpenAI instance that does not use customer data for model retraining. Whatever you use, confirm retention terms and consider local models for sensitive process IP.

Anti-pattern: generating and deploying in one step. Any workflow in which model output flows to a controller without independent verification and a human decision is an incident waiting for a date.

Practical Recommendations

Start where the risk is lowest and the evidence is best: explanation, documentation, test drafting, and code review aids. These deliver real time savings, consistent with Beckhoff’s own framing around onboarding, documentation and maintenance, without placing the model in the control path. Move to drafting BPCS logic only once the gate chain exists and has been exercised on known-good and known-bad examples.

Build the pipeline from boring parts. Use the real vendor compiler plus MATIEC as a conformance check, a static rule engine for your coding standard, a model checker such as PLCverif for properties on the patterns it supports, and your vendor simulator for scenario tests. Keep the model replaceable. The gates are the durable investment; models will change every few months.

Define the scope in writing. Name the block types the generator may produce, such as interlocks and sequencers, and name those it may not. Anything touching a safety instrumented function is out of scope. Get your functional safety assessor to review the policy before you scale it.

A short adoption checklist:

  • Classify every target program as safety function, BPCS, or non-control, and apply the Figure 4 boundary.
  • Write properties and scenario tests before generation and freeze them.
  • Use independent gates: vendor compile, static rules, model checking, simulation.
  • Cap repair retries and escalate to a human on exhaustion.
  • Produce an evidence pack per block: properties, tool versions, traces, logs.
  • Pin and record model and tool versions; commit code, not prompts, as the artifact.
  • Give the agent no credentials to live controllers; deploy through a separate human path.
  • Audit prompt and retrieval inputs for injection, and review data-retention terms.
  • Measure reviewer catch rate with seeded defects.
  • Track iterations per requirement and use the data to fix weak specifications.

If you also run vision models on the shop floor, the same discipline of staged validation applies; see our notes on defect detection with edge AI for how validation sets and drift monitoring work in that adjacent domain.

Frequently Asked Questions

Can an LLM write IEC 61131-3 Structured Text that is safe to run?

It can write Structured Text that compiles and often satisfies simple properties, but safe to run is a property of the whole process, not the generator. In published research the verifiable rate sits well below the compile rate, and the test sets are small. Treat output as an untrusted draft, run it through compilation, static rules, model checking and simulation, and have a qualified engineer approve it before any deployment.

Is LLM-generated code allowed under IEC 61508?

The standard does not prohibit it and does not address it. It requires evidence that software tools do not introduce undetected faults, and that the software lifecycle, verification and traceability are followed at the target SIL. A model offers no specification or fixed behavior to qualify, so the defensible approach is independent verification of every output. For safety instrumented functions, keep the model assistive only and ask your assessor to confirm.

What is PLCverif and how does it help with AI-generated PLC code?

PLCverif is CERN’s open-source model checking platform for PLC programs, available openly since September 2020. It translates PLC code into a formal model and checks requirements against it, returning a counterexample trace when a property fails. In an AI pipeline it acts as an independent gate that finds logic errors a compiler cannot, and its traces make useful repair feedback for the generator.

Which PLC languages work best for AI code generation?

Structured Text works best, since it is textual and well represented in public code, and Siemens SCL is the target of Siemens’ copilot. Ladder Diagram and Function Block Diagram carry meaning in graphical layout, so generation requires a serialization format and is less mature. Verification tooling also leans toward ST. If you need graphical output, generate and verify ST, then review any translation separately.

Do PLC AI copilots send my code to the cloud?

It depends on the product and configuration. Siemens describes optional access to a private Azure OpenAI Service instance that does not use customer data for retraining. Beckhoff states that TwinCAT CoAgent works with providers such as OpenAI or Anthropic and can use locally hosted models for fully air-gapped operation. Check retention, residency and contract terms for each tool, and prefer local models for sensitive process logic.

How should I benchmark LLM PLC code generation for my plant?

Build a private benchmark from your own requirements, with frozen properties and scenario tests, and measure compile rate, verifiable rate and simulation pass rate separately. Public benchmarks such as the 23-task set in Agents4PLC are small and textbook-scale. Report iterations per task and reviewer catch rate on seeded defects, and re-run the suite whenever the model version, prompt template or toolchain changes.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *