AI Did Not Make Palletisers Smarter. It Moved the Configuration Problem Elsewhere.

AI Did Not Make Palletisers Smarter. It Moved the Configuration Problem Elsewhere.

AI Did Not Make Palletisers Smarter. It Moved the Configuration Problem Elsewhere.

Leon Potgieter, robotics visual systems designer

Leon Potgieter

Leon Potgieter

Robotics Visual System Designer

Robotics Visual System Designer

Robotics Visual System Designer

Palletising is the most automated task at the end of a production line, and the one where the gap between AI marketing and shipped capability is widest. The real change is not robot intelligence. It is that the decision about how a cell should behave has moved from an integrator's laptop to a touchscreen next to the conveyor, and almost nobody has designed that screen properly.

Palletising is the most automated end-of-line task and the one where AI claims drift furthest from what ships. What actually changed, and who now has to decide.

Previous/Next Articles

A new box arrives on a Tuesday.

Here is the scene I keep coming back to, because it is the one that actually costs money.

A food packer runs a line into a palletising cell. Cases come off the sealer, a robot picks them, a pattern builds on a Euro pallet, and a wrapper takes it away. The cell has been running the same 400 by 300 case for two years. Cycle time is fine. Everyone is happy.

Then marketing changes the carton. The new one is 20mm taller, and the board is slightly weaker, so the old five-layer pattern crushes the bottom row. In the way this used to work, and still works in most plants I see, the sequence goes like this: the production supervisor calls the integrator, the integrator schedules a visit, someone drives out, someone opens the teach pendant, someone rewrites the pattern, and the approach poses, someone runs a validation, and the line runs at reduced capacity or not at all for somewhere between two days and two weeks.

Nobody in that chain is incompetent. The robot is not slow. The gripper is not wrong. The problem is that the knowledge required to change the cell's behaviour lives with a person who is not in the building.

That is the cost that matters at the end of the line. Cycle time is a solved, boring number: a palletiser that does 15 or 20 picks a minute has been available for a long time, and the marginal gain from making it faster is small compared to the cost of the line standing still. Changeover time is the number that moves a P&L. Judge every claim about AI in palletising against that number and no other.

Three different jobs wearing one phrase

"AI palletising" is three separate technologies that mature at completely different rates, and collapsing them is how buyers get sold the mature one and delivered the immature one.

The first job is pattern generation and load planning. Given a set of case dimensions, weights, crush strengths and a pallet footprint, produce a stacking pattern that is stable, within the height and overhang limits, sensible about centre of gravity, and ideally sequenced so heavy goes low.

The second job is perception. Know what is actually arriving: which case, at what pose, in what orientation, in what order, on a conveyor or in a trailer or on a mixed inbound pallet.

The third job is grasp-and-motion policy. Decide how to pick the thing and where to move it, including the awkward cases where the object is not rigid, not uniform, not where it was supposed to be, or partly occluded by its neighbour.

These three get bundled into one word in a brochure. In reality, one of them was largely solved before deep learning existed, one has genuinely been transformed, and one is not yet a production technology at all. Treating them separately is the entire point of this article.

Pattern generation is optimisation, and it was already fine.

This is the part vendors label AI most loudly, and it is the part that needs it least.

Building a stable mixed load is a constrained packing problem. You have bin packing with stability constraints, crush-strength limits per SKU, interlocking requirements, overhang rules, centre-of-gravity targets, and a sequencing constraint if the cases arrive in a fixed order. Constraint solvers and heuristic packers have done this competently for decades. The maths is not new, the objective function is well defined, and the output is deterministic, which is exactly what you want when a pallet has to survive a forklift and a truck ride.

Calling that AI is not a lie, depending on how generously you define the term. It is just not informative. If a solver produces the same pattern every time for the same inputs, the interesting property is determinism, not learning.

Learning earns its place in a narrower, more interesting sense: dynamic re-planning when reality disagrees with the plan. Cases arrive out of sequence. A SKU is short. A case is rejected upstream and the pattern now has a hole in layer three. A static solver replans from scratch and may produce a completely different stack, which is a problem when half of it is already built. A system that can produce a good next placement given a partially built and partially wrong stack, quickly enough to not break takt, is doing something genuinely useful. Photoneo describe exactly this in their mixed palletising work: pallet patterns and load balancing determined dynamically rather than fixed up front (Photoneo).

So my position on pattern generation: real capability, mostly pre-existing, modest genuine improvement from learning at the re-planning edge, and the loudest AI branding in the category. Treat the word on the slide as marketing and ask instead what happens when the sequence is wrong.

Perception is where the change is real.

If you want to see what actually changed in the last decade of palletising, look at the camera.

2D vision on a conveyor works when the world is cooperative. It fails, quietly and expensively, on glossy film, on reflective foil, on dark packaging that gives the sensor nothing to work with, and on out-of-sequence arrival where the system cannot rely on knowing what is coming next. It also struggles with the geometry of the task: you need to cover a large area, a conveyor plus a palletising zone plus possibly an infeed buffer, while still resolving edges to millimetre precision. Those two requirements pull in opposite directions (Photoneo).

Structured light 3D scanning changed the economics of that problem. Photoneo's mixed palletising setup runs at a throughput of 300 to 1,000 cases per hour, and their demonstration handled around 500 cases across roughly 20 SKUs, building stable stacks up to six feet high. The sensing wasn't one clever camera: three PhoXi 3D Scanner XL units covered the conveyor and palletising areas, a PhoXi 3D Scanner L captured fine detail, and three vision controllers behind them (Photoneo).

That hardware list is the part people skip, and it tells you what is actually going on. The gain came from sensing plus compute, with learned models doing classification and segmentation on top of good depth data. It did not come from a model that understood boxes.

Depalletising is harder than palletising, and the reason is simple: when you build a pallet you know what you put where, and when you take one apart you do not. On the build side, you have an internal model of the stack that you authored yourself. On the teardown side, you are looking at a stack somebody else built, possibly badly, possibly shifted in transit, possibly with a slip sheet or a layer of shrink film in the way, possibly mixed SKU with three case sizes in one layer. The perception problem goes from verification to discovery. Anyone quoting you the same cycle time for both is quoting you the easy one.

Grasp and motion policy is where the gap is widest.

Now the part where the marketing has run furthest ahead.

The dream is a vision-language-action model: one general policy that sees the scene, understands the instruction, and produces motion, so that a new product needs no programming at all. If that worked, changeover would go to zero. This is the capability being implied when a vendor tells you their palletiser "learns".

I recommend a case study to anyone being sold that story. Siemens ran a vision-language-action deployment on their own factory floor in Erlangen, and published it with the failure numbers intact (arXiv). The task was not exotic: pick transparent accessory bags out of cluttered bins and insert them into cardboard package cavities so the contents settle below the closing plane of the box. That last clause is the whole difficulty. It is not enough to get the bag in. The bag has to end up flat enough for the lid to close.

They adapted a pretrained π₀.₅ model with iterative fine-tuning, using LoRA and full fine-tuning, teleoperating with a Meta Quest 3 on UR7e robots with custom 3D-printed fingertips. They collected 900 episodes in a mock lab setup, then 2,535 episodes on the actual factory floor across about 10 hours, in three rounds with constraints progressively removed (arXiv).

The failure distribution is the useful part. Contents remaining on top of the product rather than settling below the closing plane accounted for 65% of failures. Multiple bags grasped at once: 23%. Bag not fully inserted: 15%: poor or failed grasps: 15%. Final success rates fell short of expectations, and the detail that should stop a room: trial 1, the one with the most constraints in place, outperformed the less constrained trials 2 and 3. The team's own conclusion included that camera placement and hardware ergonomics substantially limited what the policy could perceive and therefore learn (arXiv).

Read that as an engineer rather than as a sceptic. This is a well-resourced team, a real industrial task, thousands of real episodes, a strong base model, and the result is a system that works better when you constrain it more. That is not a failure of the research. It is an accurate picture of where the technology is.

My rating for foundation-model manipulation as a palletising production technology, today: 3 out of 10. It earns three points because the data pipeline is real, transfer from lab to floor demonstrably happens, and the failure modes are legible rather than mysterious. It does not earn more because reliability is short of production, because the system degrades as you remove constraints rather than generalising into them, and because a palletising cell needs a number like one failure in ten thousand picks, not a percentage that sounds encouraging in a paper.

Anyone selling learned manipulation as shipping palletising capability in 2026 is ahead of the evidence. Not lying, necessarily. Ahead.

The market is not buying the story yet.

Set the technical picture next to what people are actually installing.

IFR counted 542,000 industrial robot installations globally in 2024, with operational stock reaching 4,664,000 units, up 9% year on year (IFR). Market value for installations hit an all-time high of US$16.7 billion (IFR). Those are healthy headline numbers, and they get quoted constantly.

The regional detail is less comfortable. Europe installed about 85,000 units in 2024, down 8%. Germany, the largest European market at 32% of the total, installed 26,982 units, down 5%. The USA fell 9% to 34,200 (IFR). Asia carried the growth at 74% of installations, with China alone at 295,000 units.

So in the exact markets where AI-enabled automation is being marketed hardest, unit installations fell while claims accelerated. I do not read that as proof the technology is empty. I read it as buyers being sensibly conservative about capital when the promised capability is hard to verify at the point of sale. A plant manager cannot run a benchmark on a robot cell the way you can on a laptop. They buy on references and on what the integrator will contractually commit to, and almost nobody will contractually commit to a learned policy's success rate.

That scepticism is rational, and it is also a design brief. If the buyer cannot verify the capability before purchase, the interface has to let them verify it after installation, continuously, in a way that is quick to read. Trust in these systems is built on the screen, or not at all.

Certification is the ceiling, not model capability.

A hard boundary surrounds all of this that most AI-in-robotics writing ignores.

ISO 10218-1 and ISO 10218-2 were republished in March 2025, after nearly eight years of work involving experts from more than 20 countries (Automate). Several changes matter directly to what is learned.

ISO/TS 15066:2016 has been folded into the two parts and no longer stands alone as the collaborative reference. The terminology changed from "collaborative robot" and "collaborative operation" to collaborative application, for a precise reason: only an application can be validated as collaborative, never a robot in isolation. "Safety-rated monitored stop" became "monitored standstill". Part 2 shifted emphasis from the robot system to the robot application, explicitly including workpieces, task programs and supporting machinery. Cybersecurity requirements are now in scope, and functional safety requirements are explicit rather than implied (Automate, ISO).

Put that against a policy network. Validation is per-application, and it covers the task program. A behaviour that cannot be characterised cannot be validated. A behaviour that changes after validation invalidates the validation. If your cell's motion comes from a model whose output distribution you cannot bound, you don't have a certification problem you can engineer around later; you have one that sits in front of commissioning.

This is why I expect learned behaviour in production palletising cells to stay firmly inside the perception layer for a while, with deterministic, bounded, verifiable motion downstream. A vision system proposing a pose that a conventional motion planner then executes within a validated envelope is certifiable. A policy generating joint commands from pixels is a much harder conversation with a notified body. The IFR's own 2026 trends list puts safety, standards alignment, liability frameworks and cybersecurity risk from cloud-connected AI as a top-five theme, which tells you the industry knows (IFR).

The ceiling on AI in cells is regulatory and epistemic before it is technical. That is not a complaint. Palletisers move 25kg cases at speed near people.

Where this lands on a screen

Everything above lands on a screen, and that is where I have spent the last five years.

I design the HMI UX and UI for the software that runs Unchained Robotics' MalocherBot cell. It is built as three separate interfaces, and they exist because there are three genuinely different humans. There is an engineering layer, where the system itself is built. There is a configuration layer, where a cell gets defined and changed. There is an operator layer, which is what the person on the floor actually sees. The split is not organisational tidiness. It is the answer to a question that changed when cells became adaptive.

When a cell is fixed, the operator's question is "is it running?" A green light answers it. When a cell can adapt, the question becomes "why did it choose that?", and no light answers that. This is the real shift AI caused in industrial HMI, and it is much less glamorous than the demos: every capability you add to the cell becomes a decision somebody has to make, understand, or approve.

Every capability we add to a cell has to survive contact with a person in a noisy warehouse wearing gloves, under bad light, with four seconds of attention and one hand free.

Some of what that forced me to design:

Show the proposed pattern before committing. When the system generates a stacking pattern for a new case, the operator sees a 3D preview of the full stack, layer by layer, before picking a single case. Not a specification table. The stack is rotatable, with the layer under discussion highlighted and the rest ghosted. A supervisor who has been building pallets for fifteen years can look at a proposed stack for two seconds and tell you it will lean. That expertise is enormously valuable and completely inaccessible if the only representation is a parameter list.

Show confidence, but not as a number. A detection confidence of 0.87 is meaningless to a shift worker and, honestly, close to meaningless to me without knowing the distribution behind it. What is legible is a three-state treatment on the detected case outline in the live view: confident, uncertain, refusing. Uncertain means the cell slows and asks. Refusing means it stops and shows you what it is looking at. The number exists in the engineering layer, for the engineer. On the operator screen it is a colour and a shape, because the operator's job is to answer "do I trust this pick", not to calibrate a model. I have written more about where that boundary sits in how HMI design changed after AI.

Make changeover a two-minute task. This is the centre of it all. The MalocherBot Palettierer comes in four sizes, supports cobots and industrial robots, requires no coding, and can be deployed in under a day. That promise is worthless if changing a case size still requires a specialist. So the changeover flow in the configuration layer is: select the product, enter or scan the new dimensions and weight, review the generated pattern, review the gripper approach, confirm. No pendant. No code. The target I design against is that a supervisor who has never opened the configuration interface can get through it with the printed card next to the panel.

Design for the actual body in the actual room. Gloved hands mean minimum touch targets far larger than any consumer guideline, and no gestures requiring two fingers. Poor lighting and reflective panel glass mean high contrast, no thin type, and no light grey on white. Noise means never relying on an audible alert alone. A shift worker who did not attend a training course means no hidden affordances, no long-press, no meaning carried by colour alone. Vibration and dust mean the panel gets touched with the side of a hand as often as a fingertip.

Never let the operator go passive. This is the one I argue about most. If the cell handles everything silently until it cannot, the person you need in the loop has been out of it for three hours and will be slow and wrong exactly when it matters. So the operator screen surfaces what the cell is deciding even when no decision is required from the human, in a low-attention way: the next three placements, visible, always. It costs screen area, and it buys situation awareness. I go through the heuristics behind that in making automation heuristics legible.

Working across KUKA, Doosan and Nachi has made one thing obvious: the robot brand barely matters to the operator, and matters enormously to the configuration layer. The person on the floor should not be able to tell which arm is in the cell from the screen. The person configuring it needs the differences surfaced precisely: reach envelopes, payload derating with the gripper attached, singularity behaviour, how each vendor's safety-rated functions map onto the cell's monitored standstill zones. Two audiences, opposite requirements, which is the whole argument for separating the configuration layer from the operator layer rather than shipping one interface with a permissions toggle.

Four seconds

The thing I would tell anyone specifying a palletising cell in 2026: ignore the AI claims on the first page and ask three questions instead. What happens when a case arrives out of sequence? How long does a case-size change take, performed by someone already on site? And show me the screen where a person approves what the system proposed.

The capability curve in perception is genuinely steep, and mixed-SKU depalletising that was researched in 2018 is quotable work now. The capability curve in learned manipulation is real but nowhere near production reliability, and the honest published evidence says so. The capability curve in pattern generation was already high before anyone called it AI.

None of those three is the binding constraint. The binding constraint is that an adaptive cell generates decisions, and every decision has to be legible to a person standing at a panel between two other jobs in four seconds, or the adaptivity becomes a liability rather than a feature. Build the model you like. The cell only goes as fast as that.

Sources

Leon Potgieter designs HMI and visual systems for industrial robotics. He has spent the last five years as the visual systems designer for Unchained Robotics in Germany, working on the operator, configuration and engineering interfaces behind the MalocherBot cell.

Previous/Next Articles

more aloud

more aloud

Leon Potgieter

Robotics Visual System Designer

The AI Tooling I Actually Use for HMI Visuals, and the Line I Do Not Cross
A working account of the generative tooling in my HMI pipeline: what it does for cell renders, motion and variants, and what it must never be allowed near.

Leon Potgieter

Robotics Visual System Designer

The AI Tooling I Actually Use for HMI Visuals, and the Line I Do Not Cross

A working account of the generative tooling in my HMI pipeline: what it does for cell renders, motion and variants, and what it must never be allowed near.

Leon Potgieter

Robotics Visual System Designer

The Brand System Should Start at the Panel, Not at the Fair Wall
Most brand work is built for the first ninety seconds, not for five years. The interface faces the hardest constraints, so it should set the rules.

Leon Potgieter

Robotics Visual System Designer

The Brand System Should Start at the Panel, Not at the Fair Wall

Most brand work is built for the first ninety seconds, not for five years. The interface faces the hardest constraints, so it should set the rules.

Leon Potgieter

Robotics Visual System Designer

Perception Does Not Check Its Sources
Gestalt cannot become obsolete because it describes the viewer, not the screen. Which is why it is now dangerous: it will confidently group a guess with a fact.

Leon Potgieter

Robotics Visual System Designer

Perception Does Not Check Its Sources

Gestalt cannot become obsolete because it describes the viewer, not the screen. Which is why it is now dangerous: it will confidently group a guess with a fact.

Leon Potgieter

Robotics Visual System Designer

The HMI Was a Mirror. Now It Has to Be an Argument
For forty years, the industrial HMI represented machine state. AI broke that contract. The screen's job is now accountability, and that is a different design discipline.

Leon Potgieter

Robotics Visual System Designer

The HMI Was a Mirror. Now It Has to Be an Argument

For forty years, the industrial HMI represented machine state. AI broke that contract. The screen's job is now accountability, and that is a different design discipline.

Leon Potgieter

Robotics Visual System Designer

Four months to the Machinery Regulation: what HMI teams must fix
The EU Machinery Regulation applies on 20 January 2027. Several of its requirements end on a screen. Here is what HMI teams should fix now.

Leon Potgieter

Robotics Visual System Designer

Four months to the Machinery Regulation: what HMI teams must fix

The EU Machinery Regulation applies on 20 January 2027. Several of its requirements end on a screen. Here is what HMI teams should fix now.

Leon Potgieter

Robotics Visual System Designer

Slop Is a Decision Failure: Taste, Perception and Why Plausible Is Not Correct
Slop isn't an aesthetic failure; it is output made without a cost function. What taste actually is, why perception is trainable, and where AI still loses.

Leon Potgieter

Robotics Visual System Designer

Slop Is a Decision Failure: Taste, Perception and Why Plausible Is Not Correct

Slop isn't an aesthetic failure; it is output made without a cost function. What taste actually is, why perception is trainable, and where AI still loses.

Leon Potgieter

Robotics Visual System Designer

Show the Runners-Up: UX for Machines That Decide
Automation heuristics are not hard to understand, they are invisible. A working method for exposing a machine's ranked decisions at the right depth in the HMI.

Leon Potgieter

Robotics Visual System Designer

Show the Runners-Up: UX for Machines That Decide

Automation heuristics are not hard to understand, they are invisible. A working method for exposing a machine's ranked decisions at the right depth in the HMI.

Robotics Visual System Designer

Leon/Potgieter

Leon/Potgieter

Robotics Visual Systems Designer; bridging the gap between how automation technology works and how the world understands it.

©2026 Leon Potgieter - All work, all rights.

Offline

Leon Potgieter
Koringberg
Western Cape

South Africa

Leon Potgieter - Koringberg, Western Cape, South Africa