Let AI read your operator interviews. Never let it be the operator.

Let AI read your operator interviews. Never let it be the operator.

Let AI read your operator interviews. Never let it be the operator.

Leon Potgieter, robotics visual systems designer

Leon Potgieter

Leon Potgieter

Robotics Visual system Designer

Robotics Visual system Designer

Robotics Visual system Designer

In 2026, LLMs matched human analysts at coding interview data and still failed at inventing it. That split should decide where AI sits in HMI research.

In 2026, LLMs matched human analysts at coding interview data and still failed at inventing it. That split should decide where AI sits in HMI research.

Previous/Next Articles

Twice in the last few weeks I have sat through a demo where a "synthetic operator" answered questions about a robot cell it had never stood next to. In industrial B2B, the pitch lands, because recruiting six real machine operators, on shift, across three plants and two countries, is genuinely hard. A plant manager has to sign off. Somebody has to cover the line. The operator you need most is always the busiest. Against that, a persona that answers in four seconds and never cancels looks like a bargain.

I think it is the worst bargain on offer to an HMI team right now, and 2026 gave me numbers to argue it with rather than taste. They split cleanly in two directions.

The 2026 evidence draws a line, and the line is not where vendors put it

In April 2026, PLOS Digital Health published a blinded comparison of large language models against human analysts doing thematic analysis. Two general-purpose models and one specialised tool coded a focus-group transcript, and blinded human coders and an expert consensus panel judged the results.

On deductive coding, the models matched the humans. Agreement was 93.5 percent for the LLMs (95% CI 92.5 to 94.5) against 92.7 percent for the humans (95% CI 91.6 to 93.9), with an identical kappa of 0.34 and near-identical Gwet's AC1 scores of 0.93 and 0.92. All three models cleared non-inferiority against the human analysts (p < 0.0001), and two of them, ChatGPT-5 and Claude 4 Sonnet, reached statistical superiority on the deductive task. On inductive theme generation, the picture was thinner: only ChatGPT-5 cleared non-inferiority, and then only at p = 0.043.

The error figures matter as much as the agreement figures. Strict hallucination was low at 1.2 percent, but the comprehensive error rate was 12.4 percent, and the authors are explicit that human verification remains necessary. They are equally explicit about their study's limits: one transcript, on a topic sitting squarely inside the training data, judged against panel consensus rather than objective ground truth.

One transcript in one domain does not license a blanket claim, and the authors say so. But applying an existing codebook across forty operator interviews is the same class of task the models just performed at human parity, at a speed no human team will match, provided somebody verifies the output line by line and expects roughly one item in eight to carry an error. That is a real capability, and it is worth money.

Now look at what happens when you ask a model to produce the data instead of read it.

Figure 1: LLMs matched human coders on deductive thematic coding, but roughly one output in eight carried an error. Source: Hill et al., PLOS Digital Health, April 2026.

Synthetic participants fail in exactly the place industrial research lives

Jim Lewis and Jeff Sauro published a review of the synthetic-user experiments in April 2026. Their tally was 9 encouraging findings against 14 discouraging ones. The interesting part is not the headline count. It is which categories scored zero encouraging findings. Matching expected variance came out at zero encouraging findings against three discouraging. Good qualitative depth came out at zero against three. Two further categories, representativeness and matching regression weights, also scored zero.

The underlying studies are blunt. Park and colleagues attempted 14 classic social science studies with GPT-3.5 and got six sets of unanalysable data, five failed replications and three successes. Shrestha and colleagues put 43 policy questions to GPT-4 across three countries and found about 70 percent of the synthetic answers significantly different from human ones, even though the means looked reasonably aligned. Almeida and colleagues tested four models across eight psychology studies of legal and moral reasoning, found GPT-4 the closest match to human responses, and still reported systematic differences, with a tendency for models to exaggerate effects.

The clearest single result is older and comes from political science. In Political Analysis (2024), Bisbee and colleagues had ChatGPT adopt demographic personas and rate 11 sociopolitical groups, then compared the output against American National Election Study data. Averages tracked the benchmark closely. Everything below the average fell apart. Standard deviations in the synthetic estimates were far smaller than in the real survey. Forty-eight percent of regression coefficients were statistically different from the ANES estimates, and among those, 32 percent carried the wrong sign. Power analyses based on the synthetic data underestimated the sample size you would actually need by nearly an order of magnitude. The responses also shifted substantially over three months as the model was updated.

Take those findings together and the failure mode is specific. Synthetic participants give you a plausible centre and a collapsed distribution.

Figure 2: The generation side of the ledger. Left: encouraging against discouraging findings in the 2026 review, overall and in the two categories that scored zero against three discouraging findings each. Right: share of synthetic survey regression coefficients that differed significantly from the real survey benchmark, and the share of those differing coefficients that carried the wrong sign. Sources: Lewis and Sauro, MeasuringU, April 2026; Bisbee et al., Political Analysis, 2024.

In HMI, the collapsed distribution is the part you needed

A plausible centre is near worthless in interface work for robot cells. The average operator does not break your error-recovery flow. The outlier does.

The outlier is the setter in cut-resistant gloves who cannot hit a small target. It is the night-shift operator who was never trained on the cell and learned it from a colleague in a few minutes. It is the person who has retaught the same pick point so often that they have stopped reading your confirmation dialogue. It is the technician who reaches your fault screen already angry, line down, supervisor behind him.

Every one of those is a tail case, and in my experience every one of them produces a design change. A model that reliably reproduces the group mean and reliably flattens the spread is optimised to hide precisely the people whose behaviour you need to see. Worse, it hides them while producing output that reads like research, complete with quotes, which is the failure mode that survives a review meeting.

The Nielsen Norman Group named a second problem in 2024, and I haven't seen it solved since. Models want to please. Rosala and Moran call solution validation with synthetic users "incredibly risky" because "AI loves to please, every idea is often seen as a good one." Their synthetic users "seem to care about everything" instead of ranking needs. In a discipline whose entire job is deciding which of nine competing demands gets the top of the screen, a participant who values everything equally is not a participant. It is noise with good grammar.

Context is the finding, and it does not exist in the model

ISO 9241-210:2019 exists because human-centred design is not a survey exercise. It sets requirements and recommendations for applying human-centred design across the lifecycle of computer-based interactive systems, and Annex B provides a checklist to support claims of conformance. My reading is that it requires you to produce an account of how you actually understood the environment your interface runs in.

A language model can't recover that environment. The ambient noise level makes your audio alarm decorative. It is a panel mounted for the tallest person on the commissioning team and operated by the shortest person on the shift. In my experience, the laminated cheat sheet taped beside the screen is the most valuable artefact on any HMI field visit because it records every place the interface failed. Shift handover lands at the same moment as changeover, so your setup wizard gets used by the most distracted person in the building.

The part no model has access to. AI-generated illustration.

Shivani Kapania and colleagues put this to 19 qualitative researchers in work published at CHI 2025. The researchers started sceptical, were surprised at how similar the narratives looked when the model was given the interview probe, then over several turns identified what was missing: responses lacking in "palpability and contextual depth", the foreclosure of participants' consent and agency, and a real risk of delegitimising qualitative methods altogether. They call the substitution the "surrogate effect". That phrase is the right one. You are not saving research. You are producing a surrogate for it and then making decisions as though you had done it.

Where AI has actually earned a place in my process

I am not arguing for less AI in HMI research. I am arguing for putting it on the correct side of the line.

It belongs on the analysis side, hard. Codebook application across a large interview set, now with published evidence behind it. First-pass clustering of free-text service tickets. Triaging telemetry into candidate problem areas before a human decides which are real. Stress-testing an interview guide before you burn a plant visit on it.

It also belongs in preparation, the one the Nielsen Norman Group endorsed: synthesising available material about a user group into something digestible, learning a new domain, piloting a guide, and generating hypotheses you then test with actual people.

It does not get to be the person. Not for concept validation, not for prioritisation, not for a persona printed and pinned to a wall for two years.

Figure 3: The band I put each activity in, and the four questions I run before letting a model stand in for anybody.

A substitution test you can run in a meeting

When somebody proposes replacing a research activity with a model, four questions settle it.

  1. Am I reading data or creating it? Reading is supported by evidence. Creating is not. If the answer is creating, stop here.

  2. Does the decision depend on the spread or the average? Interface decisions almost always depend on the spread, and synthetic output has a documented variance problem in both the survey work and the review above. Anything hinging on outliers, edge cases or minority workflows is out of scope.

  3. Is the finding physical? Reach, noise, glare, gloves, posture, PPE, mounting height, line-of-sight to the cell. If the answer lives in the body or the room, no model has it.

  4. Who verifies? On the analysis side, name the human and budget the hours. The published comprehensive error rate was 12.4 percent, so assume you will find real errors and plan for them rather than discovering them in review.

A fifth rule, for safety-relevant screens: nothing generated stands unverified. An emergency-stop confirmation, a safe-restart sequence, a zone-override dialog. Those get real operators or they do not ship. ISA already publishes a technical report on this, ISA-TR101.02-2019, on HMI usability and performance, addressing design and validation approaches for process automation interfaces. Use it.

The economics are the actual risk

The International Federation of Robotics reported on 24 September 2026 that the global operational stock of industrial robots reached a record 5 million units, up 9 percent, on more than 600,000 new installations during 2025, an 11 percent rise. Germany installed fewer than 25,000 units, an 8 percent decline.

That combination is where I think the danger sits. Growth of that size means more operators, in more plants, using interfaces built by teams they will never meet, while a contracting home market puts European vendors under margin pressure and sends them looking for lines to cut. Field research is a soft target: expensive, slow, producing nothing that looks like progress, and now facing a product category that promises to replace it for a subscription.

Opinion, not finding: teams that cut plant visits in 2026 and 2027 will not feel the cost for around eighteen months, then feel it all at once, in commissioning, in support-ticket volume, and in the quiet operator workarounds nobody reports upward because the line is running.

Conclusion

My reading of the current evidence is narrow and, I think, useful. Models have become good enough at reading qualitative data that refusing to use them there wastes your team's time, provided a human verifies what comes back. They remain bad at being people, and they are worst at it in exactly the dimension industrial interface design depends on: the spread rather than the centre.

So spend the speed you gain on the analysis side on something specific: more hours on a shop floor, watching somebody who has never met you try to recover from a fault. That is the trade I would make, and I suspect it is the only one that still produces interfaces operators trust at 3 a.m.

Sources

Previous/Next Articles

more aloud

more aloud

Leon Potgieter

Robotics Visual System Designer

AI Did Not Make Palletisers Smarter. It Moved the Configuration Problem Elsewhere.
Palletising is the most automated end-of-line task and the one where AI claims drift furthest from what ships. What actually changed, and who now has to decide.

Leon Potgieter

Robotics Visual System Designer

AI Did Not Make Palletisers Smarter. It Moved the Configuration Problem Elsewhere.

Palletising is the most automated end-of-line task and the one where AI claims drift furthest from what ships. What actually changed, and who now has to decide.

Leon Potgieter

Robotics Visual System Designer

The AI Tooling I Actually Use for HMI Visuals, and the Line I Do Not Cross
A working account of the generative tooling in my HMI pipeline: what it does for cell renders, motion and variants, and what it must never be allowed near.

Leon Potgieter

Robotics Visual System Designer

The AI Tooling I Actually Use for HMI Visuals, and the Line I Do Not Cross

A working account of the generative tooling in my HMI pipeline: what it does for cell renders, motion and variants, and what it must never be allowed near.

Leon Potgieter

Robotics Visual System Designer

Your AI robot render may now count as a deepfake under EU law
Since 2 August 2026, a realistic AI image of a real robot can fall under the EU deep fake rules. Here is how to keep robot visuals true.

Leon Potgieter

Robotics Visual System Designer

Your AI robot render may now count as a deepfake under EU law

Since 2 August 2026, a realistic AI image of a real robot can fall under the EU deep fake rules. Here is how to keep robot visuals true.

Leon Potgieter

Robotics Visual System Designer

The Brand System Should Start at the Panel, Not at the Fair Wall
Most brand work is built for the first ninety seconds, not for five years. The interface faces the hardest constraints, so it should set the rules.

Leon Potgieter

Robotics Visual System Designer

The Brand System Should Start at the Panel, Not at the Fair Wall

Most brand work is built for the first ninety seconds, not for five years. The interface faces the hardest constraints, so it should set the rules.

Leon Potgieter

Robotics Visual System Designer

Perception Does Not Check Its Sources
Gestalt cannot become obsolete because it describes the viewer, not the screen. Which is why it is now dangerous: it will confidently group a guess with a fact.

Leon Potgieter

Robotics Visual System Designer

Perception Does Not Check Its Sources

Gestalt cannot become obsolete because it describes the viewer, not the screen. Which is why it is now dangerous: it will confidently group a guess with a fact.

Leon Potgieter

Robotics Visual System Designer

The HMI Was a Mirror. Now It Has to Be an Argument
For forty years, the industrial HMI represented machine state. AI broke that contract. The screen's job is now accountability, and that is a different design discipline.

Leon Potgieter

Robotics Visual System Designer

The HMI Was a Mirror. Now It Has to Be an Argument

For forty years, the industrial HMI represented machine state. AI broke that contract. The screen's job is now accountability, and that is a different design discipline.

Leon Potgieter

Robotics Visual System Designer

Four months to the Machinery Regulation: what HMI teams must fix
The EU Machinery Regulation applies on 20 January 2027. Several of its requirements end on a screen. Here is what HMI teams should fix now.

Leon Potgieter

Robotics Visual System Designer

Four months to the Machinery Regulation: what HMI teams must fix

The EU Machinery Regulation applies on 20 January 2027. Several of its requirements end on a screen. Here is what HMI teams should fix now.

Leon Potgieter

Robotics Visual Systems Designer

Natural-language robot programming needs a proof screen, not a chat box
Typing "pick the small parts" into a robot is easy. Knowing what the robot understood is the hard part. Design the proof screen first.

Leon Potgieter

Robotics Visual Systems Designer

Natural-language robot programming needs a proof screen, not a chat box

Typing "pick the small parts" into a robot is easy. Knowing what the robot understood is the hard part. Design the proof screen first.

Leon Potgieter

Robotics Visual System Designer

Proof is the new production value in robotics content
AI-generated posts now dominate the LinkedIn feed. In robotics, the content that still earns trust is the kind that proves itself on camera.

Leon Potgieter

Robotics Visual System Designer

Proof is the new production value in robotics content

AI-generated posts now dominate the LinkedIn feed. In robotics, the content that still earns trust is the kind that proves itself on camera.

Leon Potgieter

Robotics Visual System Designer

Robot HMIs need a validation test borrowed from medical devices
Many medical device makers validate critical tasks with real users. Robot HMIs, in my view, rarely do. Here is how to borrow the method.

Leon Potgieter

Robotics Visual System Designer

Robot HMIs need a validation test borrowed from medical devices

Many medical device makers validate critical tasks with real users. Robot HMIs, in my view, rarely do. Here is how to borrow the method.

Leon Potgieter

Robotics Visual System Designer

Slop Is a Decision Failure: Taste, Perception and Why Plausible Is Not Correct
Slop isn't an aesthetic failure; it is output made without a cost function. What taste actually is, why perception is trainable, and where AI still loses.

Leon Potgieter

Robotics Visual System Designer

Slop Is a Decision Failure: Taste, Perception and Why Plausible Is Not Correct

Slop isn't an aesthetic failure; it is output made without a cost function. What taste actually is, why perception is trainable, and where AI still loses.

Leon Potgieter

Robotics Visual System Designer

Show the Runners-Up: UX for Machines That Decide
Automation heuristics are not hard to understand, they are invisible. A working method for exposing a machine's ranked decisions at the right depth in the HMI.

Leon Potgieter

Robotics Visual System Designer

Show the Runners-Up: UX for Machines That Decide

Automation heuristics are not hard to understand, they are invisible. A working method for exposing a machine's ranked decisions at the right depth in the HMI.

Robotics Visual System Designer

Leon/Potgieter

Leon/Potgieter

Robotics Visual Systems Designer; bridging the gap between how automation technology works and how the world understands it.

©2026 Leon Potgieter - All work, all rights.

Offline

Leon Potgieter
Koringberg
Western Cape

South Africa

Leon Potgieter - Koringberg, Western Cape, South Africa