In 2026, LLMs matched human analysts at coding interview data and still failed at inventing it. That split should decide where AI sits in HMI research.
Twice in the last few weeks I have sat through a demo where a "synthetic operator" answered questions about a robot cell it had never stood next to. In industrial B2B, the pitch lands, because recruiting six real machine operators, on shift, across three plants and two countries, is genuinely hard. A plant manager has to sign off. Somebody has to cover the line. The operator you need most is always the busiest. Against that, a persona that answers in four seconds and never cancels looks like a bargain.
I think it is the worst bargain on offer to an HMI team right now, and 2026 gave me numbers to argue it with rather than taste. They split cleanly in two directions.
The 2026 evidence draws a line, and the line is not where vendors put it
In April 2026, PLOS Digital Health published a blinded comparison of large language models against human analysts doing thematic analysis. Two general-purpose models and one specialised tool coded a focus-group transcript, and blinded human coders and an expert consensus panel judged the results.
On deductive coding, the models matched the humans. Agreement was 93.5 percent for the LLMs (95% CI 92.5 to 94.5) against 92.7 percent for the humans (95% CI 91.6 to 93.9), with an identical kappa of 0.34 and near-identical Gwet's AC1 scores of 0.93 and 0.92. All three models cleared non-inferiority against the human analysts (p < 0.0001), and two of them, ChatGPT-5 and Claude 4 Sonnet, reached statistical superiority on the deductive task. On inductive theme generation, the picture was thinner: only ChatGPT-5 cleared non-inferiority, and then only at p = 0.043.
The error figures matter as much as the agreement figures. Strict hallucination was low at 1.2 percent, but the comprehensive error rate was 12.4 percent, and the authors are explicit that human verification remains necessary. They are equally explicit about their study's limits: one transcript, on a topic sitting squarely inside the training data, judged against panel consensus rather than objective ground truth.
One transcript in one domain does not license a blanket claim, and the authors say so. But applying an existing codebook across forty operator interviews is the same class of task the models just performed at human parity, at a speed no human team will match, provided somebody verifies the output line by line and expects roughly one item in eight to carry an error. That is a real capability, and it is worth money.
Now look at what happens when you ask a model to produce the data instead of read it.

Figure 1: LLMs matched human coders on deductive thematic coding, but roughly one output in eight carried an error. Source: Hill et al., PLOS Digital Health, April 2026.
Synthetic participants fail in exactly the place industrial research lives
Jim Lewis and Jeff Sauro published a review of the synthetic-user experiments in April 2026. Their tally was 9 encouraging findings against 14 discouraging ones. The interesting part is not the headline count. It is which categories scored zero encouraging findings. Matching expected variance came out at zero encouraging findings against three discouraging. Good qualitative depth came out at zero against three. Two further categories, representativeness and matching regression weights, also scored zero.
The underlying studies are blunt. Park and colleagues attempted 14 classic social science studies with GPT-3.5 and got six sets of unanalysable data, five failed replications and three successes. Shrestha and colleagues put 43 policy questions to GPT-4 across three countries and found about 70 percent of the synthetic answers significantly different from human ones, even though the means looked reasonably aligned. Almeida and colleagues tested four models across eight psychology studies of legal and moral reasoning, found GPT-4 the closest match to human responses, and still reported systematic differences, with a tendency for models to exaggerate effects.
The clearest single result is older and comes from political science. In Political Analysis (2024), Bisbee and colleagues had ChatGPT adopt demographic personas and rate 11 sociopolitical groups, then compared the output against American National Election Study data. Averages tracked the benchmark closely. Everything below the average fell apart. Standard deviations in the synthetic estimates were far smaller than in the real survey. Forty-eight percent of regression coefficients were statistically different from the ANES estimates, and among those, 32 percent carried the wrong sign. Power analyses based on the synthetic data underestimated the sample size you would actually need by nearly an order of magnitude. The responses also shifted substantially over three months as the model was updated.
Take those findings together and the failure mode is specific. Synthetic participants give you a plausible centre and a collapsed distribution.

Figure 2: The generation side of the ledger. Left: encouraging against discouraging findings in the 2026 review, overall and in the two categories that scored zero against three discouraging findings each. Right: share of synthetic survey regression coefficients that differed significantly from the real survey benchmark, and the share of those differing coefficients that carried the wrong sign. Sources: Lewis and Sauro, MeasuringU, April 2026; Bisbee et al., Political Analysis, 2024.
In HMI, the collapsed distribution is the part you needed
A plausible centre is near worthless in interface work for robot cells. The average operator does not break your error-recovery flow. The outlier does.
The outlier is the setter in cut-resistant gloves who cannot hit a small target. It is the night-shift operator who was never trained on the cell and learned it from a colleague in a few minutes. It is the person who has retaught the same pick point so often that they have stopped reading your confirmation dialogue. It is the technician who reaches your fault screen already angry, line down, supervisor behind him.
Every one of those is a tail case, and in my experience every one of them produces a design change. A model that reliably reproduces the group mean and reliably flattens the spread is optimised to hide precisely the people whose behaviour you need to see. Worse, it hides them while producing output that reads like research, complete with quotes, which is the failure mode that survives a review meeting.
The Nielsen Norman Group named a second problem in 2024, and I haven't seen it solved since. Models want to please. Rosala and Moran call solution validation with synthetic users "incredibly risky" because "AI loves to please, every idea is often seen as a good one." Their synthetic users "seem to care about everything" instead of ranking needs. In a discipline whose entire job is deciding which of nine competing demands gets the top of the screen, a participant who values everything equally is not a participant. It is noise with good grammar.
Context is the finding, and it does not exist in the model
ISO 9241-210:2019 exists because human-centred design is not a survey exercise. It sets requirements and recommendations for applying human-centred design across the lifecycle of computer-based interactive systems, and Annex B provides a checklist to support claims of conformance. My reading is that it requires you to produce an account of how you actually understood the environment your interface runs in.
A language model can't recover that environment. The ambient noise level makes your audio alarm decorative. It is a panel mounted for the tallest person on the commissioning team and operated by the shortest person on the shift. In my experience, the laminated cheat sheet taped beside the screen is the most valuable artefact on any HMI field visit because it records every place the interface failed. Shift handover lands at the same moment as changeover, so your setup wizard gets used by the most distracted person in the building.

The part no model has access to. AI-generated illustration.
Shivani Kapania and colleagues put this to 19 qualitative researchers in work published at CHI 2025. The researchers started sceptical, were surprised at how similar the narratives looked when the model was given the interview probe, then over several turns identified what was missing: responses lacking in "palpability and contextual depth", the foreclosure of participants' consent and agency, and a real risk of delegitimising qualitative methods altogether. They call the substitution the "surrogate effect". That phrase is the right one. You are not saving research. You are producing a surrogate for it and then making decisions as though you had done it.
Where AI has actually earned a place in my process
I am not arguing for less AI in HMI research. I am arguing for putting it on the correct side of the line.
It belongs on the analysis side, hard. Codebook application across a large interview set, now with published evidence behind it. First-pass clustering of free-text service tickets. Triaging telemetry into candidate problem areas before a human decides which are real. Stress-testing an interview guide before you burn a plant visit on it.
It also belongs in preparation, the one the Nielsen Norman Group endorsed: synthesising available material about a user group into something digestible, learning a new domain, piloting a guide, and generating hypotheses you then test with actual people.
It does not get to be the person. Not for concept validation, not for prioritisation, not for a persona printed and pinned to a wall for two years.

Figure 3: The band I put each activity in, and the four questions I run before letting a model stand in for anybody.
A substitution test you can run in a meeting
When somebody proposes replacing a research activity with a model, four questions settle it.
Am I reading data or creating it? Reading is supported by evidence. Creating is not. If the answer is creating, stop here.
Does the decision depend on the spread or the average? Interface decisions almost always depend on the spread, and synthetic output has a documented variance problem in both the survey work and the review above. Anything hinging on outliers, edge cases or minority workflows is out of scope.
Is the finding physical? Reach, noise, glare, gloves, posture, PPE, mounting height, line-of-sight to the cell. If the answer lives in the body or the room, no model has it.
Who verifies? On the analysis side, name the human and budget the hours. The published comprehensive error rate was 12.4 percent, so assume you will find real errors and plan for them rather than discovering them in review.
A fifth rule, for safety-relevant screens: nothing generated stands unverified. An emergency-stop confirmation, a safe-restart sequence, a zone-override dialog. Those get real operators or they do not ship. ISA already publishes a technical report on this, ISA-TR101.02-2019, on HMI usability and performance, addressing design and validation approaches for process automation interfaces. Use it.
The economics are the actual risk
The International Federation of Robotics reported on 24 September 2026 that the global operational stock of industrial robots reached a record 5 million units, up 9 percent, on more than 600,000 new installations during 2025, an 11 percent rise. Germany installed fewer than 25,000 units, an 8 percent decline.
That combination is where I think the danger sits. Growth of that size means more operators, in more plants, using interfaces built by teams they will never meet, while a contracting home market puts European vendors under margin pressure and sends them looking for lines to cut. Field research is a soft target: expensive, slow, producing nothing that looks like progress, and now facing a product category that promises to replace it for a subscription.
Opinion, not finding: teams that cut plant visits in 2026 and 2027 will not feel the cost for around eighteen months, then feel it all at once, in commissioning, in support-ticket volume, and in the quiet operator workarounds nobody reports upward because the line is running.
Conclusion
My reading of the current evidence is narrow and, I think, useful. Models have become good enough at reading qualitative data that refusing to use them there wastes your team's time, provided a human verifies what comes back. They remain bad at being people, and they are worst at it in exactly the dimension industrial interface design depends on: the spread rather than the centre.
So spend the speed you gain on the analysis side on something specific: more hours on a shop floor, watching somebody who has never met you try to recover from a fault. That is the trade I would make, and I suspect it is the only one that still produces interfaces operators trust at 3 a.m.
Sources
Hill C, Dahil A, Simpson G, Hardisty D, Keast J, Pinn CK, Dambha-Miller H. "Large language models for thematic analysis in healthcare research: A blinded mixed-methods comparison with human analysts." PLOS Digital Health, 3 April 2026. https://journals.plos.org/digitalhealth/article?id=10.1371%2Fjournal.pdig.0001189
Lewis J, Sauro J. "A Review of Experiments with Synthetic Users." MeasuringU, 14 April 2026. https://measuringu.com/review-of-experiments-with-synthetic-users/
Bisbee J, Clinton JD, Dorff C, Kenkel B, Larson JM. "Synthetic Replacements for Human Survey Data? The Perils of Large Language Models." Political Analysis, vol. 32, no. 4, pp. 401-416, 2024. https://www.cambridge.org/core/journals/political-analysis/article/synthetic-replacements-for-human-survey-data-the-perils-of-large-language-models/B92267DC26195C7F36E63EA04A47D2FE
Kapania S, Agnew W, Eslami M, Heidari H, Fox SE. "Simulacrum of Stories: Examining Large Language Models as Qualitative Research Participants." Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 25 April 2025, pp. 1-17. DOI 10.1145/3706598.3713220. Preprint: https://arxiv.org/abs/2409.19430
Rosala M, Moran K. "Synthetic Users: If, When, and How to Use AI-Generated 'Research'." Nielsen Norman Group, 21 June 2024. https://www.nngroup.com/articles/synthetic-users/
ISO 9241-210:2019, "Ergonomics of human-system interaction, Part 210: Human-centred design for interactive systems." https://www.iso.org/standard/77520.html
International Federation of Robotics. "Five Million Robots now Operate in Factories Globally." 24 September 2026. https://ifr.org/ifr-press-releases/news/world-robotics-2026
International Society of Automation. ISA-101 series of standards, including ISA-101.01-2015 and ISA-TR101.02-2019 "HMI Usability and Performance." https://www.isa.org/standards-and-publications/isa-standards/isa-101-standards













