Many medical device makers validate critical tasks with real users. Robot HMIs, in my view, rarely do. Here is how to borrow the method.
Many medical devices have to show, based on risk, that real users can perform critical tasks safely. In my view, robot HMIs rarely do. Here is how to borrow the method.
A robot cell typically goes through a factory acceptance test and a site acceptance test. Both do one thing well: prove the machine does what the specification says. Cycle time, reach, safety functions, I/O. In my view, they rarely prove that the people who will run, set, and maintain the cell can do the risky parts of their job on the interface without making a mistake.
In my view, the operator is the least tested component in most robot cells. Another industry has spent years building a method to test exactly that.
On 29 May 2026, the US Food and Drug Administration published its final guidance on the content of human factors information in medical device marketing submissions. I am not suggesting robot builders should file anything with the FDA. I am suggesting HMI teams in automation should steal the method underneath it, because it answers a question our industry mostly leaves open: how do you know an interface is safe to use, not just safe to run?
What medical device makers are expected to show
The core idea is the critical task. The FDA's human factors guidance (first issued 2016, revised August 2026) defines it as "a user task which, if performed incorrectly or not performed at all, would or could cause serious harm to the patient or user". Manufacturers identify these tasks through a use-related risk analysis, then design the interface to make errors on them unlikely.
Then they test in two phases. Formative evaluation happens during development: user interface evaluation "conducted with the intent to explore user interface design strengths, weaknesses, and unanticipated use errors." Human factors validation testing happens at the end, with representative users performing the tasks under simulated use. The guidance states that, in general, "the minimum number of participants should be 15" per distinct user population, and asks for simulated-use conditions "sufficiently realistic so that the results of the testing are generalizable to actual use". Elsewhere, it lists use-environment factors such as lighting level and noise level.
IEC 62366-1, the usability engineering standard for medical devices, works with the same pair of ideas, formative and summative evaluation, and expects safety-related use scenarios to be part of the validation plan.
The 2026 final guidance did not make this heavier. It made it smarter. According to the FDA's notice, it provides "a risk-based framework" for deciding which human factors information a submission needs. Emergo by UL's summary describes a decision point that weighs user interface use history, familiarity and complexity, and the adequacy of existing risk mitigations, and notes that manufacturers may in some cases provide a robust justification instead of new validation data, even when critical tasks are new or affected.
That is a mature position. Test hardest where errors hurt people. Justify, with evidence, where the interface is familiar and proven.
Why robotics should care
The data on serious robot accidents is thin, but what exists points at the moments an interface mediates.
A NIOSH researcher, Larry Layne, reviewed US robot-related fatalities from 1992 to 2017 in the American Journal of Industrial Medicine. He identified 41 cases. Most involved stationary robots (83 percent). In 78 percent, the robot struck the person while operating under its own power. Maintenance of the robot was mentioned in 24 of the 41 fatalities, 58.5 percent.
A 2024 study in Applied Ergonomics by Sanders, Şener and Chen analysed 77 robot-related accidents in OSHA Severe Injury Reports from 2015 to 2022. For the 54 accidents involving stationary robots, the injuries were mainly finger amputations and fractures to the head and torso.
To be clear about the limits: neither study says an interface caused any of these accidents. My reading is narrower. They show where people get hurt: during maintenance, around stationary cells, when a robot moves under its own power. Those are moments the HMI often mediates. The mode change. The restart after someone has been inside the cell. The recovery after a stop. The setting that decides how fast the robot moves when a person is near it.
If those are the high-consequence tasks, they deserve the same treatment as a nurse programming an infusion pump: identified, designed against, and validated with the people who will actually do them.
[Image: Bar chart of three shares from 41 US robot-related fatalities 1992 to 2017: stationary robots 83 percent, struck while robot under own power 78 percent, maintenance mentioned 58.5 percent] (fig-1.png, see review page)

Figure 1: Where fatal robot accidents happened, 41 US cases, 1992 to 2017. Source: Layne, American Journal of Industrial Medicine, 2023.
Five users or fifteen?
Every UX designer knows the number five. Jakob Nielsen's 2000 article argued that testing with five users finds about 85 per cent of usability problems, and that many small tests yield better results than one big one. His own example: rather than one study with fifteen users, run three studies with five.
FDA guidance generally recommends fifteen per user group. That looks like a contradiction. It is not.
Nielsen's five is a formative number. In his words, its purpose is "to improve the design and not just to document its weaknesses." The FDA's fifteen is a validation number. Its purpose is to show that, in the end, intended users can use the device without serious use errors or problems under expected use conditions. You need both. You iterate with small rounds, then you prove the result with a larger one.
Nielsen also noted that distinct user groups change the maths: he suggested three to four users per group when testing two groups. In a robot cell, the groups are distinct by any definition. An operator, a setter, a maintenance technician and an integrator's programmer do different tasks, with different training, often in different languages.
My view is that robotics rarely does either phase formally. Too often, the engineers who built the interfaces review them, then the customer tests them in production.
[Image: Diagram comparing three formative rounds of five participants with a summative validation of fifteen participants per user group] (fig-2.png, see review page)

Figure 2: Formative rounds improve the design; summative validation proves it. Numbers from Nielsen (2000) and FDA human factors guidance (2016, revised 2026).
The digital twin is your simulated-use lab
Medical device validation usually happens under simulated use, not in a live hospital. Robotics already has the equivalent and mostly uses it for something else.
Virtual commissioning, in the VDMA definition quoted in a 2026 paper by Ferle and colleagues in the International Journal of Advanced Manufacturing Technology, is "a simulation-based validation methodology for mechatronic systems in which the control system is connected to a simulation model of a component, machine, or production system to perform early, development-accompanying tests before physical realisation."
If the control system is connected to a simulation, the HMI can be too. In my view, that makes the virtual commissioning model the cheapest usability lab an automation company will ever own. You can put a setter in front of the real interface, connect it to a simulated cell, and ask it to recover from a collision stop months before the hardware exists.
The twin has limits, and the FDA's list of environmental factors shows them. A simulation does not reproduce the noise of a hall, glare on a panel, or gloves on a touchscreen. Use the twin for formative rounds. Run the final validation at the real cell, or a physical mock-up, under realistic conditions.
A critical-task validation method for robot HMIs
Here is the method I would adapt from medical devices. It is a design practice, not a legal requirement in machinery law, and it does not replace the risk assessment.
Derive critical tasks from the risk assessment. List every task where an operator error or omission could cause serious harm. Typical candidates: switching operating modes, restarting after someone has entered the cell, recovering from a protective stop, changing safety-relevant parameters, teaching positions near fixtures, maintenance access.
Name the user groups. Operator, setter, maintenance technician, programmer. Each has its own task list because they do different things on the same screens.
Run formative rounds in the digital twin. Run three rounds with about five participants, changing the design between rounds. Record use errors and close calls, not opinions.
Validate with a larger group under real conditions. Borrow the FDA bar of 15 per user group where you can, at the real cell or a mock-up, with realistic noise, light and protective gloves. The pass criterion is simple: no use errors on critical tasks that could lead to harm, and a root cause for every error you observe.
Justify reuse honestly. For a modified HMI, use a stricter version of the FDA's logic: if a critical task's interface is unchanged, familiar and already proven in the field, document why it does not need retesting. If it changed, test it, even though the FDA allows a robust justification in some cases.
Write it down. A short human factors report next to the risk assessment: critical tasks, user groups, results, fixes.
Sixty sessions for four user groups sounds expensive. I bet it costs less than one retrofit of a confusing restart flow across an installed base.
[Image: A usability test session with a participant using a tablet while an observer takes notes] (fig-3.png, see review page)

Figure 3: Validation is observation of real users on critical tasks, not a design review. AI-generated illustration.
Where AI fits, and where it does not
AI tools can make this cheaper. They can help transcribe and tag session recordings, cluster observed errors and draft the report. That is analysis work, and it is where the time goes.
They cannot replace the participants. A validation test shows that real people from a defined user group can perform a task safely. In my view, a simulated user cannot provide that evidence, because the evidence is the real person's behaviour.
Conclusion
Critical-task validation is not paperwork for its own sake. In my view, the method exists because use errors on a small number of tasks can cause serious harm, and because asking engineers whether their own interface is clear does not answer the question.
Robot HMIs share that profile. In my view, a few tasks carry most of the risk, and the people doing them are not the people who designed them. We already own the simulation tools to test those tasks early. What is missing is the habit. I would start it on the next project, with one user group and one critical task, and see what the first five sessions show.
Sources
US Food and Drug Administration, "Content of Human Factors Information in Medical Device Marketing Submissions; Guidance for Industry and Food and Drug Administration Staff; Availability", Federal Register, 29 May 2026. https://www.federalregister.gov/documents/2026/05/29/2026-10734/content-of-human-factors-information-in-medical-device-marketing-submissions-guidance-for-industry
Emergo by UL, "Key Updates in the Final FDA Guidance: Content of Human Factors Information in Medical Device Marketing Submissions (2026)". https://www.emergobyul.com/news/key-updates-final-fda-guidance-content-human-factors-information-medical-device-marketing
US Food and Drug Administration, "Applying Human Factors and Usability Engineering to Medical Devices", guidance issued 3 August 2026 (originally 3 February 2016). https://www.fda.gov/media/80481/download
Johner Institute, "Usability validation: Compliant with IEC 62366-1 and FDA". https://blog.johner-institute.com/iec-62366-usability/usability-validation/
Layne LA. "Robot-related fatalities at work in the United States, 1992-2017." American Journal of Industrial Medicine, 2023. DOI 10.1002/ajim.23470. https://stacks.cdc.gov/view/cdc/230667/cdc_230667_DS1.pdf
Sanders NE, Şener E, Chen KB. "Robot-related injuries in the workplace: An analysis of OSHA Severe Injury Reports." Applied Ergonomics, vol. 121, 2024. https://www.sciencedirect.com/science/article/abs/pii/S0003687024001017
Nielsen J. "Why You Only Need to Test with 5 Users." Nielsen Norman Group, 18 March 2000. https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/
Ferle F, Chen S, Kuhn AM, et al. "Continued use of virtual commissioning models: A novel approach toward digital twins for automated production systems." The International Journal of Advanced Manufacturing Technology, vol. 142, 2026. https://link.springer.com/article/10.1007/s00170-025-17195-y













