Industry 5.0 promotes increased operator involvement in production through the human-in-the-loop paradigm, enabled by robotic systems that ensure safe and effective human-robot collaboration (HRC). In parallel, Vision Language Models (VLMs) are increasingly used as high-level decision-making components in robotics due to their zero-shot reasoning capabilities. In this paper, we analyze and compare Claude-4.5 Sonnet, Gemini-2.5 Pro, and GPT-5.1 across a set of robot-assisted assembly tasks, evaluating their performance in terms of latency and accuracy. With reference to a HRC scenario involving the assembly of a toy car, we define two representative tasks: recognition of the assembled components and prediction of the next component to be assembled. Experimental results demonstrate the feasibility of the approach as well as the strengths and limitations of the compared VLMs.
An Experimental Comparison of Visual Language Models in Robotic-Assisted Assembly
Rrapi F.;Portelli B.;Serra G.;Scalera L.
2027-01-01
Abstract
Industry 5.0 promotes increased operator involvement in production through the human-in-the-loop paradigm, enabled by robotic systems that ensure safe and effective human-robot collaboration (HRC). In parallel, Vision Language Models (VLMs) are increasingly used as high-level decision-making components in robotics due to their zero-shot reasoning capabilities. In this paper, we analyze and compare Claude-4.5 Sonnet, Gemini-2.5 Pro, and GPT-5.1 across a set of robot-assisted assembly tasks, evaluating their performance in terms of latency and accuracy. With reference to a HRC scenario involving the assembly of a toy car, we define two representative tasks: recognition of the assembled components and prediction of the next component to be assembled. Experimental results demonstrate the feasibility of the approach as well as the strengths and limitations of the compared VLMs.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


