Evaluating Large Language Models with RAG Capability: A Perspective from Robot Behavior Planning and Execution

Jin Yamanaka; Takashi Kido

doi:10.1609/aaaiss.v3i1.31254

Authors

Jin Yamanaka Fujitsu Research of America
Takashi Kido Teikyo University

DOI:

https://doi.org/10.1609/aaaiss.v3i1.31254

Keywords:

Impact of GenAI on Social and Individual Well-being

Abstract

After the significant performance of Large Language Models (LLMs) was revealed, their capabilities were rapidly expanded with techniques such as Retrieval Augmented Generation (RAG). Given their broad applicability and fast development, it's crucial to consider their impact on social systems. On the other hand, assessing these advanced LLMs poses challenges due to their extensive capabilities and the complex nature of social systems. In this study, we pay attention to the similarity between LLMs in social systems and humanoid robots in open environments. We enumerate the essential components required for controlling humanoids in problem solving which help us explore the core capabilities of LLMs and assess the effects of any deficiencies within these components. This approach is justified because the effectiveness of humanoid systems has been thoroughly proven and acknowledged. To identify needed components for humanoids in problem-solving tasks, we create an extensive component framework for planning and controlling humanoid robots in an open environment. Then assess the impacts and risks of LLMs for each component, referencing the latest benchmarks to evaluate their current strengths and weaknesses. Following the assessment guided by our framework, we identified certain capabilities that LLMs lack and concerns in social systems.

Evaluating Large Language Models with RAG Capability: A Perspective from Robot Behavior Planning and Execution

Authors

DOI:

Keywords:

Abstract

Downloads

Published

Issue

Section

Information