PriMera Scientific Engineering (ISSN: 2834-2550)

Research Study

Volume 9 Issue 2

Improving Cooperation in Collaborative Embodied AI

Hima Jacob Leven Suprabha*, Laxmi Nag Laxminarayan Nagesh, Ajith Nair, Alvin Reuben Amal Selvaster, Ayan Khan, Raghuram Damarla, Sanju Hannah Samuel, Sreenithi Saravana Perumal, Titouan Puech, Venkataramireddy Marella, Vishal Sonar, Alessandro Suglia, and Oliver Lemon

July 28, 2026

Abstract

The integration of Large Language Models (LLMs) into multiagent systems has opened new possibilities for collaborative reasoning and cooperation with AI agents. This paper explores different prompting methods and evaluates their effectiveness in enhancing agent collaborative behaviour and decisionmaking. We enhance CoELA, a framework designed for building Collaborative Embodied Agents that leverage LLMs for multi-agent communication, reasoning, and task coordination in shared virtual spaces. Through systematic experimentation, we examine different LLMs and prompt engineering strategies to identify optimised combinations that maximise collaboration performance. Furthermore, we extend our research by integrating speech capabilities, enabling seamless collaborative voice-based interactions. Our findings highlight the effectiveness of prompt optimisation in enhancing collaborative agent performance; for example, our best combination improved the efficiency of the system running with Gemma3 by 22% compared to the original CoELA system. In addition, the speech integration provides a more engaging user interface for iterative system development and demonstrations.

Alice and Bob Voice Chat GUI Video demonstration.

Github link to codebase.

Keywords: Collaborative Embodied AI; co-operative AI; LLMs; evaluation

References

  1. Agashe S., et al. “Agent S: An open agentic framework that uses computers like a human”. arXiv (2024): arXiv:2410.08164.
  2. DeepSeek-AI., et al. “DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning”. (2025).
  3. Deitke M., et al. “Retrospectives on the Embodied AI Workshop”. arXiv (2022): arXiv:2210.06849.
  4. Grattafiori A., et al. “The Llama 3 herd of models”. (2024).
  5. Jiang AQ., et al. “Mistral 7B”. (2023).
  6. Ben Hassouna A, Chaari H and Belhaj I. “LLM-Agent-UMF: LLM-based agent unified modeling framework for seamless integration of multi active/passive core-agents”. arXiv (2024): arXiv:2409.11393.
  7. Ollama Inc. Ollama: Get up and Running with Large Language Models. (2025).
  8. Li G., et al. “CAMEL: Communicative agents for 'mind' exploration of large language model society”. (2023).
  9. Mandi Z, Jain S and Song S. “RoCo: Dialectic multi-robot collaboration with large language models”. (2023).
  10. Microsoft AutoGen Team. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Framework. (2023).
  11. Puig X., et al. “VirtualHome: Simulating household activities via programs”. (2018).
  12. Puig X., et al. “Habitat 3.0: A co-habitat for humans, avatars and robots”. (2023).
  13. Royce R., et al. “Enabling novel mission operations and interactions with ROSA: The Robot Operating System Agent”. (2025).
  14. Shinn N., et al. “Reflexion: Language agents with verbal reinforcement learning”. (2023).
  15. Gemma Team., et al. “Gemma 3 technical report”. (2025).
  16. UMass Embodied AGI Lab. CoELA GitHub Repository. (2023).
  17. Wang G., et al. “Voyager: An open-ended embodied agent with large language models”. (2023).
  18. Wang Z., et al. “Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents”. (2024).
  19. Yao S., et al. “Tree of thoughts: Deliberate problem solving with large language models”. (2023).
  20. Yao S., et al. “ReAct: Synergizing reasoning and acting in language models”. (2023).
  21. Zhang H., et al. “Building cooperative embodied agents modularly with large language models”. (2024).

Foot Notes

  1. https://github.com/Himajacob/Co-LLM-Agents?tab=readme-ov-file#a3-structured-baseprompt—forced-reasoning
  2. https://github.com/Himajacob/Co-LLM-Agents?tab=readme-ov-file#b3-cprompt2—one-shot