
Are we testing the right things in multi-agent AI? (Ruchira Dhar, University of Copenhagen) — UK AI Forum
About this event
With the rise of large language models (LLMs), multi-agent systems (MASs) are becoming increasingly popular and getting deployed for multiple tasks. Evaluating their performance has become critical for both capability assessment and AI governance. In this talk I take a look at the evaluatio



