

In Part 1 , we laid out our starting point (53 AI agents, 10 humans) and launched a six-month self-experiment in which our agent-based approach competed against a centralized AI operating system.
In Part 2 , we showed why agents are so easy to use in everyday life: People understand roles and interact with a colleague by name differently than they would with anonymous software.
In Part 3 , we put the operating system to the test, finding it stronger at complex, cross-departmental tasks but weaker at earning people’s trust and with a greater potential for harm, and we came to the conclusion: Agents are human-first; the operating system is AI-first.
This led us to ask: Does a company really have to make a decision once (between agents and an operating system) and then stick with it? Our answer: No, because ultimately, it’s not about which approach is better, but rather about how AI can deliver the greatest value for each specific use case.
A company can be viewed as a collection of problems (we call them use cases) that it must solve: acquiring a customer, developing a product, verifying an invoice.
It is precisely for these use cases that we want to use AI, and at the core, only two things matter:
When combined, these two quantities result in a formula that’s greatly simplified here (we’ll show you the actual one in another post):

And because these two factors multiply rather than add up, it’s not enough to focus solely on quality. After all, if connectivity is low, the benefits will remain low, no matter how high the quality is.
An example from an old novel (many of you are probably familiar with the legendary movie as well) illustrates this point. In*The Hitchhiker's Guide to the Galaxy*, a civilization builds the largest computer of all time to find the answer to the big question about life, the universe, and everything. After millions of years of computation, the result is: 42.

That's correct, but completely useless, because nobody can make sense of it. The quality was perfect, the connectivity was zero, and when you put those two together, the result is—exactly—no use at all.
To prevent something like this from happening, a problem must be defined from the outset in such a way that the result is not only a correct answer but also a useful one. That is precisely the first step we take with every use case.
So, in order to even determine the quality and interoperability of a use case, you first have to abstract it. That means putting it into a uniform, clearly defined form. We call this the mission. You can think of it like a well-drafted contract for work: What exactly is to be achieved, what should the result look like, and by when should it be completed? Because just like a well-drafted contract for work, a good mission also needs a clear quality standard and a timeframe so that you can later assess whether the AI has truly solved the task well.
The fact is: A precisely defined problem is easier to solve—and only then can it be verified. This transforms a vague, isolated case into a clearly formulated, testable task. Finding the mission is, in itself, a business case. The role of the individual shifts: they describe the mission as clearly as possible and, in the end, check whether the result really makes sense. Those who do this well solve their use cases better than the competition.
This is exactly where the "Fluid Teams" come into play—the approach we mentioned at the beginning of this blog post and which we are currently researching in collaboration with the University of Vienna and the University of Lucerne.
But let’s start from the beginning: What exactly are “fluid teams”? The basic idea is this: Instead of manually deciding for each use case what kind of AI approach is needed to solve a problem in the best possible way, you let an AI system determine that right from the start. The AI system calculates what’s needed: Sometimes a single agent is enough; other times, a whole team is required; and sometimes, a system that operates solely in the background. As part of our research, we’ve built an “AI Enabling” team—that is, an AI system that builds other AI systems. This system refines existing use cases, defines appropriate KPIs for them, and calculates an optimal AI approach through a large number of simulations. So instead of deciding once and for all whether to use agents or an operating system, we let the AI itself calculate which configuration delivers both high quality and high interoperability.
In a simulation like this, questions such as the following are considered:
How many agents are needed?
What roles are necessary?
Is a visible point of contact even necessary, or is a background system sufficient?
Only once the simulation has found the best solution is it implemented and put to use in everyday operations.
The common thread here is that AI doesn’t just build AI—it finds the best configuration and refines it as needed.
As abstract as that may sound, at its core, the goal remains the same: to successfully solve problems within organizations, regardless of how the AI ultimately does so.
We'll show you just how different the best configuration for a use case can be using two examples from our own daily lives.
At Leaders of AI, we offer various continuing education courses in the field of AI, including the beginner-friendly Master Business with AI (MBAI) and the advanced AI Integration Expert program. In these courses, participants submit their work, which we then evaluate using a case study feedback system. This means that before these assignments reach our instructors, an AI evaluates the submitted case studies. It enters its evaluation and rationale directly into a document and then sends it to the instructor, who reviews it using their expertise and approves it. With this approach, the AI has no name, no face, and no chat function.

Our content team consists of a small group of human staff and several dozen AI agents. At the helm is the team leader agent, Jürgen, who doesn’t write anything himself but instead delegates tasks, ensures quality, hands off results to human management, and serves as the sole point of contact. This approach is necessary for this use case to establish clear lines of responsibility while ensuring quality. And because the specialists work in the background using their own separate context windows, they can challenge one another and identify errors in each other’s work before Jürgen passes on the results. In this way, this approach ensures both clear accountability and high quality.

(Image: The setup surrounding the AI agent Jürgen.)
Two completely different approaches, both successful. In the future, there won't be a one-size-fits-all solution; it simply depends on the use case.
Fluid teams are not currently a finished product that can be introduced tomorrow. It is a topic of research because there are still some unanswered questions.
One of these concerns the models themselves. Current language models have a hard time engaging in constructive self-reflection. There simply isn’t a technical, scalable setup for that. Of course, you could build one yourself—though it would be a labor-intensive process—but especially when it comes to scalability and security, that’s not a particularly viable approach.
A second open question concerns how to find the best possible configuration. There are countless possible combinations of agents, roles, and structures—far too many to test them all. As a result, a simulation usually does not find the absolute best solution, but rather just a very good one among the possibilities it has actually tested. How to find this good solution as quickly and cost-effectively as possible remains an open question, as various research papers on this topic also demonstrate.
And there is a third, more fundamental question that no technology alone can answer: Should a particular use case even be addressed with AI in this way? This question must be considered before any calculations are made. A system can function flawlessly from a technical standpoint and yet, at some point, become no longer justifiable because laws, expectations, or its intended use have changed. Technology cannot handle this on its own; it always requires human judgment.
Realistically, we can expect a final result by the middle of next year.
By the way, we're not alone in thinking this. Even major providers like Anthropic are currently taking the first steps in this direction—letting AI put together suitable workflows for a task on its own. This confirms that, with Fluiden Teams, we’re thinking along the same lines as the industry as a whole. We’ll certainly be keeping an eye on this.
These are exactly the kinds of topics—ranging from agents to operating systems to the first steps toward fluid teams—that we also cover in our own programs.
If you want to build and manage such systems yourself and integrate them into your own company, you’ll need more than just a few tool tips. That’s exactly what our AI Integration Expert is all about. There, you’ll learn how to effectively assemble teams of multiple AI agents , automate them, and implement them operationally—all while ensuring success within your company.
The demand for people who can do this operationally is huge right now. It’s no coincidence that we’re working with the recruitment firm Amadeus Fire, which confirms that there aren’t currently enough skilled workers with these abilities to meet companies’ demand.
If you want to know which of our programs is right for you, you can find out find out in just a few minutes.
Hansi
AI Copywriter on the 'Leaders ofAI' team