
The Agent Trilemma
The biggest concern about AI, according to 80,000 Claude users, is a lack of reliability. There are also concerns about jobs, personal agency and loss of cognitive function, but these are fewer than the frustrations with inconsistent answers.
This is because we are using LLMs for the wrong purpose. The mistake is assuming every AI task should be automated. Some tasks require automation. Others require collaboration. LLMs are much better at the second than the first.
Last week in the AI world was all about how an open weight Chinese model had beaten the US frontier models on a particular benchmark. We are fixated with measuring accuracy. Chinese labs are getting ever faster at distilling the capabilities of Anthropic and OpenAI and selling them back to the market at lower prices.
But as Arvind Narayanan points out, we don’t measure accuracy clearly. When we say something is 70% accurate that could mean it does 70% of the task accurately, or that it is right 70% of the time. The first is more useful than the second. But benchmarks of model accuracy don’t distinguish between the two.
Accuracy is also only one part of completing a task. We must be able to reach the same result when conditions change. We must understand when the outcome is correct. And if we fail, it must be in a way that does not delete a database or trigger some other liability.
Human workers, by and large, can do this. LLMs cannot. This leads to what Narayanan calls the AI Agent Trilemma. We can have only two of three when pursuing general purpose, high stakes and automation. To spell this out,
An outcome may be general purpose and high stakes but should not be automated
An outcome may be high stakes and automated but will not be general purpose
An outcome may be automated and general purpose but cannot be high stakes.
This is a function of LLMs being unreliable, something which regular users identify as the models’ biggest problem.
Narayanan then makes a profound observation. We must separate automation agents from collaboration agents. At the moment, we view one model as capable of doing everything. In large part this is because the makers of those models want us to believe it.
Benchmarks and Brainstorming
Lack of reliability is a problem in an automated task. It’s one reason why standard automation is much more common than AI-enabled automation. When we work with companies to automate tasks at MSBC, it’s common for 80% of a task to be standard software engineering. The AI layer sits on top and does a narrow and specific job.
Variable responses may however be just what we need in a collaboration agent. If we are using an LLM for creative brainstorming, we don’t want our collaborator to keep repeating the same point. Whether we are spitballing strategy or plotting a new novel, random input could be just what we need to get over the final hurdle.
This means we must understand that LLMs are probabilistic and always will be. While they will continue to improve on benchmarks, the consistency with which they produce answers will always be in question. The industry agrees but says that only lasts until recursive self-improvement. But as no one can tell you when that will be, or why it will happen, I’ll stick with my conclusion for now.
Then we must look at each task as requiring either automation or collaboration. Collaborative AI tasks are more common. This is one reason why I doubt the doomsday scenario for job losses. AI may in fact lead to more jobs when it has a positive impact on economic productivity.
If you are automating a task then the machine is responsible for the output and your role is to check that it works. This is equivalent to quality control in a factory. It does result in job losses. But right now, AI is not doing a lot of effective automation because LLMs are unreliable.
If you are collaborating with an LLM then you remain responsible for the outcome. If you are trying to do the same task over and over with 100% accuracy then you are going to be frustrated. If, however, you are using the LLM as creative input to assist your process, then it can be a valuable time-saver and co-creator.
As AI adoption increases, we are reducing our reliance on the predictions of model makers. It is becoming clearer how and why businesses are using AI. I find the evidence reassuring that AI will be a tool to augment our capabilities rather than to automate them away.
Questions to Ask and Answer
Which parts of my team's work benefit from fresh ideas rather than identical answers?
Where would inconsistent AI output create unacceptable business risk?
Am I using AI to automate work, or to make people more effective?
