AI Strategy

9 min read

We Should Expect More From AI Software

Why AI harnesses should make skilled people materially better at their jobs instead of pretending that memory, tools, prompts, and constant output amount to a human employee.

AK

Arsalan Khawaja

Senior professional working with an AI system designed to improve judgement rather than replace the role

Harness wrappers are the new ChatGPT wrappers

How intelligent is AI really, if after everything it can supposedly do, we have collectively decided that the most imaginative use for it is to replace senior employees with something resembling an intern?

One might have hoped for a little more ambition.

We are, after all, repeating a familiar mistake. The early days of generative AI produced an extraordinary number of ChatGPT wrappers: thin products built around a general model and presented as if the surrounding interface had somehow created a new category of software.

Now we have harness wrappers.

Give the model memory. Give it tools. Add a prompt, some workflows, a few permissions and perhaps a pleasant name. Then describe the result as a CFO, a Chief of Staff, a marketer or a product manager available for a monthly subscription.

It is difficult to take the claim entirely seriously.

A human employee is not simply a collection of outputs. People accumulate experience, form opinions, change their minds, develop judgement, understand relationships, notice things that were never written down and occasionally possess the useful quality of having a spine.

None of this means AI is unimportant. Quite the opposite. The models are remarkably capable. We can create enormous value with them without pretending that a prompt, some memory and a toolset amount to a person.

The more interesting question is why we seem so determined to make that pretence.

Organisations do not generally suffer from a shortage of activity

What exactly is the objective?

Did we collectively decide that reducing the cost of labour was a strategy in itself? When did constant output become evidence that the work was useful? And when did producing more become synonymous with producing something better?

AI companies have placed considerable emphasis on long-running tasks and long-horizon benchmarks as demonstrations of intelligence. We seem to have accepted, rather casually, that this is also a useful model for work.

But a senior role in an organisation is not simply a very long task.

A good product manager is not valuable because they can perform product-management-shaped activities continuously. A CFO is not distinguished by the number of finance-related actions they can complete before lunch. A marketer who produces twice as many campaigns is not necessarily twice as useful.

Organisations do not generally suffer from a shortage of activity.

They suffer from a shortage of good judgement.

And good judgement is selective.

It is knowing what deserves attention and what does not. It is asking the right question rather than producing twenty answers to the wrong one. It is understanding when to intervene, when to wait, when to ignore something and when a seemingly small detail matters rather more than everyone assumes.

Replacing employees with supervision routines is an odd definition of progress

This is where the current harness model becomes less convincing.

Why is our time necessarily better spent supervising AI? We verify its work, correct it, teach it how we prefer things to be done and repeat the process when the result is unsatisfactory. We absorb a considerable amount of output, much of it competent, some of it useful and a great deal of it rather ordinary.

The employee has supposedly disappeared, but a new management routine has quietly appeared in their place.

That seems an odd definition of progress.

There are many criticisms one could make of this category of software. One is that reducing a person to their visible work product is an impoverished view of what they contribute. But I am more interested in the software-design problem.

Many of these systems encourage intellectual laziness.

They make it remarkably easy to avoid understanding the work.

Before automating a role, prove that you understand it

So I would suggest a simple test for any leader considering one of these systems.

Can you clearly explain what the person in that role actually does?

Not their job description. Not the list of tasks in the handbook.

What do they really do?

Who do they speak to? What information do they rely on? Which decisions matter most? What made their best work unusually good? What does an excellent six months in that role look like? What mistakes are expensive? What should they deliberately not spend time on?

And here is the more demanding version of the test:

Can you write a genuinely insightful essay answering those questions without AI?

If you cannot, I would be slightly concerned about your readiness to automate the role.

These are not generic research questions. They are questions about your organisation, your people, your incentives and your standards. If the software vendor understands the role only through prompts generated from other prompts, one should perhaps resist giving it responsibility for redesigning the work.

A harness should improve your abilities, not excuse you from developing them

None of this is an argument against harnesses.

Used with some discipline, they are excellent tools.

They can organise your thinking, retrieve things you once found useful, make connections across projects, help with targeted research, reformat work you already understand and automate bounded tasks where the objective is clear.

That is all valuable.

But the distinction matters.

A harness should improve your abilities, not provide an excuse to stop developing them.

There is a difference between using AI to extend your thinking and using it to avoid thinking altogether.

Good software has an opinion about how work should be done

Good software has always contained an opinion about how work ought to be done.

That opinion may be modest or ambitious, but it has to exist.

A serious software product must confront questions such as: Which parts of this work should be automated? Which should remain under human control? What information should be surfaced? What can safely be ignored? How should quality be measured? How can strategically important work be improved without creating problems elsewhere?

The answer cannot simply be:

Generate something. Check it. Try again if dissatisfied.

That is not much of an operating model.

Make the existing employee materially better at the job

If an AI system claims it can replace an employee for a monthly fee, I would expect it first to demonstrate that it can make the existing employee materially better at the job.

That seems a more sensible standard.

Do not give me an artificial CFO. Give the CFO better information, better visibility and better judgement.

Do not give me a synthetic product manager producing product-management-shaped material all day. Give the actual product team a system that helps them notice what matters, preserve context, reduce administrative work and make better decisions.

We have an unusual opportunity to improve enterprise software after two decades in which much of it became cumbersome, generic and strangely detached from the work it was meant to support.

AI gives us the ability to build systems that are more adaptive, more contextual and more capable of helping people do difficult work well.

Reducing that possibility to an employee in a chat window, or to a system in which people spend their time validating an endless stream of machine output, seems an astonishingly narrow use of the technology.

We should expect rather more from the software.

Chat with us 1-on-1

Search

Search posts and videos across the site.