17 minute read

The Problem Nobody Set

What twenty-two players and a dancer alone on a stage are doing that the objective doesn’t describe, and why we have never counted it.


There is a piece of work your body performs every waking second, and you will almost certainly never think about it, unless one day it stops.

I injured my lumbar spine. What I learned afterward was not really about pain, which is the part everyone expects and the part I will spare you. It was about administration. Somewhere well below anything I could call thinking, my brain had been issuing a continuous stream of instructions to the small stabilizing muscles around my spine, adjusting them constantly, correcting for every shift of weight and every turn of my head, every hour I had been alive. Standing still turns out not to be still. It is a controlled fall, caught and re-caught, in every direction at once.

I never chose that work. I never noticed it. I became aware of it only when it started arriving late and imprecise, and the ordinary business of getting up from a chair became something I had to attend to deliberately, the way you attend to a difficult sentence.

Here is the part that took me longer to see. The work was not executing a standing instruction. Nothing in me was running an order that said stay upright. At every instant something was deciding which of an enormous number of signals mattered right now, and therefore what the problem currently was: a floor that turned out to be wet, a bag heavier on one side than expected, a step that arrived sooner than the last one. The instruction never changed. The problem changed constantly, and something reformulated it, continuously, without asking me.

Nobody set that problem. Something I have never met noticed it and re-noticed it thousands of times an hour.

The Ledger We Keep Of Difficulty

Hans Moravec noticed the consequence of this in 1988, in a book called Mind Children, and the observation now carries his name. It is comparatively easy, he wrote, to make computers perform at adult level on intelligence tests or at checkers, and difficult or impossible to give them the skills of a one-year-old at perception and mobility.

The paradox is normally presented as a fact about machines. I want to present it as a fact about us.

We rank difficulty by how hard a thing feels while we do it. Chess feels like intellect. It is effortful, available to introspection, and we can watch ourselves doing it, so it sits at the top of the ledger. Catching a set of keys tossed across a room feels like nothing at all, so it sits nowhere. The ranking is not a judgment about the problems. It is a report on which problems we happen to be conscious of solving.

Let me be careful about how much weight that carries. I am not claiming effort measures unfinished learning. Expert musicians work hard at passages they have played for decades, vigilance is exhausting after thirty years of practice, and creative work stays effortful precisely because there is no stable routine to settle into. The claim is narrower: felt effort is an unreliable proxy for computational difficulty, and we have been using it as our main one.

An objection deserves to arrive here rather than later. The computer scientist Arvind Narayanan has argued that Moravec’s paradox is not a dependable predictor of what will prove easy or hard for machines, and is better read as a statement about what the field has found worth working on. I think that is largely right, and I am not going to use the paradox as a forecast. Whatever it does or does not predict, it correctly identifies a bias in our accounting.

That informal ledger has a formal descendant, and it now runs the industry. We call them benchmarks. A benchmark is our difficulty ledger written down and made scoreable, and it inherits the same property: it can only contain what somebody thought to put in it.

The Objective Was Never the Hard Part

Let me concede the thing that arguments of this kind usually avoid.

Football has an exceptionally clean objective. Score more goals than the opposition inside the allotted time. You can write it on a napkin. There is no ambiguity about success, and a scoreboard settles it. If difficulty lived in the difficulty of stating an objective, football would be among the simplest things humans do.

So the difficulty is not in the objective. It is in the distance between having one and knowing what to do next.

Watch what that distance contains. Twenty-two people are running the stabilization work I described at the start of this essay, continuously, while sprinting, turning, being shoved and landing badly. Air moves across the stadium in ways nobody has modeled. The surface changes through the match. The light shifts. The altitude sets what their bodies can spend. And each player is reading the intentions of everyone else on the field while being read in return. A pass is a prediction about where a teammate will be in two seconds, made by a body in mid-fall, against defenders making their own predictions about that prediction, who will change what they do the moment they detect it.

I want to be precise about what the players are doing, because there is a mystical version of this claim and I do not want it. They are solving the problem. Obviously they are; the ball arrives. What they cannot do is say how. The computation is distributed through the body, learned over years, and closed to introspection, which is different from absent.

But the deepest part is not the policy either. It is that the problem does not hold still.

Somewhere around the hour mark, a match becomes a different match. The referee has been letting contact go, so the cost of a challenge has changed. Someone is carrying a knock and is now a target. The other side has gone to three at the back, and every space on the pitch means something new. Nobody announces any of this. No player is told the problem has changed. Each one works it out while still playing, revises what matters, and acts on the revision before being able to articulate it.

It is worth being exact about what changed, because adapting to a moving situation is something good control systems already do well, and I do not want to dress that up as something rarer than it is. What changed was not the difficulty. It was which information mattered to success. A defender’s body position meant one thing at twenty minutes and means something else now, and nothing in the original instruction said so, or could have.

That is the line I want to hold throughout this essay. Adjusting your behavior because conditions moved is one thing. Discovering that a different set of facts has become the relevant ones is another, and it is the second that nobody set.

The objective did not move. It was to score more goals in the first minute, and it is to score more goals in the ninetieth. What moved was everything about what that requires.

Take Away The Opponent

It would be easy to conclude this all comes from the adversary, so it is worth removing the adversary and seeing what remains.

Ballet has no opponent in the relevant sense. Nobody on that stage is trying to make a dancer fall.

I should concede at once that she faces an evaluator. Two thousand people are watching and judging, and their judgment is opaque, plural and impossible to reduce to a number. That is real pressure, and I have not removed all of it. What I have removed is the thing football is usually assumed to be hard because of.

And the difficulty does not drop.

She is running the same continuous stabilization work at the limit of what a skeleton tolerates, on a floor with its own particular give, in time with music played live and therefore slightly differently tonight, alongside other dancers whose small variations she absorbs without appearing to absorb anything. And through it all, she is deciding what tonight’s performance is. Not the choreography, which is fixed and which she knew months ago. Suppose her partner arrives a fraction early and the lift no longer sits where the phrase wants it. She has perhaps half a second, and the choice is not how to execute the lift better. It is whether tonight the moment is the lift at all, or the recovery out of it, and she has to decide which one the audience is now watching before she can have any view about how to do it well. Nothing in the choreography addresses that, because the choreography could not have known her partner would be early.

The choreography is the specification. It is complete, and it is not the performance.

Before I go on, a correction to an impression these two examples will have created. I chose a World Cup final and a stage because they are legible, not because the capacity is rare. Nothing in the argument depends on excellence. The same continuous re-noticing runs in an aide steadying someone out of a chair, whose sense of what this particular person can bear today was never written into any care plan, and in a warehouse picker rerouting around an aisle that is blocked differently than it was an hour ago. The elite cases are the ones we buy tickets for. They are not the ones where the capacity lives, and if anything, the fact that we pay handsomely for the visible version and almost nothing for the ordinary one is its own small piece of evidence about the ledger.

What A Measurement Can See

Take drone light shows, which are the cleanest case available. Thousands of aircraft, physically present, holding formation in real wind. Every objection about digital versus physical dies on contact with a drone show.

The lazy version of this argument is wrong, so I won’t make it. Those aircraft are not blindly following fixed paths. Modern autonomous systems execute policies rather than trajectories: wind arrives, and the aircraft recomputes what to do to stay where it needs to be. That is real adaptation and real engineering.

Notice what is nonetheless fixed. The image was decided before anything left the ground. Nothing in that system is built to ask whether the picture is still the right picture, or whether tonight calls for something other than what was planned. Adaptation of the route, thoroughly solved. Reformation of the problem, not attempted, and not required. The football video games and the Stanford autonomous helicopter sit in the same place for the same reason: one with authored physics and one with an objective inferred from expert demonstrations rather than written by hand, which means the specification arrived from a different direction and was no less complete.

Here is what these cases have in common, and it is not that the systems are unimpressive. Each one was measured by something built to see a particular thing, and each one succeeded at that thing. The measurement did not fail. It reported honestly on what it was designed to report on, and said much less about anything else, which is what measurements do.

The Machines That Actually Play

Now the honest part, because there is a serious effort aimed directly at this territory.

RoboCup has been running since 1997 with an openly stated goal: a team of humanoid robots capable of beating the human World Cup champions by 2050. This summer, at the competition in Incheon, two full teams of humanoid robots played an eleven-a-side match on real hardware for the first time, with B-Human of Bremen beating HTWK Robots of Leipzig four to nil. In the tournament proper, which ran across three divisions this year, B-Human took the middle division, a Tsinghua University side took the large division, and a team from Wuhan University took the small one. Observers reported clear gains in stability, walking speed, and coordination.

This is not choreography. Those robots are perceiving, deciding, and coordinating in real time against opponents doing the same. They are learning policies rather than following scripts.

And that fact demolishes an argument I am not going to make. It would be easy to claim that because we cannot write down what a footballer does, no machine can acquire it. That claim is false, and the last fifteen years have falsified it repeatedly. Nobody wrote down a specification for recognizing a face or holding a gait, and systems learned both anyway. Humans never had a written specification either. We learned from interaction, which is what these robots are now doing. Being unable to describe a capability tells you nothing about whether it can be acquired.

So the interesting question is not whether machines will get good at this. It is what their getting good at it entitles us to conclude.

Because a match at Incheon is a measurement, and like every measurement, it was built to see certain things. It sees whether a robot can stand, perceive, pass, and score against opposition. Those are real capabilities, and it reports on them honestly. What it is not built to see is whether anything on that pitch could notice that the problem had changed. Nothing in the tournament asks that question, so nothing in the result answers it.

And look at where the noticing is actually happening, because I find this the most interesting fact in the whole competition. For nearly thirty years the humans running RoboCup have been steadily reformulating the problem: changing rules, removing crutches, merging leagues, making the environment less friendly on a deliberate schedule. This year the humanoid league and the standard platform league were merged into one. Nobody was optimizing toward that. Somebody decided the problem had become a different problem.

That is problem formation, performed by an institution rather than a person, which is worth saying plainly because it means the capacity I am describing is not a special property of human bodies. It is a property of certain systems. A governing body can do it. A committee can do it, slowly, in minutes and votes. It is simply not the thing being measured on the pitch, and it never has been.

Attention Is Problem Formation

Which returns me to the ground this series has been standing on.

William James wrote in 1890 that my experience is what I agree to attend to. He was not making a point about focus or productivity. He was saying that attention is the act by which a mind decides what its situation consists of, and that the deciding comes before the experiencing.

I should be honest that I am extending him past what he claimed. James was describing conscious attention, and most of what I have been describing never reaches consciousness at all. What I am borrowing is the structure rather than the phenomenology: something selects what the situation consists of, and that selection is prior to everything downstream of it. Whether that something is available to introspection is a separate question, and in these cases the answer is mostly no. Read the structure next to a player at the hour mark, and it stops sounding like philosophy. Deciding what matters right now is not preparation for solving the problem. It is most of what solving the problem consists of.

That capacity develops the way this series keeps describing. Consequence returns to the actor, on the same continuous self, without delay. An infant reaches and misses, pulls itself up and goes down, misjudges a gap and pays for it immediately, and the cost lands on the body that made the error. Run that loop enough times, and the competence arrives and then leaves awareness. What gets delegated downward is not only the execution. It is the judgment about what matters, which is why you can cross a room with a cat in it while thinking about something else.

An objection I should meet, since I am doing what I claim cannot be done. This essay describes an automatic competence at some length, so is it really unavailable? Describing that something happens is not the same as specifying how, and the gap between those two is the whole subject. I can tell you my spine was being stabilized. I could not have told you, at any instant, what it decided mattered.

And I should be equally plain about the limits of what I have offered. I have named a capacity and given it a criterion. I have not explained it. I cannot tell you what computation performs the re-noticing, what architecture would support it, or how it differs in kind from very good context sensitivity, and I am wary of anyone who says they can. That is not evasion, or at least I hope it is not. It is the argument: this is a capacity we have not characterized well enough to build a test for, which is precisely why our tests are quiet about it.

How Narrow The Claim Is

This is not a bet against the trend. The trend in robot football is real; it has held for nearly thirty years, and I have no argument that it will break. If anything, I expect it to continue.

The claim is about what the trend measures. Every benchmark is an architecture, and its architecture decides which capabilities it can see clearly. Success inside one is direct evidence about what it was designed to measure. Anything it appears to say about capabilities outside that design is indirect, and has to be argued for rather than assumed. That is not a strange evidentiary standard. It is the ordinary one, and we suspend it constantly.

Discovering that a different set of facts has become the relevant ones is not what any of these competitions was built to measure. So thirty years of rising scores do not, by themselves, establish whether that capacity has improved. It may have. Something like it may already be present in systems trained across enough varied conditions, and if anyone wants to argue that, the argument is available and I would read it with interest. What I am objecting to is the skipped step: treating the score as though it had settled the question.

Here is what would change my mind, and I have tried to make it checkable rather than rhetorical. Show me that performance on a specified benchmark predicts how a system does when the relevant facts change, and specifically when success requires discovering that the original framing has stopped being adequate. Not adapting a route to a fixed target. Working out that the target was the wrong one and that something else now matters. If scores turn out to predict that, the two things were never separable, the inference I am calling unwarranted is warranted after all, and I will say so here.

The Question I Cannot Answer For You

Think about the last time you were partway through something and quietly realized it had become a different task than the one you started. Not that you were failing at it. That the thing worth doing had moved while you were working.

Nobody set you that problem. You noticed it mid-stride and reformulated what you were doing before you could have explained why. No test you have ever sat measured that.

I want to be careful about where this lands, because there is a warm reading of everything above that I do not intend and would rather refuse than allow. This is not an argument that something in us is unreachable, or sacred, or beyond machines by its nature. I have no idea whether it is beyond machines, and I have said so. The capacity is ordinary; it is physical; it is learned, and an organizing committee can do a slow version of it in meetings. Nothing here needs mystery to work. What it needs is for us to notice that we have been reading our scores as though they covered it.

So when the scores go up, and we say the machines are catching up, the question is not whether the scores are real. They are. The question is what our measurements were built to see, and what they have been quietly silent about the entire time.


Originally published on Substack.