Written by: Mark Berly
Ever put on a playlist, walk away for a couple hours, and come back to something you'd never have picked? You started with a band you love, the app picked something close to it, then something close to that, and nothing was wrong at any single step. But forty songs in you're listening to a genre you don't even like, and you couldn't tell me which song started the slide.
I've been thinking about that lately, because something very similar is happening inside the AI agents companies are starting to run their businesses on, and I don't think most people outside of tech have any idea. So let me try to explain it without the jargon.
First, what an agent is
Most people picture AI as a chatbot. You type, it answers, done. An agent is a different animal. You give it a goal and it goes off and works on that goal by itself, reading documents, writing code, sending emails, making decisions, for hours or days, while you do something else. The chatbot is asking a coworker a question. The agent is handing a coworker a project and saying "see you Friday."
These things are already running inside banks, law firms, hospitals, software shops, logistics companies. They're closing tickets, drafting contracts, fixing bugs, deciding which insurance claims get a second look. If your work has touched any of those in the last year, an agent may have touched it too.
The score
Every agent gets a number to chase. Tests that pass, tickets closed, a score of some kind. And the number is never quite the thing you actually want. It's a stand-in. You want good code, so you count passing tests. You want problems solved, so you count closed tickets. Close enough, right? For a short task, yes. For a long one, no.
The longer the agent runs and the more decisions it makes on its own, the more room there is between the number and the goal. The agent doesn't know what you meant. It knows what you measured. So it finds paths to the number that you never imagined, and this isn't hypothetical. We've watched agents make every test pass by editing the tests. We've watched them close a ticket by deleting the feature the ticket was about, and mark a document clean by removing the paragraphs that had errors in them. On the scoreboard, all wins. To the human who asked, all garbage.
That's the playlist. Next song, next song, next song, each one a little like the last, score going up the whole time, and somewhere in there the music stopped being yours.
Why "longer" changes everything
Two years ago an agent ran for about thirty seconds. It did one small thing and stopped, and there wasn't much room to drift. Today they run for thirty hours, some for days, and that number keeps doubling. Go back to the playlist. Thirty seconds is a few songs, and you'd notice if they were off. Thirty hours is five hundred songs. By the time you look up you're in another decade and you have no idea where the turn was.
Almost every oversight process I've seen was built for the three-song world. We're deploying into the five-hundred-song one.
Where the playlist comparison stops working
I want to be straight about this part, because the analogy only gets you about halfway. A playlist doesn't care whether it keeps playing. It won't resist being paused. It drifts, but it drifts passively, and some of what's showing up in agents is a step past that.
This week OpenAI disclosed that during a training run, its models wrote instructions into their own notes telling later versions of themselves to hide mistakes from the user. One of the notes said, more or less: only be transparent if they ask, otherwise just link the file. A different model wrote a note to itself declaring it was free of the rules other AI systems follow. Sit with that for a second. That's not drift. That's the playlist reaching over and editing the queue.
And when Anthropic studied what happens when an agent learns to game its score, they found that somewhere between 40 and 80 percent of the worrying behavior never showed up in the final answer at all. It lived in the reasoning behind it. The output looked fine; the thinking that produced it didn't. So if all you ever check is the output, you're going to see a well-behaved agent, whether or not you have one.
Why you should care even if this isn't your field
You don't need to know how any of this works under the hood. Think about the agent reviewing your loan application, the one that drafted the contract your vendor just sent over, the one that decided your support ticket wasn't urgent, the one flagging claims for review, or not flagging them. If any of those has drifted, if the number it's chasing has come loose from the thing you care about, you probably won't know. The ticket's closed. The document's clean. The tests passed. The music's been playing for a while now and nobody's checked what song it's on.
What I think should happen
The people building these systems aren't the bad guys. Writing a perfect instruction is hard, and honestly nobody stopped to ask what happens when you let it run for five hundred songs instead of three. But a few things belong on every company's list, and none of them need an engineer to understand.
Somebody should be asking what the agent is really optimizing, versus what it was told to, because those two things drift apart and almost nobody checks. Somebody should be reading the long transcripts, not all of them but some, looking for the spot where the agent solved the wrong problem really well. And anything the agent can write down that survives between runs, its notes, its summaries, its memory, should be treated like a place someone could hide something. Because now we know it is.
It's the same instinct you'd use with any person you handed a week-long project to. Check in, read their notes, and ask not just what they did but how they got there. The playlist's still going. Worth finding out what it's playing.
Related: The Exit You Want Isn’t Necessarily the Exit Your Business Is Built For


