A job candidate this week can read a guide on how to prepare for an interview with a bot. The guide exists because the bot exists. Once enough candidates read it, the bot will be scoring people on how well they prepared for the bot.
That’s the whole problem in miniature, and it isn’t limited to hiring. Any system that judges people by a rule teaches the people being judged the rule. The rule stops measuring what it was built to measure and starts measuring how fast the crowd learned it.
The Interview That Grades Itself
An automated interviewer needs a rubric. Maybe it listens for structure, keywords, or a certain pace of speech. The rubric can’t stay secret for long, because candidates compare notes and career sites publish advice, and now there’s a new genre of prep article for exactly this. Within a season, most candidates give the answers the rubric rewards.
The hiring team then sees a pool of strong-looking interviews and can’t tell them apart. The bot did its job perfectly, and that’s the trouble. It found the people who are good at bot interviews. Whether that overlaps with the people who are good at the work is now an open question, and nobody can settle it from inside the loop.
I’m not pessimistic about this, for what it’s worth. Screening at scale is a real problem, and a consistent first pass is fairer than a tired recruiter on their ninth call of the day. The tool works. It just works on a clock, and the clock starts when people learn how it works.
Fraud Models and Last Year’s War
The same clock runs faster in payments. A fraud model learns from what fraud looked like, so it’s always a portrait of the previous attacker. The next attacker has watched which transactions got declined and which sailed through, and has adjusted. Every decline is a small lesson handed to the other side.
There’s a line about strategy I keep turning over. If strategy were a science whose principles could be learned, then all the belligerents would learn them, and the result would be stalemate or attrition. Fraud teams know this feeling. You ship a better model, the loss rate drops for a quarter, and then it climbs back to roughly where it was, because the adversary is a learning system too. You didn’t win. You raised the price of the game for both players.
Nothing in that is a failure of the model. A model that scores well against last year’s fraud is still doing real work; it stops a large share of the easy attacks, and that has value. The mistake is treating it as a fence when it’s closer to a treadmill. A fence stays put, while a treadmill only counts as progress if you keep running.
Changing the Board
The more interesting move is the one that stops playing the game. Retailers figured out a version of this in the card world. When interchange fees got expensive, some pushed customers toward cheaper debit routes, which is still playing on the network’s board. Others issued their own private label cards, which have no network affiliation and only work at that retailer’s own stores.
That’s a different kind of answer. A closed card doesn’t out-score the network’s fraud models. It removes most of the surface those models were defending. The retailer knows the customer, controls both ends of the transaction, and sets its own rules, so there’s far less room for an outsider to probe the rubric. It gives up reach in exchange for control, and for a lot of retailers that trade has been worth it for decades.
I think about this when I read about data that lives somewhere it can be reached. The Guardian’s report on sensitive UK police data and its exposure is a headline about who can get to the data, and that’s an architecture question before it’s anything else. You can promise good behavior with policy, or you can build the system so the reach isn’t there. Policy is a rule that everyone learns; structure is a place to stand that the rule doesn’t touch. I don’t know the technical details beyond the reporting, so I’ll leave the specifics to people who do. The shape of the choice is clear enough.
Unfairness Is a Design Input
Some of this is simply unfair, and it’s worth saying so plainly. A defender has to be right every time and an attacker has to be right once. A candidate can rehearse a hundred interviews, while the bot never gets a second look at the person it just rejected. Life hands out those asymmetries without asking anyone, and sometimes they land in your favor.
The useful response isn’t to resent the asymmetry, and it isn’t to pretend it away either. It’s to design with it in view. If your opponent will learn every visible rule, then visible rules should carry as little weight as possible. That points toward a few habits that hold up better than a smarter score. Mix in signals the other side can’t see or cheaply fake. Put a human in the loop at the point where judgment matters most. Rotate the test before it hardens into a script. Keep the fast automated pass for what it does well and stop asking it to be the final word.
None of that is exotic. Most of it is already practiced by the better fraud teams and the more careful recruiters. What’s changing is how quickly the learning happens: when everyone has a model that can read the rubric and rehearse against it, the treadmill speeds up for everyone at once.
What a Score Is Worth
A score is a snapshot taken before anyone knew they were being photographed. The moment people learn the pose, they hold it, and the picture keeps getting sharper while it tells you less. So the question for any bot, model, or metric is how long its signal survives contact with the people it measures.
The systems that last won’t be the ones with the cleverest rubric. They’ll be the ones built so that winning the game and doing the job are still the same act. Once everyone has learned the rules, a better move stops mattering; only a different board does.

Leave a Reply