Live ops after AI
What moved, what didn't, and what we stopped doing. The value sat in churn and LTV, not in the calendar or the NPC, and most studios chase the wrong 5%.
By the founder · Higher Agency
Day-seven retention dropped on a Monday, and it took until Thursday to say why. Someone pulled the cohort by hand, cross-referenced three dashboards, escalated to a PM, and by the time anyone knew what happened the event was half over. That is the live-ops week that never makes the pitch deck. Not the strategy. The fetching.
When a vendor pitches AI for live ops, it is usually one of two things. A player-facing system, an NPC that talks back or dialogue that writes itself. Or a button that fills your event calendar for you. We have run live ops for years. Both of those are the wrong place to start, and one of them is a trap.
What actually moved was the layer under the calendar. A model that flags a cohort sliding toward churn while you can still do something about it. Forecasting that changes which offer you bother to test at all. A segment-level read on why day seven fell, arriving Monday instead of Thursday. None of it decides anything. It puts the decision in front of a person days earlier, and in a live game those days are the entire margin.
What did not move is taste. The model ranks offers, and it does not know that the one it ranked first cheapens the game. It marks a retention cliff and stays silent on the actual cause, which is usually that the new event is boring. We watched a team hand the judgment call to the system and ship a technically optimal sale that played like a shakedown. The model had no idea it had done anything wrong; the players spotted it in a day.
So we stopped automating the deciding and automated the fetching. The Monday report writes itself now. The three-dashboard cross-reference is one query. The PM who lost half a week to pulling numbers spends that week reading them instead. Nobody's judgment got replaced. Their morning did.
The tempting place to put AI in a game is the player-facing surface, the NPC and the generated dialogue. GDC's 2026 survey found 5% of the developers already using these tools work on player-facing features, against 81% on research and brainstorming. That gap is not an accident. Player-facing is the hardest thing to ship and the slowest to pay for itself, and when it goes wrong it goes wrong in public. The value was sitting in the unglamorous 81% the whole time.
An economy agent is only as trustworthy as what surrounds it. Turn one loose on a live economy with no backtest against events that actually shipped and no hard gate on the economy rules, and sooner or later it ships a sale that breaks the game. The live-ops platform we run for a top-20 mobile studio simulates every event against real player segments, checks it against the hard economy rules, and backtests it against events that shipped, under 490 automated tests. That is the difference between an agent you can leave running and one you have to stand over.
If your live game acquires players well and then leaks them somewhere your dashboards will not explain, the first move is not a model. It is finding the leak. That is what the Product & Growth Audit does: one month, an operator maps where retention and the economy are actually losing money, and hands you build-ready specs. If the fix turns out to be an agent that watches the leak so a person does not have to, the AI Build Program ships it and trains your team to run it.
Live games taught us that the treadmill is real and the judgment stays yours. We automate the fetch, not the feel. If you cannot remember the last time your live-ops read landed early enough to act on, that is the conversation we would start with.