Building Agents? Stop Treating messages[] Like a Database
Stop using messages as your agent's memory. Learn how structured state makes AI agents more reliable, efficient, and production-ready.

I hear a version of this in almost every conversation right now:
Someone on the buying committee saw a prototype get built in an afternoon—their own team, a demo, doesn't matter whose—and now they're asking why the number on our proposal doesn't look more like that.
Fair reaction. Wrong yardstick.
There’s a part of the conversation being skipped over when that happens, and it's the part I spend most of my time walking boards and steering committees through now.
Two things have changed now that agentic AI is in the picture:
That second part hasn't moved nearly as fast, and unfortunately it’s the part that decides whether you can actually replace a vendor with something your team built, not just whether you can stand up a working version once.
In our latest report, The Software Buyer's Guide to Agentic AI, my colleague Everett Zufelt and Orium’s VP of Agentic Systems, put it well: "The prototype is never the hard part anymore. The hard part is who's on call for it in eighteen months, and whether the team that built it fast is the same team responsible for keeping it running."
I've watched teams skip past that question entirely because the prototype looked good on a laptop. And I've watched the opposite mistake too: teams that refuse to even test the build path because "we already decided to buy," which just moves the same disagreement to month six of implementation instead of week two of evaluation.
That's just the build side. The buy side has its own problem, stemming from the way most RFP processes were built.
A demo used to prove something. One clean run would give you a decent read on the product, and then you moved on to the next vendor for the same thing. Agentic AI doesn't hold still in that way. You’ll still operate with the same inputs, but now there can be a completely different output each time, because the system is reasoning through the problem each time instead of following a fixed path. A demo still tells you what happened, but just this one time, and certainly not what it'll do in six months, against your actual data, under your actual constraints.
In the report, Preseetha Pettigrew, VP Global Partnerships at Contentstack, notes that she sees this constantly: "The hardest question we get from buyers isn't whether the AI can do something. It's whether it will still be doing the right thing, in the buyer's own brand voice, three releases from now. That's a different kind of proof than a demo can offer."
Then there's the part almost nobody's evaluation is scoped for, and it's the one that surprises procurement teams the most when we walk them through it. Every vendor on your shortlist is shipping some kind of AI capability right now— whether you asked for it or not, whether it's in your requirements document or not. And that AI capability inevitably reaches your data. It has a cost attached to it, whether or not anyone tells you what it is up front. That capability needs its own review, and it needs to be separate from the platform decision, not bundled into it and waved through.
I liked what Dirk Hoerig, Founder and Chief Innovation Officer at commercetools, said about this in the report. If a vendor can't show you the extension points running live, against your own data, that's the tell. As he put it, "The commitment has to be an open architecture a buyer's own team can build on, not a roadmap promise."
The questions I'd tell any buying committee to sit with before the first vendor call:
Those questions separate a smooth evaluation from one that reopens itself after signature.
We put a full guide together on this with MACH Alliance, commercetools, and Contentstack that covers what's actually new about evaluating software right now, what you owe your own organization before the first vendor call, and what your vendors owe you that most aren't being asked to provide yet.
If your board's asking the build question, or your evaluation hasn't caught up to what you're actually assessing, it's worth twenty minutes.
Stop using messages as your agent's memory. Learn how structured state makes AI agents more reliable, efficient, and production-ready.
Traditional approaches to change management weren’t working before. AI just makes the gaps impossible to ignore.
How smart companies are evolving with agent-powered delivery models, and what it takes to lead in the new era of intelligent services.