The AI models are the easy part

Getting an AI model to do something impressive takes minutes these days.
A support assistant answers questions faster than a human would. An internal tool summarises in seconds what used to take someone hours to find. Product teams ship things that would have been unreasonable to attempt a few years ago, and the first version often works the same day.
The most recent and relevant example of this last one is Asana's migration away from Enzyme, using Codex. OpenAI claims the work would have taken 5 years pre-AI, which is extremely questionable, though that's beside the point in fairness. Bottom line, something that Asana considered expensive and painful was done in 2 weeks by AI, cheap enough for them to be happy with it.
That is the part of AI everyone can see, whereas the harder part is after the demo when you still have to turn it into proper software.
AI doesn't replace the system around it
Lately there seems to be a lot of temptation (and sadly reality) where people treat the model as the product. User asks something, model thinks, model responds, done.
Most real world software doesn't work like that. It has users, permissions, and real data that has to come from somewhere. Every action has very specific consequences, some things need to happen immediately, some can happen later, and some absolutely must not happen twice. In software, we call this deterministic behaviour.
In contrast to AI, which is software with a probabilistic component inside it. Sometimes that component is incredibly powerful, like it can understand an email, simplify and explain something messy, pull information out of a long/boring document, write code, reason about a request, or decide which tool to use. That in itself doesn't mean that AI should be responsible for everything around it.
If I already know that customer X has three unpaid invoices, I don't need GPT-5.6 to work that out for me. A simple database query tells me that already. Whereas if I need to decide whether an email is a cancellation request, an angry complaint or somebody asking for a copy of an invoice, that's a very good place for a model. The engineering judgement is knowing where the line between the two sits.
Decide what doesn't need a model
We're increasingly seeing a pattern emerge as models become easier to use where everything starts getting pushed through it. A perfectly deterministic step becomes another instruction in the system prompt. Check whether the customer has an active subscription, work out which plan they're on, decide whether they're allowed to do this.
Your app already knows all three answers, and you've just swapped something fast, cheap and deterministic for something slower, metered and capable of being wrong. The opposite mistake is real too, where trying to handle genuinely messy human input with 400 `if` statements isn't clever engineering either.
The best AI systems are hybrid, and use normal software vs models where they each shine respectively. The fact that AI can do something does not mean it should.
The demo has no idea what happens when things go wrong
The happy path is very easy to build now, which is one reason AI demos are so seductive. Give the model the right input under the right conditions and it just feels like magic! But in reality, when an app is running in production, and the model takes 40 seconds instead of 4 seconds, what happens then?
When the provider returns an error halfway through a workflow, what happens then? The user clicks the button 3x more because nothing is happening, and you've just run the same £2 operation 4 times in a row. The model confidently returns malformed data, or calls the right tool with the wrong arguments, or produces an answer that is technically valid, yet completely useless.
These are everyday software problems, and they have software answers: retries, timeouts, queues, validation, idempotency, fallbacks, permissions, logs and sensible UX around failure. None of that makes for an exciting demo, and it is also exactly the difference between a demo and a product.
The worst failures are those that return 200 OK
Traditional software has a useful habit of breaking loudly. Something throws an exception, a request fails, an alert fires, someone swears at Datadog.
AI fails beautifully instead. The API returns successfully, latency looks normal and nothing shows up in the logs. Instead the answer is just bad. Maybe a prompt changed, maybe the wrong data was retrieved, or the most likely, your customers started giving it inputs you didn't anticipate. The system can happily keep returning `200 OK` while becoming less useful every day.
So monitoring AI works differently. You still care about uptime and latency, but you have to care about output quality as well. Is it answering the right question? Did it extract the right information? Would a human agree with the decision? Are users correcting it more often than they did last month?
Answering those means testing real outputs, keeping examples of good and bad behaviour, evaluating changes properly and actually looking at what the system does in production. Otherwise your first quality alert is a customer emailing to say the thing has been wrong for two weeks.
Then there is the small matter of the bill
AI is often hilariously cheap at prototype scale. You build something useful, run it 50x and spend £1.84.
Then real people start using it. 1k users slowly become 10k users, context gets larger, you add another model call because the first one wasn't quite reliable enough, agents start calling tools, failed jobs retry, you name it. By then the economics look different.
That doesn't quite make AI expensive though, because it's still absurdly cheap compared with the human work it replaces. Instead it makes cost an architectural concern, such that you should know roughly what an important workflow costs, what happens to that number at 10x the volume, and what you're getting back for the money. If an AI feature costs £4k/mo and removes £30k of work, nobody sensible should care about the £4k spend. But if it costs £4k/mo and nobody can tell whether customers would notice you switching it off, you have a different problem.
The fix is boring though. Use caching, minimise model calls, reduce context where possible, use a smaller model and most importantly use code instead of inference when the answer is already known. The model won't be offended!
Good engineers matter more than ever
There's an understandable question hanging over software engineers at the moment, which is "if AI can write the code, what is my job?"
The answer is pretty much what it always was, which is everything around the code. Deciding what should be built. Understanding the existing system before changing it. Working out what AI should and shouldn't be responsible for. Designing the interfaces between the deterministic and probabilistic parts of the product. Making sensible security decisions. Knowing what happens when something fails. Keeping the system understandable enough that somebody can change it six months later. Working out whether an impressive technical capability solves a problem anybody actually has. And, unfortunately the most painful one at the moment, reviewing what the agents produce.
Asana's Enzyme migration is a good example of that too. Engineers decided what needed to change, gave the agents an environment to work in, checked their progress and reviewed the proposed changes. AI moved the implementation boundary dramatically, but engineering judgement mattered more than ever.
This is where some of the "AI replaces software engineers" conversation gets muddled. The human effort required to turn an idea into code is collapsing, but code was never the whole job.
Let's build boring systems around clever models
Models will keep getting better! Six months from now they'll do things that are awkward today, and six months after that we'll move the line again. Predicting exactly where it stops is a waste of time. What you can do is build systems that are comfortable with the change.
Keep model providers replaceable where you can, keep business logic outside the prompt and make model outputs structured and validate them before trusting them. Measure quality, cost and more than ever, know who owns the feature after it launches.
Basically, make the system around the model boring. Boring is good, because the clever part will change every three months.
How we think about AI at Celestify
We are extremely bullish about what AI makes possible. We use it constantly, we build with it, our engineers use coding agents, and we put models into products and internal workflows where they remove crazy amounts of manual work.
We're less interested in adding AI to something because somebody wants an AI feature on the roadmap. The question we start with is "what problem are we solving? which pain are we removing?" Sometimes the answer involves an agent, sometimes it involves retrieval, a model and three integrations, sometimes it involves a single skill.md file. Sometimes, it doesn't need a model at all.
Making the impressive bit work is getting easier every month. Making the whole thing useful, reliable, secure and economically sensible is still the job.
If you've got an AI feature in that grey area and can't tell which side of the line it sits on, it's worth working that out before you build it and we're happy to help you think it through.
